Prompt trials
Measure a proposed stage prompt against 6 to 12 real tasks before you promote it, read the guardrails, and roll it back.
A retrospective can propose a change to a stage’s prompt. Applying that change by hand is one option, and it works. A prompt trial is the other option: it measures the proposed prompt against the one running today, on real finished work, before you commit to it.
A prompt trial is a different tool from a trial run. A prompt trial measures one proposed prompt against the one live today and can promote it once it clears a fixed bar. A trial run replays a task set under a different harness, model, or pinned stage versions, and leaves the decision entirely to you.
A trial takes a set of tasks your workspace already finished, and replays each one twice — once on the current prompt, once on the proposed one — under otherwise identical conditions. It pairs the two replays task by task, checks quality, time, and cost, and tells you whether the proposed prompt actually cleared a set bar. Nothing is promoted automatically: a person still decides.
What it takes to start one
- A recorded proposal whose change is a stage prompt — the kind Retrospectives can produce.
- A workspace owner or admin, signed in as a person. A trial can’t be started from an agent session.
- A leaf workspace. Trials don’t run in a group workspace.
- 6 to 12 finished tasks of the proposal’s task type, each with a real commit to replay from.
A trial’s two replays and their judge runs draw on the same account-wide optimization budget as a retrospective’s own analysis. If that budget is enabled and exhausted for the day, a replay or its judging simply waits rather than failing — see that page for how to check whether the budget is what’s holding a trial up.
Starting one
ch quality trials create --from-file trial.json
trial.json names the proposal, the stage and the version of it you’re measuring against, the launch config and model to run the replays on, the judge model, and the source tasks:
{
"proposalId": "019fb0d5-3c81-7e44-a2b9-5f0e7c31da68",
"type": "story",
"stage": "review",
"configId": "019f8b21-6e02-7a14-9c3d-1a2b3c4d5e6f",
"model": "claude-sonnet-5",
"judgeModel": "opus",
"baselineContentHash": "9f2c1d8a7b6e...",
"credentialConsent": { "secretIds": [], "secretSlots": [] },
"sources": [
{ "taskId": "019fe33a-05bc-7c18-9e6d-84f2a71b0d95", "baseSha": "a1b2c3d4e5f6..." }
]
}
type is the task type the proposal’s workflow belongs to (story, bug, and so on) — every source task must share it. stage is the exact stage instance the proposal targets, the same name Retrospectives’ ch quality optimization proposals prints as TARGET.
configId is the launch config both replays run on. That config has to be enabled, and it has to name the model you pass on its Allowed models list — an exact match, not a tier. This is stricter than the rest of the product: elsewhere an empty Allowed models list means “no restriction” (see Which model your agents run), but a trial refuses a config whose list is empty or doesn’t carry that model. If your config runs unrestricted today, add the model to its Allowed models before you create the trial.
judgeModel is fixed to opus — the same model CodeHerder’s own stage judging uses — and the request is rejected if it’s anything else.
baselineContentHash is the content hash of the stage as it stands today. It’s required whenever the stage is still one of CodeHerder’s built-ins, which is the case for any stage your workspace has never customised. Get the value from any of the source tasks: run ch workflow show <taskId> and read the stage_content_hash on the stage instance you’re trialling. For a stage your workspace owns, the field is optional — but if you send it, it still has to match.
Each source task needs the commit its work was actually built on; baseSha is that commit’s SHA.
credentialConsent names every CodeHerder secret the replays may use. Leave it out when the source launches used no secret. If a source launch used a secret, list its id in secretIds (or its device slot key in secretSlots). The list must match exactly: a secret you leave out, or one no source used, is refused with a message naming it. Only an admin can create a trial, and CodeHerder records your consent as an event that lists ids only.
Repeating the exact same request returns the trial you already started instead of starting a second one, so retrying a create that timed out is safe.
What a trial can reach
- CodeHerder: a trial can read its own frozen launch input, knowledge and skills. It cannot write a comment, a message or a task.
- CodeHerder secrets: none, unless you named them in
credentialConsent. - The AI provider: yes. A replay cannot run without it.
- The device: the replays run on a real device. They share its filesystem, its network and any git credentials it holds. CodeHerder does not isolate these.
Run trials on a device that holds nothing you would not give a replay.
Reading it
ch quality trials
lists every trial in the workspace, oldest first. Pass a trial’s ID to read it in full:
ch quality trials show <trialRef>
Both commands print JSON. show returns the frozen inputs the trial was created with, one entry per source task holding that task’s baseline and candidate measurements — state, session time, cost, and quality — plus whether the trial is eligible to promote right now, a reason for every guardrail it hasn’t cleared yet, and the limitations of the measurement itself. A missing measurement just means that replay hasn’t finished.
The guardrails
A trial only becomes eligible to promote once every one of these holds. This bar is deliberately high — a trial sitting short of eligible for a while is normal, not a bug:
- Every planned pair finishes, with a quality result and real, priced spend. A session’s cost isn’t counted until at least ten minutes after it exits, so a very recent pair can still be waiting on its bill.
- Both sides of every pair pass judging, and the candidate matches or beats its baseline’s quality level. A pair whose baseline failed judging can never make the trial eligible, however good the candidate was — so pick source tasks whose original work was sound.
- No candidate trips a blocking criterion. This one applies to the candidate side only.
- Everything the trial froze still matches — the worker’s launch config version and model, the judge model, and the rubric it grades against.
- Median session time improves by at least 10% on the candidate prompt, across the completed pairs.
- Cost doesn’t rise by more than 5%, summed across the completed pairs.
- The improvement clears a one-sided statistical test at p ≤ 0.05 — a tie between a pair’s baseline and candidate counts against the candidate, not for it.
- No batch in the trial was cancelled. A cancelled trial can never authorize a promotion, even if some of its pairs already finished cleanly.
Deciding
Recording an evaluation keeps a permanent snapshot of the trial’s result at that moment:
ch quality trials evaluate <trialRef>
Promoting applies the candidate prompt to the live stage — only if the trial is eligible:
ch quality trials promote <trialRef> --expected-version <n>
--expected-version names the stage version you reviewed the proposal against, the same number ch stage versions or the proposal’s own BASE VERSION shows. If someone has changed the stage since, promote refuses rather than overwrite their edit. If the trial isn’t eligible yet, promote records a hold instead of touching the stage.
Rolling back restores the prompt promote replaced:
ch quality trials rollback <trialRef> --expected-version <n>
Here --expected-version is the version number promote produced — not the version before it. Rollback only works while that promoted version is still the stage’s current one; if it’s been edited or promoted over again since, roll forward by hand instead.
Read the decisions recorded against a trial — the most recent ones, newest first — with:
ch quality trials decisions <trialRef>
Read this before you act on it
A trial measures one prompt, on one stage, against a small, deliberately bounded set of tasks. Clearing every guardrail tells you the candidate prompt did better on that set, under those conditions — it isn’t proof the change helps every kind of work your workspace runs. And nothing here promotes itself: an eligible trial is a green light, not an action. A person still runs promote.
Related guides
- Trial runs — replay a task set under a different harness, model, or pinned stage versions
- Retrospectives — where a stage-prompt proposal comes from
- The optimization budget — the shared cap a trial’s replays and judging draw on
- Stage judging — the same judge model, scoring day-to-day stage attempts
- Stage-attempt detail — read the individual attempts behind any stage’s numbers
- The stage library — stage versions, content hashes, and prompt history
Last updated