# Prompt trials

Source: https://codeherder.com/docs/prompt-trials/

Measure a proposed stage prompt against 6 to 12 real tasks before you promote it, read the guardrails, and roll it back.

A [retrospective](https://codeherder.com/docs/retrospectives/) can propose a change to a stage’s prompt. Applying that change by hand is one option, and it works. A **prompt trial** is the other option: it measures the proposed prompt against the one running today, on real finished work, before you commit to it.

A prompt trial is a different tool from a [trial run](https://codeherder.com/docs/trial-runs/). A prompt trial measures one proposed prompt against the one live today and can promote it once it clears a fixed bar. A trial run replays a task set under a different harness, model, or pinned stage versions, and leaves the decision entirely to you.

A trial takes a set of tasks your workspace already finished, and replays each one twice — once on the current prompt, once on the proposed one — under otherwise identical conditions. It pairs the two replays task by task, checks quality, time, and cost, and tells you whether the proposed prompt actually cleared a set bar. Nothing is promoted automatically: a person still decides.

## What it takes to start one

- A recorded proposal whose change is a **stage prompt** — the kind [Retrospectives](https://codeherder.com/docs/retrospectives/) can produce.
- A workspace **owner or admin**, signed in as a person. A trial can’t be started from an agent session.
- A **leaf workspace**. Trials don’t run in a group workspace.
- **6 to 12** finished tasks of the proposal’s task type, each with a real commit to replay from.

A trial’s two replays and their judge runs draw on the same account-wide [optimization budget](https://codeherder.com/docs/optimization-budget/) as a retrospective’s own analysis. If that budget is enabled and exhausted for the day, a replay or its judging simply waits rather than failing — see that page for how to check whether the budget is what’s holding a trial up.

## Starting one

```
ch quality trials create --from-file trial.json
```

`trial.json` names the proposal, the stage and the version of it you’re measuring against, the launch config and model to run the replays on, the judge model, and the source tasks:

```
{
  "proposalId": "019fb0d5-3c81-7e44-a2b9-5f0e7c31da68",
  "type": "story",
  "stage": "review",
  "configId": "019f8b21-6e02-7a14-9c3d-1a2b3c4d5e6f",
  "model": "claude-sonnet-5",
  "judgeModel": "opus",
  "baselineContentHash": "9f2c1d8a7b6e...",
  "credentialConsent": { "secretIds": [], "secretSlots": [] },
  "sources": [
    { "taskId": "019fe33a-05bc-7c18-9e6d-84f2a71b0d95", "baseSha": "a1b2c3d4e5f6..." }
  ]
}
```

`type` is the task type the proposal’s workflow belongs to (`story`, `bug`, and so on) — every source task must share it. `stage` is the exact stage instance the proposal targets, the same name [Retrospectives](https://codeherder.com/docs/retrospectives/) ’ `ch quality optimization proposals` prints as TARGET.

`configId` is the launch config both replays run on. That config has to be **enabled**, and it has to name the `model` you pass on its **Allowed models** list — an exact match, not a tier. This is stricter than the rest of the product: elsewhere an empty Allowed models list means “no restriction” (see [Which model your agents run](https://codeherder.com/docs/agent-models/)), but a trial refuses a config whose list is empty or doesn’t carry that model. If your config runs unrestricted today, add the model to its Allowed models before you create the trial.

`judgeModel` is fixed to `opus` — the same model CodeHerder’s own [stage judging](https://codeherder.com/docs/stage-judging/) uses — and the request is rejected if it’s anything else.

`baselineContentHash` is the content hash of the stage as it stands today. It’s **required whenever the stage is still one of CodeHerder’s built-ins**, which is the case for any stage your workspace has never customised. Get the value from any of the source tasks: run `ch workflow show <taskId>` and read the `stage_content_hash` on the stage instance you’re trialling. For a stage your workspace owns, the field is optional — but if you send it, it still has to match.

Each source task needs the commit its work was actually built on; `baseSha` is that commit’s SHA.

`credentialConsent` names every CodeHerder secret the replays may use. Leave it out when the source launches used no secret. If a source launch used a secret, list its id in `secretIds` (or its device slot key in `secretSlots`). The list must match exactly: a secret you leave out, or one no source used, is refused with a message naming it. Only an admin can create a trial, and CodeHerder records your consent as an event that lists ids only.

Repeating the exact same request returns the trial you already started instead of starting a second one, so retrying a `create` that timed out is safe.

## What a trial can reach

- **CodeHerder:** a trial can read its own frozen launch input, knowledge and skills. It cannot write a comment, a message or a task.
- **CodeHerder secrets:** none, unless you named them in `credentialConsent`.
- **The AI provider:** yes. A replay cannot run without it.
- **The device:** the replays run on a real device. They share its filesystem, its network and any git credentials it holds. CodeHerder does not isolate these.

Run trials on a device that holds nothing you would not give a replay.

## Reading it

```
ch quality trials
```

lists every trial in the workspace, oldest first. Pass a trial’s ID to read it in full:

```
ch quality trials show <trialRef>
```

Both commands print JSON. `show` returns the frozen inputs the trial was created with, one entry per source task holding that task’s baseline and candidate measurements — state, session time, cost, and quality — plus whether the trial is eligible to promote right now, a reason for every guardrail it hasn’t cleared yet, and the limitations of the measurement itself. A missing measurement just means that replay hasn’t finished.

## The guardrails

A trial only becomes eligible to promote once every one of these holds. This bar is deliberately high — a trial sitting short of eligible for a while is normal, not a bug:

- **Every planned pair finishes**, with a quality result and real, priced spend. A session’s cost isn’t counted until at least ten minutes after it exits, so a very recent pair can still be waiting on its bill.
- **Both sides of every pair pass judging**, and the candidate matches or beats its baseline’s quality level. A pair whose *baseline* failed judging can never make the trial eligible, however good the candidate was — so pick source tasks whose original work was sound.
- **No candidate trips a blocking criterion.** This one applies to the candidate side only.
- **Everything the trial froze still matches** — the worker’s launch config version and model, the judge model, and the rubric it grades against.
- **Median session time improves by at least 10%** on the candidate prompt, across the completed pairs.
- **Cost doesn’t rise by more than 5%**, summed across the completed pairs.
- **The improvement clears a one-sided statistical test at p ≤ 0.05** — a tie between a pair’s baseline and candidate counts against the candidate, not for it.
- **No batch in the trial was cancelled.** A cancelled trial can never authorize a promotion, even if some of its pairs already finished cleanly.

## Deciding

Recording an evaluation keeps a permanent snapshot of the trial’s result at that moment:

```
ch quality trials evaluate <trialRef>
```

Promoting applies the candidate prompt to the live stage — only if the trial is eligible:

```
ch quality trials promote <trialRef> --expected-version <n>
```

`--expected-version` names the stage version you reviewed the proposal against, the same number `ch stage versions` or the proposal’s own BASE VERSION shows. If someone has changed the stage since, promote refuses rather than overwrite their edit. If the trial isn’t eligible yet, promote records a hold instead of touching the stage.

Rolling back restores the prompt promote replaced:

```
ch quality trials rollback <trialRef> --expected-version <n>
```

Here `--expected-version` is the version number promote produced — not the version before it. Rollback only works while that promoted version is still the stage’s current one; if it’s been edited or promoted over again since, roll forward by hand instead.

Read the decisions recorded against a trial — the most recent ones, newest first — with:

```
ch quality trials decisions <trialRef>
```

## Read this before you act on it

A trial measures one prompt, on one stage, against a small, deliberately bounded set of tasks. Clearing every guardrail tells you the candidate prompt did better on that set, under those conditions — it isn’t proof the change helps every kind of work your workspace runs. And nothing here promotes itself: an eligible trial is a green light, not an action. A person still runs `promote`.

## Related guides

- [Trial runs](https://codeherder.com/docs/trial-runs/) — replay a task set under a different harness, model, or pinned stage versions
- [Retrospectives](https://codeherder.com/docs/retrospectives/) — where a stage-prompt proposal comes from
- [The optimization budget](https://codeherder.com/docs/optimization-budget/) — the shared cap a trial’s replays and judging draw on
- [Stage judging](https://codeherder.com/docs/stage-judging/) — the same judge model, scoring day-to-day stage attempts
- [Stage-attempt detail](https://codeherder.com/docs/stage-attempts/) — read the individual attempts behind any stage’s numbers
- [The stage library](https://codeherder.com/docs/stage-library/) — stage versions, content hashes, and prompt history
