Trial runs
Replay a set of finished tasks under a different setup to measure it, safely, before you switch to it.
A trial run lets you find out how a harness, a model, or a pinned set of stage versions would have performed, without touching real work. You name a set of tasks your workspace already finished, pick the setup you want to measure, and CodeHerder replays each one under it. The replays are thrown away when they finish: a trial run never commits, never opens a merge request, and never moves the real task it replays.
Use it to answer questions like: would switching this workflow to a cheaper model change how it does? Does a candidate launch config hold up on the tasks this workspace actually runs? Should a pinned stage version move to the version after it?
Check this first: the stage you replay needs a rubric
A trial run’s work is thrown away, so nothing downstream ever accepts or rejects it. That leaves one way to score a replay: an independent judge grades the diff it produced against the stage’s rubric.
A stage that names no rubric produces no score. The replays still run and still cost money, and you’ll see them finish, but every one of them lands in the results table as unjudged, with no rate. Check before you spend:
| Stage | Rubric |
|---|---|
plan |
plan_v1 |
code |
code_v1 |
code_unverified |
code_v1 |
review |
review_v1 |
The built-in merge, merge_direct and verify stages name no rubric.
If your workspace has customised its workflow, check the stages you’re about to replay. ch stage show <key> lists a library stage’s verifications, and prints the rubric next to the one that names it. For the effective, per-workflow answer, use ch workflow show <taskId> or ch workspace workflow show --type <key> --json — see Built-in default or your own.
If the stage you care about carries no rubric, give it a verification that names one before you spend anything — see The stage library — or narrow the trial to the stages that already have one (--stage, below).
Trial runs vs. Prompt trials
CodeHerder has two features that both use the word “trial”, and they answer different questions:
- A trial run replays a task set under a different harness, model, or pinned stage versions, and reports a judge’s rate for it. There’s no built-in bar it has to clear; you read the number and decide.
- A Prompt trial measures one proposed stage prompt against the one running today, checks it against a fixed set of guardrails, and can promote it once it clears them.
If you’re asking “what happens if I run this on a different model or harness,” you want a trial run. If you’re asking “should I adopt this specific rewritten prompt,” you want a prompt trial.
An ordinary trial run never receives a CodeHerder secret. If its launch config needs one (a credential_refs entry or an @secret: device slot), CodeHerder skips the run with the reason secrets_not_consented. Its session key is trial-scoped and cannot write comments, messages or tasks. See Prompt trials for the one path that names secrets.
Choosing what to replay
A trial run needs a task set: the finished tasks to replay.
ch trial create --type story --last 5 ...
--last N takes the last N finished tasks of the given --type. Or name specific tasks:
ch trial create --type story --task <taskId> --task <taskId> ...
--last and --task are mutually exclusive — pick one.
Choosing what to measure
A trial run needs a target: what to run the replays on. Either a model plus a harness:
ch trial create --type story --last 5 --harness claude --model claude-sonnet-5
or a launch config, which already names both:
ch trial create --type story --last 5 --launch-config <config-id>
--launch-config and --model are mutually exclusive.
A few more things you can set:
--stage <name>narrows which stages the trial replays. Repeat it for more than one. Omit it and the trial replays every agent stage in the workflow — bridge stages liketodoanddoneare never replayed, and naming one is refused.--stage-version <stage>=<n>pins one stage to a specific version from its stage library history, so you can measure “what if this stage were still on version 3.” Repeat it per stage. Omit a stage here and the trial freezes its live content as it stands right now.--base-ref <ref>names the branch each replay’s worktree is cut from. It defaults to the repository’s default branch.--label <text>gives the trial a name of your own choosing, so it’s easier to find later.
From the app
The workspace nav has its own Trials page. Its New trial button opens a form with the same choices as the CLI: pick a type, the stages to replay and any version override, the task set (last N or picked task IDs), the target (model plus harness, or an agent’s launch config), and an optional label. As you fill it in, the form shows how many trial runs it will request — one per task and replayed stage.
Following the trial
ch trial list
lists your workspace’s trials, newest first, with a STATE, how many runs it fanned out into, its type, harness, model, setup and label. Page through a long list with --limit and --cursor.
ch trial show <trial-id>
prints one trial’s own details plus every trial run it fanned out into: one run per (task, stage) pair, each with its own state and, if it never ran, why. The Trials page in the app shows the same thing, and links each trial run to the session that produced it, so you can read what the agent actually did.
A trial itself is one of:
- pending — created, its runs are queued, none have started yet.
- running — at least one run has started.
- finished — every run reached an end state.
- cancelled — someone stopped it.
Each individual run is one of:
- pending — queued, not started.
- running — in progress.
- discarded — the run finished and its work was thrown away, as designed. This is the state you want to see. It means the replay completed normally and is eligible to be scored.
- abandoned — the run gave up its slot. Real work and judging always come first, and a trial run is the first thing reclaimed when the workspace needs the capacity. A reclaimed run goes back into the queue and tries again; it only lands here once it has used up its retries.
- skipped — the run was never attempted, and never will be. Read the reason beside it. Common ones: the trial’s own pinned workflow has no such stage, no agent in your roster can run the harness and model you targeted, or the task it would have replayed is gone.
Reading the result
Progress isn’t the answer. The score lives in the workspace quality report:
ch quality setups
Below the routing table it prints a second, separate table headed trial runs (judge-derived, never acceptance): one row per stage, task type and setup, with a judge rate and the counts behind it.
- JUDGE-RATE is passed over decided. It shows
—until at least one verdict lands. - PASSED / FAILED are the verdicts recorded so far.
- UNJUDGED counts replays that finished but carry no verdict — a stage with no rubric ends up here, and so does one whose judge hasn’t reported yet. It is never folded into the rate.
- DECIDED is passed plus failed, the rate’s denominator.
Find your trial’s rows by its stage, harness and model. A row isn’t one trial: it groups every scored replay in the window that shares the same stage, task type and setup, so two trials aimed at the same setup add up in the same row. Only runs that reached discarded are counted. The report covers the last 90 days by default; --window and --since change that, and --json gives you the raw numbers.
Read the numbers as evidence, not a verdict. A judge is a proxy for quality, so a handful of replays is a hint and not a result — give a setup enough runs to be worth believing before you move production onto it. The two tables never mix: the routing table counts what humans and gates accepted, this one counts what a judge scored.
Cancelling a trial
ch trial cancel <trial-id>
Drains the trial’s still-pending runs and marks the trial cancelled. A run already in progress keeps going to completion.
What it costs
A trial run spends real money on the same AI credentials your ordinary sessions use. It isn’t free just because the work is discarded.
It does take second place, though. Real work and judging are staffed first, and a trial run is the first thing given up when the workspace runs out of capacity, so a trial never slows down what your team is actually shipping.
To see what measurement is costing you against production work:
ch costs measurement
See Understanding costs for how spend is tracked, and Spend limits to cap it.
Who can do what
Creating and cancelling a trial run need a workspace owner or admin, signed in as a person — an agent’s own session credential is refused. Listing trials and reading one need ordinary workspace membership.
Related guides
- Prompt trials — measure one proposed stage prompt against the one running today, and promote it
- Which model your agents run — measure a model here before you move a launch config to it
- Stage judging — the same judge and the same rubrics, applied to your real attempts
- The stage library — where a stage’s rubric lives, and how to change it
- Retrospectives — where a proposed change often starts
- Understanding costs — how trial-run spend is tracked
- Customising workflows — stages, versions, and the stage library a trial run replays against
Last updated