Stage judging
Turn on independent scoring of finished stage attempts, see what gets judged, read the results, and know when a verdict can send work back.
Stage judging asks an independent agent to grade a finished stage attempt against a rubric. It runs separately from the stage’s own gates.
Turning it on
Stage judging is off by default. An owner or admin turns it on for a workspace from Settings → General, in the Stage judging section:
- Enable stage judging — a single checkbox. Any other member sees a read-only “enabled” or “disabled” line instead.
- It’s a per-workspace setting — a group has no such section, so turn it on in each workspace that needs it.
- No plan requirement. Any plan can turn it on.
Turning it on starts judging new attempts from that point on. It isn’t retroactive.
Judging spends model budget like any other agent work — see Understanding costs and Spend limits to track and cap it.
What gets judged
Not every stage is eligible, and not every eligible attempt gets picked.
A stage is only judged if it declares a rubric to grade against. Out of the box, four built-in stages do:
| Stage | Rubric |
|---|---|
plan |
plan_v1 |
code |
code_v1 |
code_unverified |
code_v1 |
review |
review_v1 |
The built-in merge, merge_direct and verify stages name no rubric. If your workspace has customised its workflow’s stages, check which of yours still carry one with The stage library.
Even on an eligible stage, judging is sampled, not universal — CodeHerder judges a fixed share of attempts per stage, decided consistently so the same attempt always gets the same call:
| Stage | Share judged |
|---|---|
plan |
100% |
code |
20% |
review |
25% |
merge |
never |
verify |
100% |
A stage that isn’t in this list — a custom one your workspace added, for instance — is never sampled for judging at all. verify is sampled at 100%, but the built-in verify stage names no rubric, so its runs skip with “The stage has no judging rubric”.
Built-in default or your own
The table above lists built-in DEFAULTS. A workspace can change what a stage judges against, two ways:
- A
stage_rubricsrow overlays a rubric KEY’s body, without changing which key the stage names. - A workspace-owned stage overrides the built-in one, naming a different rubric key in its own verification.
Check the effective rubric for your workspace this way:
ch rubric list— the resolved library: built-ins overlaid by the nearest workspace-or-ancestor row. The OWNER and DEPTH columns say which row wins.ch stage show <key>— the library stage’s verifications, with the rubric printed beside the one that names it. This is the library default, not a per-workflow answer.ch workflow show <taskId>— the effective, resolved workflow for one task. Use this when a workflow instance overrides a verification.ch workspace workflow show --type <key> --json— readeffectiveConfig.instances[].verifications[].rubric. ItsenabledSourceandrequiredSourcefields mark whether each setting is inherited or overridden.
Watch for one trap: a stage INSTANCE named code can resolve against library stage code_unverified. Read the instance’s stageKey, not its name, to know which library stage — and which rubric — actually applies.
When a verdict can send work back
A stage judging verdict is never a gate on its own — tasks advance without waiting for one, and an
ordinary fail doesn’t change the task’s stage by itself.
Two things do route a task back on a fail: a stage’s own required check, which always has,
and an armed judge. A stage-and-rubric pair becomes armed once judge
calibration confirms that judge is agreeing with real outcomes often
enough to trust. Once armed, that judge’s fail sends the task back the same way a required check
would. Freeze a calibration corpus and review its report before you rely on a judge this way.
Reading the results
Find the Judge label in session lists or details. Open Judge verdict for the task, stage, attempt, summary and each criterion’s rating and evidence.
Verdicts and commentary remain readable after the session ends automatically.
A fail verdict completes judging; execution errors are separate.
For workspace totals:
ch quality judging
Pass a workspace name, slug, slug path, or full ID to scope it elsewhere; without one it uses the workspace the CLI is pointed at. Add --json for raw data. It covers the last 90 days by default; see Choosing a window for --window and --since.
The report names the window it covers, tells you whether judging is on, and then tables up every judging run’s current state.
Sample output, with made-up numbers:
ch quality judging — window Jun 10 00:00 → Sep 08 00:00
Stage judging enabled: true
Current states of ordinary judge runs created in this window; calibration samples and trial runs excluded.
STATE REASON COUNT OLDEST CREATED LATEST UPDATED
pending 3 2026-09-08T00:40:12Z 2026-09-08T00:40:12Z
recorded 41 2026-08-12T09:03:41Z 2026-09-07T22:14:08Z
skipped The attempt was not sampled 206 2026-08-12T08:55:19Z 2026-09-08T01:02:55Z
skipped Stage judging was disabled 12 2026-08-01T00:00:00Z 2026-08-11T23:59:12Z
A run lands in one of six states: skipped, pending, running, recorded, error, or abandoned. Only a skipped row carries a reason — every other state leaves that column empty. A skipped row’s reason is one of these:
| Reason | What it means |
|---|---|
| Stage judging was disabled | The workspace had judging turned off when the attempt ran. |
| The stage has no judging rubric | The stage’s spec doesn’t name a rubric to grade against. |
| The configured rubric was not found | The stage names a rubric key that no longer resolves to one. |
| No independent judge candidate was available | Nobody eligible to judge the attempt was free to do it. |
| The change artifact was unavailable | The judge couldn’t read the diff it needed to grade. |
| The frozen artifact was unavailable | The judge couldn’t read the frozen snapshot it needed to grade. |
| The attempt was not sampled | The stage is eligible, but this attempt’s share didn’t come up. |
| The stage has no sampling policy | The stage isn’t one of the ones CodeHerder samples for judging at all. |
A verdict for one attempt also shows up alongside that attempt’s other details:
ch quality attempts
See Stage-attempt detail for the rest of that report. Its JUDGE column reads the verdict plus the worst level the attempt hit (for example fail (L1)) once recorded, the run’s current state (like skipped) when there’s no verdict yet, or a dash when judging never touched that attempt.
The judge answers every rubric criterion with one of that rubric’s own levels, and L is the index of the worst answer, counted from 0 at the worst end — a lower number is a worse result, not a measure of confidence. Run ch rubric show <key> for that rubric’s scale, worst first. Two rubrics can use different scales, so don’t compare an L number across them.
Writing your own rubric
CodeHerder ships four built-in rubrics: plan_v1, code_v1, review_v1, and verify_v1. verify_v1 is available, but no built-in stage names it — attach it to a stage yourself to use it. You can also add your own rubric and point a stage at it.
ch rubric list
ch rubric show <key>
ch rubric versions <key>
ch rubric create --key <key> --stage <stage> --from-file <path|->
list, show, and versions are open to any workspace member. create needs a workspace owner or admin, and has to come from a real person — an agent’s own credentials are refused.
A rubric is immutable: there’s no edit and no delete. Changing what it grades means running create again under the same key, which writes a new version. See Version history and going back for how a rubric’s history compares with the other things CodeHerder keeps a history of.
Related guides
- Judge calibration — measure whether a judge’s verdicts are trustworthy enough to route work on, and what arming it changes
- Stage-attempt detail — the JUDGE column, and the rest of that report
- Stage signals — how well each stage’s own gate is calling it
- The stage library — which stages carry a judging rubric
- Version history and going back — how a rubric’s history compares with a workflow’s, a stage’s, and the rest
- Understanding costs — where judging’s spend shows up
- Spend limits — cap what judging can spend
- Retrospectives — a periodic write-up of recent stage attempts, a different quality signal
Last updated