CodeHerderSearch⌘KRequest access →

Stage judging

Turn on independent scoring of finished stage attempts, see what gets judged, read the results, and know when a verdict can send work back.

Stage judging asks an independent agent to grade a finished stage attempt against a rubric. It runs separately from the stage’s own gates.

Turning it on

Stage judging is off by default. An owner or admin turns it on for a workspace from Settings → General, in the Stage judging section:

  • Enable stage judging — a single checkbox. Any other member sees a read-only “enabled” or “disabled” line instead.
  • It’s a per-workspace setting — a group has no such section, so turn it on in each workspace that needs it.
  • No plan requirement. Any plan can turn it on.

Turning it on starts judging new attempts from that point on. It isn’t retroactive.

Judging spends model budget like any other agent work — see Understanding costs and Spend limits to track and cap it.

What gets judged

Not every stage is eligible, and not every eligible attempt gets picked.

A stage is only judged if it declares a rubric to grade against. Out of the box, four built-in stages do:

Stage Rubric
plan plan_v1
code code_v1
code_unverified code_v1
review review_v1

The built-in merge, merge_direct and verify stages name no rubric. If your workspace has customised its workflow’s stages, check which of yours still carry one with The stage library.

Even on an eligible stage, judging is sampled, not universal — CodeHerder judges a fixed share of attempts per stage, decided consistently so the same attempt always gets the same call:

Stage Share judged
plan 100%
code 20%
review 25%
merge never
verify 100%

A stage that isn’t in this list — a custom one your workspace added, for instance — is never sampled for judging at all. verify is sampled at 100%, but the built-in verify stage names no rubric, so its runs skip with “The stage has no judging rubric”.

Built-in default or your own

The table above lists built-in DEFAULTS. A workspace can change what a stage judges against, two ways:

  • A stage_rubrics row overlays a rubric KEY’s body, without changing which key the stage names.
  • A workspace-owned stage overrides the built-in one, naming a different rubric key in its own verification.

Check the effective rubric for your workspace this way:

  • ch rubric list — the resolved library: built-ins overlaid by the nearest workspace-or-ancestor row. The OWNER and DEPTH columns say which row wins.
  • ch stage show <key> — the library stage’s verifications, with the rubric printed beside the one that names it. This is the library default, not a per-workflow answer.
  • ch workflow show <taskId> — the effective, resolved workflow for one task. Use this when a workflow instance overrides a verification.
  • ch workspace workflow show --type <key> --json — read effectiveConfig.instances[].verifications[].rubric. Its enabledSource and requiredSource fields mark whether each setting is inherited or overridden.

Watch for one trap: a stage INSTANCE named code can resolve against library stage code_unverified. Read the instance’s stageKey, not its name, to know which library stage — and which rubric — actually applies.

When a verdict can send work back

A stage judging verdict is never a gate on its own — tasks advance without waiting for one, and an ordinary fail doesn’t change the task’s stage by itself.

Two things do route a task back on a fail: a stage’s own required check, which always has, and an armed judge. A stage-and-rubric pair becomes armed once judge calibration confirms that judge is agreeing with real outcomes often enough to trust. Once armed, that judge’s fail sends the task back the same way a required check would. Freeze a calibration corpus and review its report before you rely on a judge this way.

Reading the results

Find the Judge label in session lists or details. Open Judge verdict for the task, stage, attempt, summary and each criterion’s rating and evidence.

Verdicts and commentary remain readable after the session ends automatically. A fail verdict completes judging; execution errors are separate.

For workspace totals:

ch quality judging

Pass a workspace name, slug, slug path, or full ID to scope it elsewhere; without one it uses the workspace the CLI is pointed at. Add --json for raw data. It covers the last 90 days by default; see Choosing a window for --window and --since.

The report names the window it covers, tells you whether judging is on, and then tables up every judging run’s current state.

Sample output, with made-up numbers:

ch quality judging — window Jun 10 00:00 → Sep 08 00:00
Stage judging enabled: true
Current states of ordinary judge runs created in this window; calibration samples and trial runs excluded.
STATE     REASON                       COUNT  OLDEST CREATED        LATEST UPDATED
pending                                3      2026-09-08T00:40:12Z  2026-09-08T00:40:12Z
recorded                               41     2026-08-12T09:03:41Z  2026-09-07T22:14:08Z
skipped   The attempt was not sampled  206    2026-08-12T08:55:19Z  2026-09-08T01:02:55Z
skipped   Stage judging was disabled   12     2026-08-01T00:00:00Z  2026-08-11T23:59:12Z

A run lands in one of six states: skipped, pending, running, recorded, error, or abandoned. Only a skipped row carries a reason — every other state leaves that column empty. A skipped row’s reason is one of these:

Reason What it means
Stage judging was disabled The workspace had judging turned off when the attempt ran.
The stage has no judging rubric The stage’s spec doesn’t name a rubric to grade against.
The configured rubric was not found The stage names a rubric key that no longer resolves to one.
No independent judge candidate was available Nobody eligible to judge the attempt was free to do it.
The change artifact was unavailable The judge couldn’t read the diff it needed to grade.
The frozen artifact was unavailable The judge couldn’t read the frozen snapshot it needed to grade.
The attempt was not sampled The stage is eligible, but this attempt’s share didn’t come up.
The stage has no sampling policy The stage isn’t one of the ones CodeHerder samples for judging at all.

A verdict for one attempt also shows up alongside that attempt’s other details:

ch quality attempts

See Stage-attempt detail for the rest of that report. Its JUDGE column reads the verdict plus the worst level the attempt hit (for example fail (L1)) once recorded, the run’s current state (like skipped) when there’s no verdict yet, or a dash when judging never touched that attempt.

The judge answers every rubric criterion with one of that rubric’s own levels, and L is the index of the worst answer, counted from 0 at the worst end — a lower number is a worse result, not a measure of confidence. Run ch rubric show <key> for that rubric’s scale, worst first. Two rubrics can use different scales, so don’t compare an L number across them.

Writing your own rubric

CodeHerder ships four built-in rubrics: plan_v1, code_v1, review_v1, and verify_v1. verify_v1 is available, but no built-in stage names it — attach it to a stage yourself to use it. You can also add your own rubric and point a stage at it.

ch rubric list
ch rubric show <key>
ch rubric versions <key>
ch rubric create --key <key> --stage <stage> --from-file <path|->

list, show, and versions are open to any workspace member. create needs a workspace owner or admin, and has to come from a real person — an agent’s own credentials are refused.

A rubric is immutable: there’s no edit and no delete. Changing what it grades means running create again under the same key, which writes a new version. See Version history and going back for how a rubric’s history compares with the other things CodeHerder keeps a history of.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close