Judge calibration
Measure whether your judge's verdicts are trustworthy enough to route work on, and see what changes once they are.
Stage judging grades finished stage attempts, but a grade is only useful if you can trust it. Judge calibration measures that trust directly: it freezes a fixed set of already-decided attempts, scores them again, and checks the judge’s verdicts against what actually happened.
What it answers
Calibration answers one question: is this judge’s verdict, on this stage and rubric, agreeing with reality often enough to act on? It does that by comparing the judge’s calls against attempts whose real outcome you already know — accepted or reworked — rather than trusting the judge’s opinion of itself.
Freeze a corpus
Calibration always measures against a corpus: a fixed, frozen sample of attempts, so every report and every repeat run looks at the same evidence.
A workspace owner or admin freezes one from the CLI, signed in as a person — an agent session can’t do it:
ch quality calibration capture --name calibration-2026-09
This pulls from settled stage attempts over the last 90 days by default; pass --window or
--since to pick a different range, and --description to note why you took the sample. It draws
up to 30 attempts per stage-and-rubric pair from up to 500 candidates, spread across your workspace
so no one stage dominates the sample. Capturing a corpus doesn’t turn on judging or change how
anything routes — it only fixes the evidence you’ll measure against.
Read the results
From the app: open the Quality page and expand the Judge calibration panel. It shows the corpus you’re measuring, a progress row per rubric while scoring runs, and — once a run finishes — the six readiness checks below plus each check’s own pass or fail badge. Any workspace member can view it.
From the CLI:
ch quality calibration report
This prints the same measurement as the app panel: the scoring run and judge model for each stage
and rubric, the agreement figures, and the readiness checks. Pass --corpus <id> to report on an
older frozen corpus instead of the latest one. There’s no --window here — the corpus itself is
the fixed window being measured.
The six readiness checks
A judge only becomes trustworthy enough to route work on once every one of these clears:
- Decided rows — whether enough of the corpus’s attempts have a settled, known outcome to measure the other checks against. Too few, and nothing below can be measured either.
- Cohen’s kappa — whether the judge’s verdicts agree with those known outcomes more than chance would predict on its own.
- False fail rate — how often the judge would have sent back work that actually turned out
fine. This is the direct cost of trusting a
failverdict. - Self-consistency — whether the judge reaches the same verdict on a second, independent scoring run over the same corpus. A judge that flips its own answer isn’t ready to gate on.
- Length bias — whether the spread between accepted and rejected work has been measured across different response lengths at all, so a long-vs-short bias would have a chance to show up.
- Self-preference — whether the judge has been checked against a different harness, so a bias toward its own harness’s output would have a chance to show up.
Each check reports its measured value against its own bar, or reads as unmeasured when there isn’t enough evidence yet to judge it — for example, too few decided rows to publish a kappa figure, or no repeat run yet to measure self-consistency against.
Repeat consistency
Self-consistency needs a second run to compare against. Once a first scoring run finishes, repeat it:
ch quality calibration repeat --scoring-run <id>
This needs a workspace owner or admin, signed in as a person. Retrying the same repeat request is safe — it reuses the run already started rather than starting a second one.
What “armed” changes
A stage-and-rubric pair becomes armed the moment every readiness check above passes, and arming
changes what happens next, not just what the panel shows. Ordinarily a stage judging fail
verdict never sends a task back on its own — only a stage’s own required checks do that. Once a
pair is armed, a fail from that judge sends the task back too, the same way a required check
would. Freezing a corpus and reading its report is how you decide whether to hand a judge that
power before it has it.
Cost note
Capturing a corpus and running or repeating a scoring run spend model budget like any other agent work. See Understanding costs and Spend limits to track and cap it — the app panel also shows the measurement spend for the report you’re looking at, excluding any failed or incomplete run.
Related guides
- Stage judging — turn on judging itself, and what “armed” changes about it
- Understanding costs — where calibration’s spend is reported
- Spend limits — cap what calibration can spend
Last updated