# Judge calibration

Source: https://codeherder.com/docs/judge-calibration/

Measure whether your judge's verdicts are trustworthy enough to route work on, and see what changes once they are.

[Stage judging](https://codeherder.com/docs/stage-judging/) grades finished stage attempts, but a grade is only useful if you can trust it. **Judge calibration** measures that trust directly: it freezes a fixed set of already-decided attempts, scores them again, and checks the judge’s verdicts against what actually happened.

## What it answers

Calibration answers one question: is this judge’s verdict, on this stage and rubric, agreeing with reality often enough to act on? It does that by comparing the judge’s calls against attempts whose real outcome you already know — accepted or reworked — rather than trusting the judge’s opinion of itself.

## Freeze a corpus

Calibration always measures against a **corpus**: a fixed, frozen sample of attempts, so every report and every repeat run looks at the same evidence.

A workspace owner or admin freezes one from the CLI, signed in as a person — an agent session can’t do it:

```
ch quality calibration capture --name calibration-2026-09
```

This pulls from settled stage attempts over the last 90 days by default; pass `--window` or `--since` to pick a different range, and `--description` to note why you took the sample. It draws up to 30 attempts per stage-and-rubric pair from up to 500 candidates, spread across your workspace so no one stage dominates the sample. Capturing a corpus doesn’t turn on judging or change how anything routes — it only fixes the evidence you’ll measure against.

## Read the results

**From the app:** open the **Quality** page and expand the **Judge calibration** panel. It shows the corpus you’re measuring, a progress row per rubric while scoring runs, and — once a run finishes — the six readiness checks below plus each check’s own pass or fail badge. Any workspace member can view it.

**From the CLI:**

```
ch quality calibration report
```

This prints the same measurement as the app panel: the scoring run and judge model for each stage and rubric, the agreement figures, and the readiness checks. Pass `--corpus <id>` to report on an older frozen corpus instead of the latest one. There’s no `--window` here — the corpus itself is the fixed window being measured.

## The six readiness checks

A judge only becomes trustworthy enough to route work on once every one of these clears:

- **Decided rows** — whether enough of the corpus’s attempts have a settled, known outcome to measure the other checks against. Too few, and nothing below can be measured either.
- **Cohen’s kappa** — whether the judge’s verdicts agree with those known outcomes more than chance would predict on its own.
- **False fail rate** — how often the judge would have sent back work that actually turned out fine. This is the direct cost of trusting a `fail` verdict.
- **Self-consistency** — whether the judge reaches the same verdict on a second, independent scoring run over the same corpus. A judge that flips its own answer isn’t ready to gate on.
- **Length bias** — whether the spread between accepted and rejected work has been measured across different response lengths at all, so a long-vs-short bias would have a chance to show up.
- **Self-preference** — whether the judge has been checked against a different harness, so a bias toward its own harness’s output would have a chance to show up.

Each check reports its measured value against its own bar, or reads as unmeasured when there isn’t enough evidence yet to judge it — for example, too few decided rows to publish a kappa figure, or no repeat run yet to measure self-consistency against.

## Repeat consistency

Self-consistency needs a second run to compare against. Once a first scoring run finishes, repeat it:

```
ch quality calibration repeat --scoring-run <id>
```

This needs a workspace owner or admin, signed in as a person. Retrying the same repeat request is safe — it reuses the run already started rather than starting a second one.

## What “armed” changes

A stage-and-rubric pair becomes **armed** the moment every readiness check above passes, and arming changes what happens next, not just what the panel shows. Ordinarily a stage judging `fail` verdict never sends a task back on its own — only a stage’s own required checks do that. Once a pair is armed, a `fail` from that judge sends the task back too, the same way a required check would. Freezing a corpus and reading its report is how you decide whether to hand a judge that power before it has it.

## Cost note

Capturing a corpus and running or repeating a scoring run spend model budget like any other agent work. See [Understanding costs](https://codeherder.com/docs/costs/) and [Spend limits](https://codeherder.com/docs/spend-limits/) to track and cap it — the app panel also shows the measurement spend for the report you’re looking at, excluding any failed or incomplete run.

## Related guides

- [Stage judging](https://codeherder.com/docs/stage-judging/) — turn on judging itself, and what “armed” changes about it
- [Understanding costs](https://codeherder.com/docs/costs/) — where calibration’s spend is reported
- [Spend limits](https://codeherder.com/docs/spend-limits/) — cap what calibration can spend
