CodeHerderSearch⌘KRequest access →

Setup comparison

Put two setups side by side on the same stages and see which one actually held up.

Stage signals rates a stage overall. ch quality setups breaks that rate down by which setup ran it. Setup comparison answers the narrower question those two leave open: given exactly two setups, which one is doing better on the same stages?

From the app

Open Quality and find the Setups panel. Pick a row, open its Compare with… menu, and choose a sibling setup — one that ran the same stage on the same task class — to open the comparison with both sides already filled in. The comparison page also has its own Window, Setup A, and Setup B pickers, so you can start from there and choose both sides yourself. The Quality page links to it directly as well.

Any workspace member can open it. There’s no extra role or plan requirement.

From the CLI

ch quality compare --setup <a> --setup <b>

Pass --setup exactly twice: the first is setup A, the second is setup B. Each value is a setup selector — at least 8 hex characters of a setup’s own hash, or the full hash.

The easiest way to get both selectors is to skip the hashes: the Compare with… menu and the comparison page’s own pickers fill both sides in for you. To build the command yourself, run ch quality setups --json and read a cell’s setupHash. The plain ch quality setups table doesn’t print that value — its STAGE-HASH column is a different hash, covering the stage’s prompt alone.

An optional workspace name, slug, slug path, or full ID can come first; without one, the command uses the workspace the CLI is pointed at.

The window flags work the same way they do everywhere else in this report family: --window today|yesterday|week|month, or --since with a duration (3d, 12h) or a full RFC 3339 timestamp — the two are mutually exclusive, and without either the report covers the last 90 days. Add --json for the raw data instead of the table.

Reading the report

The header names both setups — harness, model, and stage content — then a cost line:

SETUP A: a1b2c3d4  harness=claude model=sonnet stage-content=1a2b3c4d
SETUP B: e5f6a7b8  harness=codex  model=gpt    stage-content=5e6f7a8b

cost/completed-task:  A=$1.42  B=$1.08

The value right after SETUP A: is that setup’s own hash, so you can feed it straight back to --setup on your next run. The stage-content= value is the other hash mentioned above, and won’t resolve as a selector.

Below that, one row per stage, with the same figures for each side:

STAGE   A-RATE  A-REWORKED  A-DECIDED  A-JUDGE-RATE  A-JUDGE-DECIDED  B-RATE  B-REWORKED  B-DECIDED  B-JUDGE-RATE  B-JUDGE-DECIDED
code    92%     4           50         88%           40               78%     4           18         81%           16
review  95%     3           60         —             0                90%     2           20         —             0
  • RATE — first-pass rate: the share of decided attempts that came out clean, the same figure ch quality setups reports for one setup alone.
  • REWORKED — how many of that setup’s attempts at this stage got sent back.
  • DECIDED — the denominator behind RATE: accepted plus reworked attempts. A stage with a small DECIDED count hasn’t produced enough evidence to trust the rate next to it.
  • JUDGE-RATE / JUDGE-DECIDED — a pass rate and its own denominator, from a different population: the setup’s trial runs at this stage, graded by a judge. These count discarded replays, not the real attempts RATE and DECIDED count, so the two pairs never share a denominator and can disagree without either being wrong. A rate cell with nothing decided prints an em dash (—), never 0%.

Read this before you act on it

  • This adds no new score. Every number here is read straight off the same per-stage, per-class, per-setup cells ch quality setups already shows — comparison just narrows the view to two setups at a time.
  • Trial runs fill the judge columns — nothing else does. A setup with no trial runs at a stage shows nothing in JUDGE-RATE and JUDGE-DECIDED, however much real work it did. Replay that setup as a trial run to get those two columns to say something.
  • A small DECIDED count means the gap between A and B is probably noise, not a real difference. Don’t call a winner off a handful of attempts, and read each pair against its own denominator.
  • The two halves of a row measure different things. RATE and DECIDED report real work, done under whatever conditions it happened to meet. The judge pair reports replays run on purpose. Read them as two pieces of evidence about one setup, not as one figure checked twice.
  • Stage signals — the per-stage rate this report breaks down into two setups
  • Stage-attempt detail — the individual attempts behind every cell
  • Trial runs — the replays behind this report’s judge columns, and how to run one
  • Prompt trials — measure a proposed stage prompt against real finished work
  • Stage judging — independent scoring of real finished attempts, read on the attempt itself rather than here
  • Understanding costs — where cost/completed-task is reported in full
  • Cost and rework per completed task — the task-level cost and first-pass numbers this report doesn’t break down by model or harness

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close