Cost and rework per completed task
See what a completed task actually cost, how often it was right first time, and what the window spent regardless of outcome.
Two questions come up before any other quality question: what does finishing a task actually cost, and how often does an agent get it right on the first try? The outcome-cohort report answers both, for tasks that completed in a chosen window.
Where to find it
In the web app, open Quality. The cohort panels sit right below the window filter, above judging activity and the setups list.
From the CLI:
ch quality cohorts
Pass a workspace name, slug, slug path, or full ID as the first argument to scope it elsewhere; without one, the command uses the workspace the CLI is pointed at. See Using the ch CLI for how ch resolves a workspace reference. Add --json to get the raw data instead of the table. Any member of the workspace can run this report — there’s no extra role or plan requirement.
The report always covers the workspace you name plus every workspace nested beneath it — there’s no flag to scope it to just the one workspace.
Two scopes that never reconcile
The report has two halves, and they are not the same tasks:
- The completion cohort is every task that finished in the window — whichever way it finished. Each task counts once, with its full lifetime cost, even if it took several attempts to get there.
- Window spend is every dollar the workspace spent in the window, whether or not a task owns it. Work still in progress counts. So does work that was cancelled. So does spend that belongs to no task at all — a session you start on a device, or a session you start in your own checkout. See Task-bound vs. task-free spend for that split.
Read these as two different questions — “what did finished work cost?” and “what did we spend this window?” — not as two views of one number. They will not add up, and that’s deliberate. The cohort tells you the price of finished work; window spend keeps everything else visible instead of hiding it until it completes. Alongside the dollar total, the report counts how many tasks the window touched. Treat that as a companion figure: the total is not a sum over those tasks.
The five buckets, and their one denominator
Every completed task in the cohort lands in exactly one bucket:
| Bucket | What it means |
|---|---|
| First-pass | Every stage was decided on the first attempt, and the verdict was accepted. |
| Reworked | At least one stage got sent back before the task finished. |
| Pending | A stage is still waiting on a verdict. |
| Censored | A stage was judged but its verdict was withheld. |
| Unjudged | No stage on this task has a settled verdict at all. |
First-pass rate is first-pass tasks divided by decided tasks — first-pass plus reworked only. Pending, censored, and unjudged tasks sit outside that denominator entirely: none of the three counts as a success, and none counts as rework. That keeps the rate honest when a chunk of the cohort simply hasn’t been judged yet.
Some tasks also carry a confounded flag. That means the task’s outcome only came to light after it had passed more than one gate, so the result can’t be pinned on a single stage — usually because the task skipped an earlier gate and a later one caught the problem. Confounded tasks still land in one of the five buckets above; the flag is a caution about attribution, not a sixth outcome.
Reading the daily trend
Below the headline numbers, the report breaks completions out by day: how many finished, the cost per completion, and the rework rate. A day with no completions is a real gap, not a continuation of the day before — don’t read it as “nothing changed.”
If the window’s cohort hit the report’s task cap, the earliest days in that trend under-report, because the cap keeps the newest completions and drops older ones. Treat the early part of a capped window as a floor, not a finished count.
Breaking it down
Add --by with a comma-separated list to split the cohort by workflow, repo, terminal stage, or how the task’s workflow was chosen:
ch quality cohorts --by repo,terminalStage
The valid keys are type (workflow type), repo, terminalStage, and source (whether the workflow was composed, hand-pinned, or the type’s default — see Composition outcomes). Leave --by off and the report shows every breakdown at once, not none of them.
Choosing a window
The default window is the last 30 days. Change it with --window today, --window yesterday, --window week, or --window month, or use --since for a custom range:
ch quality cohorts --since 14d
--since accepts a relative duration (Nd, Nh, or Nm) or a full RFC 3339 timestamp — the same grammar as ch costs --since. --window and --since are mutually exclusive.
Comparing two periods
Add --compare to line the window up against the one immediately before it, of the same length:
ch quality cohorts --compare
This is descriptive, not a statistical test — it won’t tell you whether a change is significant, only what changed. Before reading anything into a cost or rate difference, check the task-mix shift the comparison prints alongside it: if the window simply completed more of a workflow that’s naturally more expensive, that alone explains a lot of the gap.
What this report can’t tell you
A task runs across several sessions and, often, several models — so there’s no single “the model that did this task” to report. This page never breaks the cohort down by model or harness. For that comparison, see Setup comparison, which puts two setups side by side on the stages they actually ran.
Before you act on a number
The report tells you up front when it’s undercounting: a handful of newly finished tasks that haven’t been fully accounted for yet, and — on a busy window — the task cap mentioned above. Read that note before treating any figure as final; CodeHerder catches up on its own, so a small undercount usually clears on the next run.
Related guides
- Monitoring your agents — the per-model, per-stage first-pass rate, a different number from the one on this page
- Setup comparison — the model and harness breakdown this report deliberately leaves out
- Composition outcomes — whether letting the agent choose a task’s workflow is paying off
- Understanding costs — what the cost figures here draw from, and how spend is tracked more broadly
Last updated