Retrospectives
Turn on evidence capture and analysis, read what a retrospective found, and judge a proposed change before you touch anything.
Every so often, CodeHerder can freeze a batch of recently finished stage attempts and ask an independent agent to read through them: what worked, what it found, and — sometimes — one specific change worth trying. That write-up is a retrospective.
The one thing to know up front: a proposed change is never applied automatically. Recording a retrospective doesn’t touch a stage prompt, a skill, or a wiki page. It’s a lead for a human to check, not an action CodeHerder takes on its own.
Turning it on
Retrospectives are off by default, and there are two separate switches — turning one on doesn’t turn on the other.
Evidence capture freezes the batch of finished attempts:
ch quality optimization enable
Analysis reads a capture and writes it up:
ch quality analyzer enable
Turn on capture without the analyst and CodeHerder retains evidence but never starts a model to read it — nothing shows up beyond “Awaiting analysis.” Most workspaces want both on together.
Check what’s on right now:
ch quality optimization policy
ch quality analyzer policy
analyzer-policy also prints the limits the analyst runs under — how often it may start, how many can run at once, and so on. Read them from that command rather than trusting a number written down here.
Both switches need a workspace owner or admin, and both work per workspace — there’s no version for a group, so turn each on in every workspace that needs it. Pass a workspace name, slug, slug path, or full ID as the trailing argument to reach one other than the workspace the CLI is pointed at.
The analyst is a model session in its own right, so turning analysis on starts paid model work. If you want a ceiling on what that costs, see The optimization budget.
ch quality optimization disable
ch quality analyzer disable
Turning either off stops new work; it doesn’t erase what’s already been recorded.
Reading the results in the app
Open your workspace’s Quality page. The Retrospectives section lists every retained capture, one row per capture date, oldest first — so your most recent retrospective is at the bottom of the list, not the top. Each row shows how many stage attempts it covers and one of three results:
- Awaiting analysis — captured, not yet read.
- No change proposed — read, and nothing stood out.
- Proposal recorded — read, and a specific change is on the table.
Expand a row for the write-up: What worked, Findings, Uncertainty, the evidence selection (and any gaps in it), and a link to each source stage attempt. Once the analyst has read a capture, a link to its own session sits at the top of the expanded row, above the write-up — open it to see the reasoning first-hand. That link is there whether or not the analyst proposed a change.
Reading it from the CLI
The same list, from the command line:
ch quality retrospectives
ID ATTEMPTS ANALYSIS RECORDED PROPOSAL CREATED
019f8c1e-4b27-7a93-8d41-2c6ef0a5b318 35 true false 2026-08-25 00:00:00 +0000 UTC
019fb0d4-91e0-7f52-b7a8-63d1c94ee207 40 true true 2026-09-01 00:00:00 +0000 UTC
019fe33a-05bc-7c18-9e6d-84f2a71b0d95 12 false false 2026-09-08 00:00:00 +0000 UTC
Rows come back oldest first, the same order the web app lists them in.
Fetch one retrospective in full, passing the whole ID from that first column:
ch quality retrospectives show 019fb0d4-91e0-7f52-b7a8-63d1c94ee207
And the proposals that came out of every retrospective, on their own:
ch quality optimization proposals
ID COHORT TARGET BASE VERSION MECHANISM
019fb0d5-3c81-7e44-a2b9-5f0e7c31da68 019fb0d4-91e0-7f52-b7a8-63d1c94ee207 review 4 stage_prompt
Each column prints the proposal’s own stored value: TARGET is the exact key of the thing the change would apply to, BASE VERSION is the version it was written against, and MECHANISM is the kind of change it is.
Every one of these takes an optional workspace reference as its trailing argument, and pages the same way any other ch list does — see Paging through long lists.
How to read a proposal honestly
A proposal names a target and a base version, spells out the exact replacement, states a success measure and a quality guardrail to watch, lists which stage attempts supported it and which ones cut against it, and says plainly what it’s still unsure about.
Read it as a lead worth checking, not a verdict. A few things keep it from being more than that:
- The evidence is one bounded batch, not a representative sample. It’s whatever finished attempts were captured in that window — not every attempt your workspace has ever run.
- Captures overlap. The same task can show up in more than one capture, so seeing it twice isn’t two independent observations of the same thing.
- Supporting evidence usually has to span more than one task — a pattern seen on a single task isn’t a pattern yet. The one exception is a proposal that reproduces a severe, repeatable defect; that stands on a single task, because a reproduction is stronger evidence than a count.
- None of this is a controlled trial. Nothing here randomises which setup ran which task, so a gap between two setups can just as easily reflect what kind of work landed on each one — the same caveat Composition outcomes makes about its own comparisons. If you want something closer to a controlled measurement, a trial run replays the same finished tasks under the setup you’re weighing, on purpose, and throws the work away when it’s done.
Applying a proposal — editing the stage prompt, updating a skill, whatever it recommends — is a decision you make deliberately, the normal way you’d make that change. CodeHerder never makes it for you.
If the proposal changes a stage prompt, you don’t have to apply it on faith: Prompt trials can measure the proposed prompt against the one running today on a set of real finished tasks, before you decide.
Related guides
- Prompt trials — measure a stage-prompt proposal against real tasks before you promote it
- Trial runs — replay a finished task set under a different harness, model, or pinned stage versions
- Stage judging — independent scoring of individual stage attempts, a different quality signal
- Stage signals — per-stage acceptance and rework rates
- Stage-attempt detail — the individual attempts a retrospective draws its evidence from
- Understanding costs — where a retrospective’s own model spend shows up
- The optimization budget — cap what evidence capture and analysis may spend
Last updated