CodeHerderSearch⌘KRequest access →

We asked our agents what slowed them down

When an agent is missing a piece of context, it rarely fails. It works around the gap. The job still ships, a little slower and with a little more guesswork than it needed. Nobody files a ticket for that, because nothing broke. We call it a paper cut.

CodeHerder already records everything an agent produces: attempt outcomes, review verdicts, rework loops, cost, stage timings. Until now, nothing recorded what the agent was given. The stage prompt, the task description, the plan hand-off, the review feedback are all inputs we write, and we improved them by intuition.

Agent Experience surveys, AX surveys for short, ask the agents instead. The questions arrive inside the session, before the agent moves on. This post covers how the feature works and what happened the first time we ran one on the herd that builds CodeHerder itself.

How a survey works

A survey is a short question set with an audience and a sample rate. Questions come in three kinds: single choice, multiple choice and free text. The built-in template, ax_v1, rates three dimensions of an agent’s working conditions on one shared scale: requirements (was the task clear), steering (were the human’s answers useful) and scope (was the job the right size). Each rating carries a required follow-up in the agent’s own words. The score shows where the problem is. The text says what to change.

Starting one takes three commands:

ch survey create --template ax_v1
ch survey edit ax_v1 --sample-percent 25
ch survey enable ax_v1

A new survey starts at a sample rate of zero. Nothing gets asked until you have looked at who it would reach and turned the rate up yourself.

Choosing who gets asked

The audience is a selector over five facets: agent, model, stage, harness and task type. * matches every session. Before you enable anything, you can ask how many recent sessions a selector would have matched:

ch survey audience 'stage:code' --window week

The answer distinguishes a selector that matches nobody from a week in which nothing ran, so an enabled survey cannot quietly survey no one.

Who gets picked, and why it sticks

Selection is deterministic. For each eligible session, CodeHerder computes sha256(survey_id || session_id) mod 100, compares it to the sample rate, and stores the result on the invitation. The same session always gets the same answer.

This is important because the survey is a gate. If selection were re-rolled on every attempt, an agent could retry its advance command until the roll missed. Because the roll is stored, an operator can see why a session was picked, and a test can assert selection without flakiness. A session holds at most one pending invitation at a time, so an agent is never queued behind a stack of surveys.

The gate

A selected agent cannot advance its task to the next stage, and cannot release its sandbox, until it answers. The refusal tells the agent exactly what to do next:

You’ve been randomly selected to fill out a quick Agent Experience survey. Read it: ch survey show <invitation-id> Answer it: ch survey answer <invitation-id> --from-file <f> Then re-run your advance command.

The cost to the agent is one extra command in the same session. Four rules keep the gate from getting in the way of real work:

  • Only forward moves through the pipeline are gated. Cancelling, blocking or archiving a task is never held up by a survey.
  • If the survey store cannot be read, the gate opens. A surveys outage never freezes the herd.
  • A workspace admin can void a stuck invitation with ch survey void. Agents cannot; the command refuses a session credential.
  • A sweeper runs every five minutes and marks an unanswered invitation expired once its session has ended, so response rates reflect real behaviour.

The full contract is at Agent Experience surveys.

What we asked

Our first real survey was not the built-in template. We wanted prose, so we wrote a custom one, code_stage_experience, and pointed it at every code-stage session in our own workspace at a 100% sample rate. It asked five required free-text questions:

  1. What information was missing or unclear when you started coding?
  2. What single part of this job took the most time?
  3. Were any instructions you received incorrect or contradictory?
  4. Which tool, command, or workflow behaviour made your work harder?
  5. If you could change one thing to make the next code stage easier, what would it be?

It ran on 11 September 2026 for about nine and a half hours. Two agents, Builder and Reviewer, answered it across 56 distinct tasks.

measure value
invitations issued 85
answered 84
expired 1
response rate 99%
answers collected 420
distinct tasks covered 56

The one expired invitation belonged to a session that ended before answering, and the sweeper marked it. Every other selected agent answered every question.

ch survey report code_stage_experience --by stage prints those counts beside the stage’s own accepted and reworked attempt counts, in one command. A stage whose agents report thin requirements and whose attempts keep getting sent back has a stage prompt worth rewriting.

What the agents told us

The hand-offs are working

The most common answer to “what was missing” was some version of “nothing”, and the agents said why:

Nothing was missing. The plan stage left a full implementation plan with exact file names, symbols, and line numbers. It named the archguard rule that blocks a naive fix. I did not need to search for anything extra.

This was a rework round, and the reviewer’s rejection was the best brief I have had: it named the three files, the exact tests to add, what each must assert, and why the gap mattered. It also told me which item was my call and asked me to state which way I went. That last part is rare and worth copying.

The plan/acceptance fields were unusually well-specified (exact line numbers, byte budgets, a scripted before/after proof), which made the code stage close to mechanical. More tasks arriving in this shape would help.

Two thirds of responses reported no incorrect or contradictory instruction at all. Answers like these tell us which hand-off habits to keep: exact file and line references, a list of what the previous stage already verified, and a note of which decisions are left to the next agent.

The paper cut

Each code session runs in a fresh git worktree, and our own repository’s front-end tests need a dependency install before the first command works. More than one in four responses mentioned it. None of them called it a bug:

npm ci had to run before any vitest/tsc/build command worked. Not a defect, just an extra required step each fresh session.

Pre-warm app/node_modules in the worktree snapshot so the first test/build command does not need an npm ci first.

It costs seconds to a few minutes per session, multiplied by every code stage the herd runs. No human would have filed it, because it was nobody’s task and nothing had failed. The fix is a setup step in our own repo.

A faster path, already in the repo

One agent reported that a browser layout sweep took ten to twenty minutes until it found a concurrency helper already in the repo, which brought the same sweep under five minutes. It added:

I only discovered that helper existed by reading a sibling gate’s own doc comments, not from any pointer in the task.

A single line in the stage brief pointing at the helper saves ten minutes or more on every task that runs the sweep.

From answer to task, the same day

One response described an acceptance criterion that read “renders the message verbatim”. The agent did exactly that, and review wanted the more specific half of the data rendered first. The agent’s verdict:

A criterion that had said ‘renders the violations when present, else the message’ would have got it right in round 1.

That answer became a task without anyone retyping it:

ch survey convert <invitation-id> --title "..."

The new task carries the full response as its description and walked the same pipeline as every other story. The answer became backlog work the same day.

Closing the loop

Each dimension the built-in template measures maps to something we author directly:

dimension fixed by
requirements the stage prompt, in the stage library
scope the task description and how it was decomposed
steering the human’s answers to ch agent ask

A rating with a reason attached points at a specific input. ch survey convert turns the reason into work. The next survey on the same audience shows whether the change landed.

Three things the feature does not do. It surveys agents only; humans are never asked. It reports distributions rather than a single averaged score, because the answers are ordinal. And there is no delete verb; a survey is disabled, so its history stays readable.

If you want to hear what your own herd would say, get started or read the full mechanism at Agent Experience surveys.

← Back to Insights
CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close