We asked our agents what slowed them down
Our first Agent Experience survey put five questions to every code-stage agent for a day. 84 of 85 answered. Here is how it works and what the herd told us.
Topic
Trust a fleet of agents by watching what they did, not by hoping it went well.
Observability is the record of what each agent actually did: which files it touched, which commands it ran, and why it made the call it made. For a lone developer this is a diff. For a fleet of agents running unattended, it is the only way to answer "what happened here" hours or days later.
This matters most when something goes wrong. A vague summary from the agent is not enough. You need the command it ran, the file it read before it decided, and the moment a task changed state. CodeHerder keeps that trail, so a review or an incident never starts from a blank screen.
This hub covers what we log, and what we chose not to. It also covers the debugging sessions where the activity feed turned a guess into a five-minute fix.
Request access →Our first Agent Experience survey put five questions to every code-stage agent for a day. 84 of 85 answered. Here is how it works and what the herd told us.
Keeping CLAUDE.md small is best practice, but ours hit 231KB across 37 correct commits. Splitting it cut cost per turn 28%, turns per task by a third.
Agent cost doesn't track codebase size. Swapping our judgment stage's reasoning model moved a story's cost 2.4x, and its lead time nearly 3x.
944 user stories shipped in a week, with no human writing the code. Here's what the median one cost, and what each quality gate caught.
Two months of CodeHerder data: 1.2 million agent API calls, 110 billion tokens, 5,742 finished tasks, and a median story costing $7.54, done in 50 minutes.
Token spend is easy to lose track of with more than one agent running. Here's how to keep it visible per task, agent, and model.
Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.