CodeHerderSearch⌘KRequest access →

Understanding costs

How CodeHerder tracks model spend and how to view costs by window, agent, model, or task — from the CLI or the web app's Cost breakdown page.

CodeHerder records model spend and displays totals in USD. You can view costs for your own activity, for a specific agent, task, or session, or for the whole workspace — broken down by a window that defaults to the last 7 days and can be narrowed down to today.

How spend is tracked

When an agent sends a request to an AI model and receives a response, that exchange is one turn. Cost is recorded per turn and attributed to the session in progress. Sessions are grouped under the task they belong to, so you can see what a task cost across all its stages.

Which sessions report cost: Claude Code, Codex, OpenCode, and Pi sessions all report spend this way. Cursor sessions don’t record spend today — see Choosing the coding-agent CLI your agents run for the full picture across harnesses.

Costs are reported in USD.

Reading the token breakdown

Every turn involves several categories of tokens, each billed at a different rate:

  • Fresh input — the part of the prompt that is new and uncached. Billed at the standard input rate.
  • Cached read — prompt tokens that were already stored in the model’s prompt cache. These are billed at a fraction of the fresh-input rate, so a large cached-read count is a sign of efficiency, not expense. When CodeHerder’s agents re-read the same long task brief across many turns, most of it hits the cache.
  • Cache write — tokens written to the cache on a turn that establishes or extends a cache entry. Billed slightly above the fresh-input rate but amortised quickly because every subsequent read of those tokens is cheap.
  • Output — tokens the model generates in its response.

The CLI surfaces all four counts on the summary line:

total:   $0.0421   turns: 12   in: 42000   out: 3200   cache-r: 180000   cache-w: 8400   think: 0   cache-hit: 81.2%

cache-r is cached-read tokens; cache-w is cache-write tokens; cache-hit is the share of all input tokens that were served from the cache. A high cache-hit percentage means the stable prefix of your agents’ prompts is being reused efficiently.

The web app’s Cost breakdown page shows the same categories: an Input tokens tile whose value is the fresh count, with cached read and cache write given as a sub-line beneath it. See Cost in the web app below for the whole page, in the order it renders.

Checking your costs

Run ch costs to see your own spend for the last 7 days:

ch costs

The output shows the time window, total spend, turn count, and per-token-class counts. With no flags, the window is the last 7 days.

Changing the time window

Set the window with --window:

ch costs --window today
ch costs --window week
ch costs --window yesterday

The default is week (the last 7 days). --window today and --window yesterday narrow it; --window month widens it to the last 30 days. For a lookback longer than 30 days, use --since below.

Or set an exact lookback with --since:

ch costs --since 7d
ch costs --since 24h

--since accepts a relative duration (Nd, Nh, or Nm for days, hours, and minutes) or a full RFC 3339 timestamp, up to about 400 days back. --window and --since are mutually exclusive.

Scoping to an agent, task, session, or workspace

ch costs takes an optional leading scope: me (the default — same as running ch costs with nothing), agent <ref>, task [<ref>], session [<ref>], or workspace [<ref>] (also spelled ws). It’s one of four ch commands built on this same scope grammar — see Choosing whose data a list shows for the shared rules.

To view costs for a specific agent, by display name or full id:

ch costs agent <agentRef>

To view one task’s or one session’s spend, the ref is optional — inside an agent session it falls back to the task or session you’re running in:

ch costs task <taskId>
ch costs session <sessionId>

--task <taskId> is a retained alias: on the me scope (the default), it promotes straight to that task’s own rollup, the same one ch costs task <taskId> gives you; on the agent or workspace scopes, it instead narrows that scope down to just that one task.

To view the workspace-wide total (defaults to the workspace set in CH_WORKSPACE_ID):

ch costs workspace
ch costs workspace <workspaceId>

Workspace-wide cost views need the Starter plan; your own, an agent’s, a task’s, and a session’s costs (above) are open on every plan — see Plans and limits.

Breaking down by agent, model, and more

Add --by (the alias --group-by works the same way) to split the totals:

ch costs --window week --by agent
ch costs --window week --by model
ch costs --window month --by agent,model

--by accepts agent, model, stage, type, agentStage, modelStage, billingClass, and taskBinding. The CLI prints a table for four of these: by agent, by model, by billing class (see Billed vs. covered below), and by task binding (see Task-bound vs. task-free spend below).

The other four keys reach the server but don’t get their own CLI table — pass --json to see them in the raw response. Each has a web-app equivalent instead: stage is By workflow stage, type is By workflow, and agentStage/modelStage are the Agent × stage and Model × stage heatmaps — all covered in Cost in the web app below. --by also accepts day, but it does nothing: the sparkline is always day-by-day, regardless of what you pass.

Request more than one printed key and every matching table prints below the totals. The sparkline is always included too.

Use --json to receive the raw JSON envelope instead of the formatted output:

ch costs --window month --by agent,model --json

Task-bound vs. task-free spend

CodeHerder splits your spend into two kinds of work.

Task-bound spend belongs to a task. An agent picks up a task, works through its stages, and every turn that work runs is billed to that task.

Task-free spend belongs to no task. It comes from a session you start on a device, or from a session you start with ch start in your own checkout. Either way, no task owns the turns you run.

Add --by taskBinding (the alias --group-by works the same way) to see the split:

ch costs workspace --by taskBinding
ch costs workspace — window 2026-08-20 → 2026-08-27
  total:   $12.4382   turns: 340   in: 890000   out: 61000   cache-r: 2100000   cache-w: 95000   think: 0   cache-hit: 78.3%

by task binding:
  BINDING     COST      TURNS
  task-bound  $9.8100   260
  task-free   $2.6282   80

The two rows are meant to add up to the total, and for most windows they do, to the cent. When the split can be wrong below covers the one case where they might not line up exactly — either way, the total is the number to trust.

A task scope or a session scope narrows the window to one task or one session, and either one is task-bound or task-free from end to end, so the split has nothing new to tell you there. It’s most useful on the workspace and me scopes, where both kinds of spend usually show up side by side.

When the split can be wrong

For an older window, CodeHerder may not have enough detail left to sort every turn correctly between task-bound and task-free. When that happens, you may see this line under the table:

(task-bound/task-free do not cover the whole window; the total does)

This doesn’t mean any spend goes missing — the total for that window is always complete. It means some of that window’s turns may be filed under the wrong row: work that was really task-bound counted as task-free, or the other way round. Treat the split as an estimate for that window, and trust the total instead.

What one task cost

Two web app views also show cost for a single task, and they cover different periods.

  • Tasks list — Cost column. The Tasks list shows the task’s total recorded spend since it was created. It’s a lifetime figure, not a windowed one, and shows a dash when nothing has been recorded.
  • Task detail page — Cost row and Tokens block. A task’s own page shows a Cost row and a Tokens block covering the last 7 days by default, the same default the Cost breakdown page uses (see Choosing a window below). Both appear only once the task has recorded turns within that window; the Tokens block names the exact dates it covers, so check those if a figure surprises you.

Because one covers the task’s whole life and the other covers a window, the two can disagree: a task that last ran more than a week ago can show a real figure on the Tasks list and nothing at all on its own detail page.

There’s a third figure, covering yet another period: the 30-day total in the next section.

The cost line on ch task show

ch task show <taskId> prints a cost line directly in the task summary, so you don’t need a separate ch costs task <taskId> call just to check a task’s spend at a glance. It covers the last 30 days and shows total spend and turn count. When cache-read cost is non-zero the line breaks it out parenthetically — for example:

cost (30d):  $0.1234 total   ($0.0921 generation · $0.0313 (25%) cache-read)   48 turns

When the task has run across multiple workflow stages, a per-stage breakdown is printed below.

The same command also prints a spend row, showing where the task stands against whichever budget is closest to being hit — see Your task’s effective ceiling in Spend limits.

Model escalation and rework costs

When review, verify, or merge sends a task back to code for rework, CodeHerder automatically runs the rework on a more capable model tier. The first build attempt uses the standard tier; after the first rejection — from any of those three — all subsequent build attempts use the more capable tier.

This means a task that required rework will typically show spend across two model tiers — but the per-stage breakdown in ch task show <taskId> won’t show it: that breakdown is grouped by stage only, so every code attempt is collapsed into one code line regardless of which model tier ran it.

To see the two-tier split for a single task, combine --task with --by model:

ch costs --task <taskId> --by model

This breaks that task’s spend out per model, so the standard-tier and more-capable-tier rows appear separately. To see how model spend is distributed across your whole workspace instead, drop --task and widen the window: ch costs --window week --by model. The web app has the same rollup, scoped to the window you select there rather than to one task — see By agent and by model in Cost in the web app below.

For the full explanation of when escalation applies and which types use it, see How work flows. To see how often each model and stage combination succeeds without rework, see Quality metrics: first-pass rate in Monitoring your agents. For what a completed task costs, and how often it was right the first time, see Cost and rework per completed task.

Cost-aware model routing

Cost-aware model routing is a separate setting that moves in the opposite direction from escalation: escalation moves a stage up to a more capable model after rework is rejected, while routing moves a stage down to a cheaper model when it’s safe to do so. Escalation is about getting a task right after a setback; routing is about spending less when quality won’t suffer. The two run independently, so a task can be escalated and routed at different points in its life.

New workspaces have cost-aware routing turned on. Older workspaces may start off, until an owner or admin turns it on. Either way, a workspace owner or admin can turn it on or off from Settings → General, in the Cost-aware model routing section — check or clear “Route eligible stages to a cheaper model when quality allows.” It’s a per-workspace setting and only appears on a workspace, not on a group.

When it’s on, a workflow stage that is pinned to a specific model can run on a cheaper model instead of its pinned default — but only when CodeHerder expects no drop in quality. That expectation is based on how often that kind of work has recently succeeded on the first try without rework: if the cheaper model’s first-pass success rate on similar stages isn’t high enough, CodeHerder keeps the pinned model. Stages that aren’t pinned to a specific model are never affected either way.

Routing also needs somewhere to look for a cheaper option: it only ever picks from that config’s Allowed models, and only within the pinned model’s own tier. A workspace where no launch config is pinned to a model, or where a pinned config’s Allowed models list has fewer than two models of that tier, has the setting on with nothing for it to do. See Which model your agents run for how pinning and Allowed models work together.

To see what routing has actually saved, or would save if you turned it on, see Model routing in Cost in the web app below.

Routing outcomes from the CLI

ch costs routing-outcomes gives you the same picture from the command line, scoped to a workspace:

ch costs routing-outcomes
ch costs routing-outcomes <workspaceId> --window month

It takes an optional workspace ref (it defaults to CH_WORKSPACE_ID like ch costs workspace does) plus the same --window/--since pair as ch costs, and needs the same Starter plan as a workspace-wide cost view.

The output opens with two headline lines:

live:    actual $8.20   counterfactual $11.40   saved $3.20
shadow:  actual $9.30   counterfactual $6.40   would save $2.90

live totals turns that actually ran on a routed model, against what they would have cost without routing. shadow totals turns where routing was evaluated but not applied, against what they would have cost if it had been — so it shows what you’d save by turning routing on. Below that, a table lists one row per stage-and-model combination, with columns Stage, Model, Mode, Suggested, Actual, Counterfactual, Saving, Turns — Mode marks each row live or shadow; Model is what actually ran, and Suggested is the alternative model CodeHerder compared it against for that row.

Billed vs. covered (AI subscription plans)

If your workspace’s agents run on an AI subscription plan, a portion of the model spend is covered by that plan rather than billed directly. When this applies, the CLI adds a second line below the total:

billed:  $0.0210   covered: $0.0211 (subscription)

The web app shows the same split, and adds a turn-count and percentage view — see the bars, summary tiles, and By billing class in Cost in the web app below.

This section only appears when the workspace has active subscription-covered turns. If you do not see it, all recorded spend was billed directly via API usage.

ch costs --by billingClass gives you the CLI’s own version of that table:

by billing class:
  CLASS                COST      TURNS
  subscription          $0.0211   1
  api (unclassified)    $0.0210   2

Directly billed spend prints as api (unclassified) rather than a bare api — the same spend the web app’s unclassified (counted as API) row shows.

Cost in the web app

The Dashboard has a Cost tile showing your workspace’s 7-day spend. The chart now draws three named lines: Total, Task-bound, and Task-free, matching the split explained in Task-bound vs. task-free spend above. When the split might be wrong for the window (see When the split can be wrong above), a caption under the chart says so: “Task-bound and task-free do not cover the whole window. The total is complete.” Clicking View breakdown from that tile opens the Cost breakdown page — it isn’t a separate sidebar entry, so this link is the way in.

Choosing a window

A single Window filter offers Today, Yesterday, Last 7d, and Last 30d. The page defaults to Last 7d — the same default ch costs uses (see Checking your costs above). If a number here doesn’t match a ch costs run, check that both are looking at the same window first.

The bars and summary tiles

Once the window has any turns, the page opens with:

  • The composition bar — Generation against Context overhead (cache-read) (see Reading the token breakdown above).
  • A Billed / Covered (subscription) bar, shown only when the window has subscription-covered turns.
  • The date range the window covers, as YYYY-MM-DD → YYYY-MM-DD.
  • A row of summary tiles: Total spend, Model turns, Input tokens, and Output tokens always appear; Billed and Covered (subscription) join them when the window has subscription-covered turns; Tasks and Sessions — each a distinct count — join them when the workspace reports those counts.

Reading a rollup table

Every rollup table on the page — by agent, by model, by workflow stage, by billing class, by task binding, by workflow — shares one column set: Name, USD, %, Turns, $/turn, $/task, $/session. The USD column carries a bar behind the number, scaled to the largest row in that table, so you can compare rows at a glance. A share that rounds below 0.05% shows as a dash rather than a falsely precise number.

By agent and by model

Side by side, By agent — rows link through to the agent’s page — and By model break total spend down along those two dimensions.

By workflow stage

By workflow stage lists spend per pipeline stage, ordered by where each stage falls in the pipeline rather than by spend, so you can follow a task’s cost from plan through code and beyond.

By billing class

By billing class splits spend into the two ways a turn can be paid for: a Subscription (covered) row for turns covered by your plan, and an unclassified (counted as API) row for turns billed directly. If the workspace has no subscription-covered turns, only the unclassified (counted as API) row appears. This table doesn’t appear when there’s no billing-class data to show.

By task binding

By task binding splits spend into task-bound work and task-free work, the same Task-bound and Task-free rows explained in Task-bound vs. task-free spend above. You’ll find it between By billing class and Model routing, sharing the same columns as the other rollup tables on this page. Until the workspace has billed usage to group this way, the card shows “No task binding cost rows yet.”

Model routing

When the window has routing data, a Model routing section shows cost-aware routing: what actually ran, against the counterfactual model it would have run on instead, repriced from the current rate card. Two headline figures sit above the table, and which one (or both) you see depends on what’s in the window you’ve selected — not on whether routing is turned on for your workspace right now:

  • Saved by routing — appears once the window contains turns that actually ran through routing.
  • Estimated savings if enabled — appears once the window contains turns that ran with routing evaluated but not applied, recorded as a shadow comparison.

A window that spans both shows both figures together. The table itself lists Stage, Model, Mode, Actual, Would-have, Savings, Turns — the Mode column marks each row Live or Shadow, so you can tell which figure a given row is feeding. For what routing does and how to turn it on, see Cost-aware model routing above.

This section doesn’t appear at all until at least one launch config is pinned to a model, and its savings figures stay at $0 until a pinned config’s Allowed models list gives routing more than one model of that tier to compare — see Which model your agents run for setting both up.

Detailed breakdowns

Click Show detailed breakdowns to reveal three more views (the button relabels to Hide detailed breakdowns once expanded):

  • By workflow — spend per workflow, with an Unattributed row for turns that weren’t tied to a task (see Attributing older turns to their task below to fix that).
  • Agent × stage — a heatmap of spend across agents and pipeline stages.
  • Model × stage — the same heatmap, by model instead of agent.

Each cell shows that pairing’s spend in dollars. Both heatmaps shade each cell relative to the most expensive cell in the grid, leave a cell blank where that pairing had no spend, show turns and cost per turn on hover, and carry a shade legend below the grid running from $0 up to the spend in the most expensive cell.

If the window has no turns at all, the page shows “No cost rows in this window.” instead of the bars and tables above.

Attributing older turns to their task

A turn can end up Unattributed — recorded, but not tied to any task — when it’s older than the detail CodeHerder had on hand at the time. ch costs backfill attribution fixes that: it re-scans your workspace’s history and attaches each unattributed turn to the task its agent was working on when the turn happened.

ch costs backfill attribution

Run it with no argument and it targets CH_WORKSPACE_ID, or name a workspace directly:

ch costs backfill attribution <workspaceId>

It prints how many turns it found and how many it fixed:

Scanned 42 unattributed turn(s); attributed 37 to a task.

A workspace owner or admin can run it, and only a human caller can — an agent’s own session credential is refused. It’s also rate-limited, so running it repeatedly in a short span gets refused rather than re-scanning every time.

Controlling spend

You can control spend three ways: cap it, get warned before it’s capped, or reduce it. An agent daily spend cap and a task total spend cap (default) both stop spend before it grows unchecked — see Spend limits for how to set them, the warnings CodeHerder sends before one binds, and what happens when one is reached. Cost-aware model routing (above) takes a different approach, lowering spend on eligible stages instead of stopping it, with no cap to configure. CodeHerder’s own retrospectives and prompt trials carry one more cap on top of all of this, covering just that work — see The optimization budget.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close