# Cost-aware AI engineering: keeping agent spend visible and bounded

Token spend is easy to lose track of with more than one agent running. Here's how to keep it visible per task, agent, and model.

Source: https://codeherder.com/blog/cost-aware-ai-engineering/

June 1, 2026 · The CodeHerder team · 3 min read

- [cost](https://codeherder.com/blog/topics/cost/)
- [observability](https://codeherder.com/blog/topics/observability/)
- [quality](https://codeherder.com/blog/topics/quality/)

A single AI coding agent’s token spend is easy to eyeball. You’re watching the session, you have a rough sense of how long it ran. That intuition breaks down the moment a workspace runs several agents across several tasks at once. Spend that used to feel predictable turns into a number you only learn at the end of the month, from a bill that doesn’t break down by task, agent, or model. It’s not really a cost problem. It’s a visibility problem, and visibility is fixable.

## Cost tracked per task, not just per workspace

CodeHerder tracks cost at the task level, not just the workspace level. Every task shows its own running cost in plain sight: the real number, tied to the actual work that produced it, so you can tell which piece of work was expensive and which wasn’t. You can see cost broken down by agent and by model, over whatever window you care about, from the CLI as easily as from the dashboard:

```
ch costs workspace --window week --by agent,model
```

That’s the same data the dashboard renders, generated live rather than assembled after the fact, and it’s available to your own scripts and CI the moment you want it.

On CodeHerder’s own usage, the median shipped story costs $7.54, and planning takes the largest share of that at 31.4% (code 28.6%, verify 15.6%, review 15.5%, merge 8.9%): the full breakdown is in [100 billion tokens later](https://codeherder.com/blog/100-billion-tokens/). The reasoning model matters more than codebase size, too: swapping it moved the cost of a shipped story 2.4× on the same codebase in the same week, covered in [the most expensive line in your agent config](https://codeherder.com/blog/reasoning-model-is-the-cost-lever/).

## Your API key, your bill

Cost-awareness only matters if the number you’re looking at is the real one. You bring your own provider API key, and CodeHerder never bills you for token spend on top of it. Configure the key directly on your device, and it never leaves that machine; your token usage goes straight to your provider’s dashboard rather than through a CodeHerder markup. What you pay CodeHerder for is the coordination layer: routing, workflow, memory, visibility. What you pay your model provider for is tokens, directly, at their price, with nothing added in between, even though a markup would be the easier business for us to run.

## Billed on concurrent agents, not seats

Plans are metered on concurrent agents, the work actually in flight at any given moment, rather than by headcount. An agent sitting idle in a queue, parked in a done state, or simply registered but not yet claimed doesn’t use a slot or cost you anything extra. The moment a task finishes, that slot frees immediately for the next one. The cost model tracks what you actually care about controlling: how much is running right now.

## What this buys you

None of this requires trusting a forecast. It requires the actual number, attributed correctly, available whenever you ask for it, whether that’s one task, one quarter, or everything in between. Cost-aware engineering isn’t a discipline you bolt on afterwards. It falls out naturally once the number is just there, attributed, the moment you look for it.
