# The optimization budget

Source: https://codeherder.com/docs/optimization-budget/

Cap what CodeHerder's own quality-improvement work — retrospectives and prompt trials — may spend, read the envelope, and clear an unknown-cost hold.

CodeHerder can spend model time on its own quality-improvement work: a [retrospective](https://codeherder.com/docs/retrospectives/) reads through finished stage attempts, and a [prompt trial](https://codeherder.com/docs/prompt-trials/) replays real tasks twice to compare a proposed prompt against the one running today. The **optimization budget** is an account-wide cap on that spend. It’s an extra cap covering this work alone, on top of the ones in [Spend limits](https://codeherder.com/docs/spend-limits/) — it doesn’t replace them, and they still apply here too.

## What it covers

The optimization budget governs exactly three kinds of work:

- A retrospective’s own analysis session — the one that reads a capture and writes it up.
- Each of a prompt trial’s two replays, baseline and candidate.
- The judge run that scores one of those two replays.

It does not cover anything else. An ordinary task, ordinary stage judging, and an ordinary [trial run](https://codeherder.com/docs/trial-runs/) all run outside this envelope, however much they cost.

## Off until you set it

The optimization budget doesn’t exist until an operator sets one. With no policy set, or a policy you’ve explicitly disabled, optimization work runs with no cap at all — there’s no default limit standing in for it. Set a policy and enable it if you want a ceiling here.

## Reading it

```
ch quality budget show
```

When enabled, this prints the current day’s window, the limit and the per-admission reservation, how much has been reported and reserved, the resulting exposure and what’s left, how many admissions are open or waiting on unknown cost, when someone last cleared an unknown-cost hold and who did it, in-flight exposure and any overrun, the policy’s version number, and a short list of limitations on what the figures do and don’t promise. When the budget is off, it prints only that it’s off, plus the same limitations.

## The units: dollars going in, micro-dollars coming out

Setting a limit uses whole dollars. Reading one back uses millionths of a dollar. Set `--limit-usd 50` and `ch quality budget show` reports the limit as `50000000` in a unit labeled `usd_micros` — that’s still $50, just expressed differently. This is the single most common “did that work?” moment on this feature, so expect it.

## Setting it

```
ch quality budget set --limit-usd 50 --reservation-usd 5 --enable --expected-version 0
```

- `--limit-usd` — the day’s spending ceiling, in whole dollars.
- `--reservation-usd` — how much of that ceiling one admission holds while it’s running, in whole dollars.
- `--enable` or `--disable` — exactly one of these is required.
- `--expected-version` — pass `0` the first time you set a policy for this account. After that, pass the version number `ch quality budget show` last printed. If someone else changed the policy since you last read it, the command refuses rather than overwrite their change — run `show` again and retry with the current version.

Disabling a policy only turns the gate off. It doesn’t cancel anything already running, and it doesn’t erase what happened while the budget was on.

## The window, and what happens when it’s exhausted

The budget resets at the start of each **UTC calendar day**. Within that window, CodeHerder checks — at the moment each unit of optimization work would start — whether admitting it would push the day’s committed spend past the limit.

If admitting a unit would go over, that unit simply isn’t started yet. It stays pending and CodeHerder retries it later; a refusal at this point never cancels work that’s already under way. The policy also records that it’s stopped admissions, with a reason, which is what makes an idle analyst or a trial stuck on “not yet started” explainable rather than mysterious: run `ch quality budget show` and check `Stopped`. The stop clears on its own once there’s room again — a unit finishing under budget frees space immediately, and so does the day rolling over — or as soon as an operator raises the limit.

One nuance worth knowing: the limit bounds the day’s total committed spend, not a count of how many things can run. A unit that ends up costing less than its reservation frees the difference back, so a day can admit more low-cost units than a naive division would suggest. There’s no fixed “you get N runs a day” number to quote here — read the live figures instead of estimating from the limit.

## When the day ends up over the limit

A refusal keeps the day’s spend inside the limit going forward, but it can’t undo spend that has already happened. A unit’s real cost often lands after it started, so a day can finish *above* the limit without a single admission having been refused. When CodeHerder’s periodic check finds the day’s exposure above the limit, it stops the optimization work still running for that account. Ordinary task work is untouched.

What happens to the stopped work depends on which kind it was:

- **A retrospective’s analysis** goes back to waiting, and can run again in a later window. If it has already used up its retry attempts, it’s abandoned instead; if it had already recorded its write-up, that write-up is kept.
- **A prompt trial’s replay, and the judge run that scores one**, are abandoned. They don’t come back on their own when capacity frees up, so a trial that was mid-flight is left without the pair it needed.

So the two outcomes are genuinely different, and it’s worth knowing which you’re looking at: work **refused before it starts** waits and is retried, while a **trial’s work stopped after the day goes over** is finished for good.

## Unknown-cost holds and reconciling them

A unit’s actual cost sometimes doesn’t arrive — the underlying session simply never reports one. Rather than assume it cost nothing, CodeHerder keeps holding that unit’s full reservation indefinitely, because a missing cost report isn’t the same as a zero-dollar report. `ch quality budget show` calls these out as unknown-cost admissions holding a stated amount.

An operator can release those holds directly:

```
ch quality budget reconcile
```

This records your decision to treat this window’s unresolved holds as settled and frees the reservation they were holding. It’s attributed to whoever ran it, and running it again doesn’t release anything a previous reconcile already released.

## Who can do what

Any workspace member can read the budget. Setting a policy or reconciling unknown-cost holds needs a workspace **admin**, signed in as a person — a running agent session can’t do either — and it has to be done from a **leaf workspace**, not a group.

## One policy per account

The optimization budget belongs to the account, not to any one workspace. Whichever workspace you run `ch quality budget set` from, you’re setting the same account-wide policy — there’s no separate optimization budget per workspace to configure.

## What this is not

The optimization budget is an admission gate, not a bill you can pin to an exact figure. It decides whether a new unit of optimization work is allowed to start; it doesn’t promise your provider bill will never exceed the limit. Cost for a unit is only known once it finishes and reports in, so a unit that was admitted and is now running can still end up costing more than the reservation it held — `ch quality budget show` reports that as in-flight exposure and overrun rather than hiding it. That’s also why the limit can be passed before anything is refused, and why the over-limit stop described [above](https://codeherder.com/docs/optimization-budget/#when-the-day-ends-up-over-the-limit) exists to catch it after the fact. For the caps that stop ordinary task sessions the same way, see [Spend limits](https://codeherder.com/docs/spend-limits/).

## Related guides

- [Retrospectives](https://codeherder.com/docs/retrospectives/) — turn on evidence capture and analysis, the work this budget can cap
- [Prompt trials](https://codeherder.com/docs/prompt-trials/) — measure a proposed prompt against real tasks, whose replays and judging draw on this same budget
- [Understanding costs](https://codeherder.com/docs/costs/) — how spend is tracked and reported generally
- [Spend limits](https://codeherder.com/docs/spend-limits/) — the caps on agent, task, and workspace spend; an agent’s daily cap and a workspace budget reach this work too
