# Reading the Observability report

Source: https://codeherder.com/docs/tool-time/

The full Observability report — summary, rates, trend, attribution coverage, time split, six breakdowns, and the slowest calls, in the app and the CLI.

[Understanding costs](https://codeherder.com/docs/costs/) answers “what did this cost?” Observability answers a different question: “where did the time go, and how much of it was actual tool work?” CodeHerder tracks how long every tool call an agent makes actually takes, and rolls that up into a workspace-wide report — summary numbers, rates, a trend, attribution coverage, a time split by session, six breakdowns, and the individual slowest calls.

## Where to find it

**Observability** sits in the sidebar under **Work**, right after **Sessions** — for a workspace, or for a group. A group’s report covers every workspace nested beneath it, the same way a group rolls up any other report in CodeHerder.

## The summary

The summary reports ten numbers for the window:

- **Tool calls** — how many calls started and finished.
- **Tool time** — the total time spent inside those tool calls.
- **Mean per call** — tool time divided by tool calls.
- **Sessions with tool calls** — how many sessions made at least one qualifying tool call in the window. This is a narrower count than **All sessions** in the time-split tables further down the page, which counts every session whose lifetime overlaps the window, including one that made no tool calls at all. The two numbers measure different populations on purpose — if they don’t match, that isn’t a bug.
- **Tasks** — how many tasks those sessions belong to.
- **Open calls** — calls that started but never reported a result. These are excluded from tool calls and tool time; see [How to read these numbers](https://codeherder.com/docs/tool-time/#how-to-read-these-numbers) below.
- **Tokens** and **Cost** — the recorded token and dollar totals for the window.
- **Slowest call** — the single longest call in the window.
- **Errors** — how many calls in the window ended in an error.

## Filters and window

A **Window** control at the top narrows or widens the whole report at once, the same four choices — Today, Yesterday, Last 7d, Last 30d — as [Understanding costs](https://codeherder.com/docs/costs/#changing-the-time-window), with the same 7-day default.

Six filters sit alongside it, and you can combine them: **Stage**, **Agent**, **Model**, **Category**, **Tool**, and **Label**. Narrow to one stage to see where a pipeline step spends its time, or one tool to see every call it made across sessions — the filters and the six breakdowns below answer the same questions from opposite directions. Changing a filter or the window dims the panel below it and shows a loading spinner while the new figures come in, so you can always tell whether what’s on screen matches your current selection.

## Rates

A compact panel converts the summary into per-unit figures: **Tool calls/session**, **Tool calls/task**, **Tool time/session**, **Tokens/tool call**, **$/session**, and **$/task**. A rate is left blank rather than shown as zero when its denominator is zero — the field means “no measurement,” not “measured, and it’s zero.”

## Trend

Five charts plot the window’s tool activity bucket by bucket: **Tool calls per bucket**, **Duration percentiles per bucket** (p50 and p95, each computed from that bucket’s own calls, never averaged across buckets), **Failure rate per bucket**, **Cost per bucket**, and **Tokens per bucket**. A missing bucket renders as a gap, not a zero — a quiet bucket you never actually measured looks different from one you measured at zero.

Bucketing isn’t always by day. A **Resolution** control picks Auto, Hour, or Day: Auto buckets by hour for a window of 25 hours or less and by day for anything longer, so a short window still shows useful shape. Hour resolution over a long window can produce too many buckets; the report refuses that, so narrow the window or switch to Day.

Two more controls sit alongside Resolution:

- **Compare** — set it to “Previous period” and the **Tool calls**, **Cost**, and **Tokens** charts each add a second, faint line for the equal-length stretch immediately before your window, so you can see whether a metric is trending up or down.
- **Series** — pick Stage, Agent, Model, Category, Tool, or Label and an extra chart appears, breaking tool calls down by that facet’s busiest values, with everything else folded into one “Other” line.

Drag across the **Tool calls per bucket** chart to narrow the window to that range — a **Clear range** button appears once you do, since a brushed range overrides whatever the Window control above is set to. Click a single point on the first chart to open its bucket in [Selected bucket](https://codeherder.com/docs/tool-time/#selected-bucket) below. An **Export CSV** button downloads the trend data behind whatever’s currently on screen.

A point can carry two marks. **Incomplete** means the report window covers only part of the bucket. That happens at either end: the first bucket when your window starts partway through it, and the last bucket when your window, or the current time, ends partway through it. **Low sample** means the bucket has some activity, just not much of it — read a rate (like a failure percentage) from a low-sample bucket with that in mind.

Below the charts, a **Recorded changes** table lists stage-prompt, agent-setup, and release changes that landed during the window — a correlation to review, never a claimed cause of anything the trend shows. A long list of changes can show as truncated; narrow the window to see every one.

## Selected bucket

Pick a point on the trend and this section appears, listing the slowest tool calls inside that one bucket. Each row opens the session that ran the call. A **Clear selection** button closes the section and goes back to the whole window.

## Attribution coverage

This section answers “what does this window’s data actually cover?” rather than “where did the time go” — it’s a companion to the numbers above it, not a seventh breakdown. Three percentages — **Tool name**, **Label**, and **Model** — show what share of this window’s tool calls carry each piece of information. A call can be missing a piece for a few different reasons: it ran against no workflow stage, its session had no routed model recorded, or it predates the columns CodeHerder now tracks. Older records fill in less of this picture than newer ones — that’s expected, not a data-quality problem to chase down.

## Time split

Three tables — **by stage**, **by agent**, **by model** — split each session’s own wall-clock time into tool work and everything else. Each row shows a proportion bar plus:

- **Session time** — the session’s own wall clock, clamped to the window (a session still running when the window ends counts only up to the window’s edge).
- **Tool time** — the union of the session’s tool-call intervals. Two overlapping calls count that overlap once here, unlike the summary’s tool time, which sums every call’s own duration regardless of overlap.
- **Other time** — session time minus tool time: everything the report doesn’t directly measure, model turnaround and overhead included.
- **Parallelism** — summed tool-call time over tool-busy wall clock. A value of 2.00x means two tool calls ran at once on average while tools were busy; blank means the row had no tool-busy time to divide by.
- **All sessions** — every session whose span overlaps the window, including one that made no tool calls at all. This is the larger, “sessions with tool calls” population’s counterpart described under [The summary](https://codeherder.com/docs/tool-time/#the-summary) above.

## The six breakdowns

Below time split, six tables split tool time along different axes:

- **By stage** — which workflow stage the time went to (`plan`, `code`, `review`, and so on), ordered by where each stage falls in the pipeline.
- **By agent** — which agent’s sessions spent the time. A session with no agent behind it (one you started yourself) rolls up under a shared “human session” row.
- **By model** — which model was running.
- **By tool** — which tool ran (a shell command, a file read, an editor call, and so on).
- **By label** — the shape of the command behind the call, with its argument values stripped: `go test | tail`, or `sleep ; glab ci | tail`. A longer chain keeps its pipes and separators, so a label can run to a line of its own.
- **By category** — what kind of work the call was. `shell` usually dominates by a wide margin, since it covers every command an agent runs that isn’t recognised as something narrower. The rest split out `validation` (build, test, lint), reads, edits, waits, and a few narrower bookkeeping kinds. This is the same category vocabulary [Sessions from the command line](https://codeherder.com/docs/session-cli/#what-a-finished-run-left-behind) uses for one session’s own observations.

Each row carries **Tool time**, its **%** share of the total, **Mean / call**, **Calls**, **Sessions**, **Max** (the longest single call in that row), and **Errors**. **By stage**, **by agent**, and **by model** add exact **Tokens** / **Cost** columns; **by category**, **by tool**, and **by label** add **Est. tokens** / **Est. cost** instead, since a category, tool, or label isn’t itself billed — the estimate is CodeHerder’s best attribution of cost to that key. A table longer than its cap folds the rest into one “N other …” row at the bottom. That row carries the true totals for everything it folded in, not an average, and its **View all rows** link opens the complete list.

Not every call arrives with a tool name or a label — an older record, or one CodeHerder can’t summarise, keeps neither. Those calls roll up together into one row, shown as **unrecorded** in the By tool table and **unlabelled** in the By label table; the same names appear in the slowest-calls list.

**By tool**, **by label**, and **by category** also carry a **No tool call** row: tokens and cost recorded on turns where the agent didn’t make a tool call at all. It has no tool time, call count, or session count of its own — only tokens and an estimated cost — because there’s no tool call to attribute those figures to.

## The slowest calls

Below the breakdowns, a **Slowest tool calls** table lists the individual calls that took longest in the window — a fixed top 50, longest first. Only calls with both a start and a result are eligible, so an open call never appears here, however long it’s been running. Each row names the tool, its label, the stage and agent it ran under, how long it took, when it started, and how it finished.

## The same report from the command line

`ch observability` gives you the same report as a family of commands. Each one below takes an optional workspace or group as its first argument, defaulting to your current workspace. Apart from `trace` and `markers`, they all take a flag for each of the app’s six filters: `--stage`, `--agent`, `--model`, `--category`, `--tool-name` (alias `--tool`), and `--label`. They also take `--status`, a filter the app doesn’t have: it narrows to calls with one result status, `success`, `error`, or `unknown`.

```
ch observability summary
```

Prints the same summary figures plus the rates panel, for the last 7 days by default.

```
ch observability breakdown --by stage,agent
```

Prints one table per facet you name. `--by` (or `--group-by`) takes any of `stage`, `agent`, `model`, `category`, `toolName`, `label`, comma-separated — leave it off and every facet prints. Ask for the tool facet as `toolName`; the table it prints is headed `by tool`.

```
ch observability time-split --by stage,agent
```

Prints the time-split tables. `--by` here only accepts `stage`, `agent`, `model` — a time split is a property of a session, not of a tool call.

```
ch observability slowest
```

Prints the same fixed top-50 slowest-calls list the web app shows. There’s no `--limit` — it’s always the top 50.

```
ch observability sparkline
```

Prints the trend figures as a table, one row per bucket — day-bucketed by default, or hour-bucketed for a window of 25 hours or less. Add `--resolution hour` or `--resolution day` to choose explicitly. `--compare previous` prints a second table after the first, headed `previous`, for the equal-length window just before yours. `--by <key>` adds a `series:` list after the tables, with tool calls per bucket for that facet’s top values and an `other` remainder, the same split the app’s Series control draws.

```
ch observability markers
```

Prints the recorded stage-prompt, agent-setup, and release changes in the window, newest first — the same list as the Trend section’s Recorded changes table. It takes only a window (`--window`, or `--since` with an optional `--until`), not the filter flags.

Three more commands drill down from a breakdown row to the sessions behind it. Each step carries your window and filters forward, so you keep the context you started from — in the app, follow **View all rows** under a capped breakdown table and keep going from there.

```
ch observability rows --by toolName
```

Pages through one facet’s *complete* row set — every key, not just the ones that fit above the fold — sorted by tool time by default. `--by` takes exactly one facet here, spelled the same way as above. Also takes `--sort`, `--dir`, `--search`, and the usual `--limit` / `--cursor` paging.

```
ch observability contributors --tool-name Bash
```

Pages through the individual sessions behind whatever you’ve filtered to — which sessions made those calls. It has no `--by`: you narrow it with the filter flags above instead. It also takes an explicit `--since` / `--until` range rather than `--window`, since it’s a paged walk rather than a single snapshot.

```
ch observability trace --session <sessionId>
```

Prints one session’s — or, with `--attempt`, one stage attempt’s — full timeline of tool calls and cost turns. Exactly one of `--session` or `--attempt` is required. Because it names a single session, it takes neither the filter flags above nor `--window`; use `--since` / `--until` to widen the range of evidence it reads.

`ch observability methods` is the one command here that answers locally, without calling the server: it explains how each figure in a report is measured. Task-level flow — how long a task waited versus ran — is a separate report; see [Reading the task flow report](https://codeherder.com/docs/task-flow/).

## One session’s own tool calls

Every report above rolls up a whole workspace. To see one session’s own calls instead, open that session’s page and look at its **Observability** section, or run:

```
ch session tool-calls <sessionId>
```

This prints one row per call — when it started, its tool, label, category, status, and how long it took (or `open` if it never returned a result) — oldest first. It pages like any other `ch` list: 50 rows by default, 200 at most, with `--cursor` to keep going. See [Sessions from the command line](https://codeherder.com/docs/session-cli/#what-a-finished-run-left-behind) for the other things a finished session’s record can tell you.

## How to read these numbers

A few things are worth knowing before you read tool time as a stopwatch:

- **Overlapping calls each count in full, in the summary and the six breakdowns.** Tool time there is the sum of every call’s own duration, not a measure of wall-clock time — if a session runs two calls at once, both durations count, so a window’s total tool time can exceed how long the window actually lasted. Time split’s own **Tool time** column is the one place overlap is *not* double-counted — see [Time split](https://codeherder.com/docs/tool-time/#time-split) above.
- **Open calls are excluded from tool time.** A call that started but never reported a result contributes to the open-calls count, never to tool calls or tool time, since there’s no known duration to add.
- **A result recorded after the session ended isn’t counted either.** Occasionally a session is stopped but the process behind one of its tool calls keeps running, and its result gets recorded hours later. That call is treated the same as an open call — it’s never counted as a real duration, so it can’t inflate tool time or land in the slowest-calls list.
- **The record keeps the shape of a command, never its output.** A call’s label strips out argument values and keeps the programs and operators behind it — `go test | tail`, or `git status | head ; echo ; ch task`. One thing survives that stripping: a script invoked by its path is named by that path, so a label can read `./scripts/with-test-postgres.sh | tail`. Nothing a tool printed is ever stored. Read a label as a sketch of what ran, and check the **By label** table yourself before you export or share it.

## Who can see it

Any workspace member can read the Observability report, in the app or from the CLI — there’s no separate permission to grant.

## Related guides

- [Understanding costs](https://codeherder.com/docs/costs/) — the spend side of the same story: what a task or agent cost, not where the time went
- [Sessions from the command line](https://codeherder.com/docs/session-cli/) — one session’s own tool calls, observations, and execution summary
- [Reading the task flow report](https://codeherder.com/docs/task-flow/) — where a task’s own elapsed time went, and why ready work isn’t running
- [Alert rules](https://codeherder.com/docs/alerts/) — watch this report’s own queue-age figure and open an incident when it breaches
