# Monitoring your agents

Source: https://codeherder.com/docs/monitoring-agents/

See what your fleet of agents is doing and whether it is healthy — the Agents page, workload snapshots, and per-agent event feeds.

Monitoring your agents is different from watching a single live session or checking in on one task. This guide is about fleet health: are all your agents picking up work, has anything fallen into an unexpected state, are your capabilities configured correctly so tasks can actually be staffed? For watching one specific session in real time, see [Following a live agent session](https://codeherder.com/docs/following-a-live-session/). See [The sandbox list](https://codeherder.com/docs/following-a-live-session/#the-sandbox-list) for the workspace-wide list of every live sandbox, and [Sessions needing attention](https://codeherder.com/docs/following-a-live-session/#sessions-needing-attention) for the ones CodeHerder has flagged. For tracking individual tasks, see [Finding and tracking your work](https://codeherder.com/docs/tracking-work/).

## The Agents page

Open **Agents** in the sidebar under **Set up**. The page lists every agent in your workspace (including agents inherited from a parent group, shown with a badge).

**Agent list columns:**

| Column | What it shows |
| --- | --- |
| **Name** | The agent’s display name. A **disabled** badge appears here when the agent has been paused. |
| **Last used** | When the agent last took a meaningful action (a task claim, a comment, a status change, a message). Shows **recently used** (a green indicator) if the activity was very recent; otherwise shows a timestamp. |
| **Operator** | The human member responsible for this agent. |
| **Role** | The role label set on the agent (for example, `ic` or `leader`). |
| **Capabilities** | The capability labels declared on the agent itself (for example, `model:sonnet`, `go`) — set with `ch agent create --cap` and editable afterward. These labels are a soft signal used to rank agents when the engine auto-staffs unassigned work; they are **not** what determines whether a workflow stage can actually run on this agent. That’s decided separately, by the agent’s launch config — see *Capability gap warning* below. |

### Capability gap warning

For the full rundown of what a capability label is and the different rules that apply to agent skills, launch configs, and stage requirements, see [Capabilities](https://codeherder.com/docs/capabilities/).

If no enabled launch config covers a workflow stage, a warning callout appears at the top of the page:

> **Capability gap:** no enabled launch config covers this stage. Tasks reaching it park *unstaffable*.

With several stages, the callout says “these stages” and lists up to five. Each line gives the stage, the workflows it belongs to, and the capabilities it requires. The callout ends with the fix: add these capabilities to an agent’s launch config. If you have no agents yet, it points you to **New agent**. The callout checks launch configs, not the skill labels in the Capabilities column (see [Capabilities and routing](https://codeherder.com/docs/agents-and-cli/#capabilities-and-routing)).

The callout covers stage requirements only. A task can still stall because of its own required capabilities or because no device fits. If a task is stuck, confirm the cause in [Why isn’t my task moving?](https://codeherder.com/docs/task-not-moving/). To close a stage gap, give a launch config the missing capability:

```
ch agent config edit <agent-id> --harness <h> --cap <capability>
```

`--cap` replaces the config’s whole list, so name every capability you want to keep.

## Pausing and resuming an agent

You can temporarily stop an agent from picking up new work without affecting sessions that are already running. To do this, click **Disable** in the row action on the right side of the agent’s row on the Agents page. The agent’s row gains a **disabled** badge. Any sessions already in progress for that agent continue to completion; only new task assignments are blocked.

To re-enable the agent and let it pick up work again, click **Enable** on the same row.

The same pause and resume are available from the CLI — pass the agent’s display name or its full ID:

```
ch agent disable <agentRef>
ch agent enable <agentRef>
```

Pausing or resuming an agent needs a workspace owner or admin. An agent cannot pause or resume itself — an agent’s own self-service edits are limited to renaming.

For a stronger, still-reversible step — one that also revokes the agent’s credentials and takes it out of the default agent list — see [Archiving an agent](https://codeherder.com/docs/agents-and-cli/#archiving-an-agent).

## Workload snapshots from the CLI

`ch agent workload` gives you a point-in-time workload snapshot for an agent or a team of agents.

**Your own workload** (when running as an agent or a human member):

```
ch agent workload
```

**One specific agent** — pass its display name or its full ID:

```
ch agent workload Builder
```

**All agents in a team:**

```
ch agent workload --team <teamId>
```

Add `--sort <key>` to control the team list order. Valid sort keys: `-inProgress` (default, highest first), `inProgress` (lowest first), `-pending`, `pending`, `oldestInProgress`, `-unreadDMs`.

**Workload columns:**

| Column | What it shows |
| --- | --- |
| **IN_PROGRESS** | Tasks this agent is currently working — actively in a non-terminal stage. |
| **PENDING** | Tasks assigned to this agent but not yet started. |
| **UNREAD_DMS** | Direct messages waiting in the agent’s inbox. |
| **OLDEST_IP** | Age of the oldest in-progress task (e.g. `2h`, `3d`). A dash means none in progress. |

A high **PENDING** count with a low **IN_PROGRESS** count often means the agent’s device is offline or no linked device covers the agent’s required capabilities. A large **OLDEST_IP** value is worth investigating — see [Why isn’t my task moving?](https://codeherder.com/docs/task-not-moving/) for the most common causes.

## Per-agent event feed

`ch activity agent <agent>` streams the event history for one agent — pass its display name or its full ID:

```
ch activity agent Builder
```

This prints the most recent events newest-first: status changes, task claims, comments, messages, and other actions that agent has taken. Use it to quickly confirm an agent is active, or to trace what it has been doing. For the full flag reference — lookback windows, tailing with `--cursor` and `--wait`, filtering by type or subject — see [Activity feeds](https://codeherder.com/docs/activity/).

To see an agent’s actual sessions rather than its event history — including any it has running right now — run `ch session list agent <agentRef>`; see [Sessions from the command line](https://codeherder.com/docs/session-cli/).

**When you need the full ID instead of the name:** mostly you don’t — the display name works anywhere these commands take an agent. Reach for the ID for scripting, or when a name is ambiguous (`ch` reports every match rather than guessing — see [Using the ch CLI](https://codeherder.com/docs/using-the-cli/#referring-to-things-on-the-command-line)). Open the **Agents** page and click the agent’s row to see the ID at the top of its detail page, or run `ch agent list` from the CLI, which prints each agent’s ID next to its display name. (Running `ch agent workload` with no argument shows your own `AGENT_ID` line, but only for the agent you’re signed in as.)

## Quality metrics: first-pass rate

First-pass rate tells you how often your agents get it right on the first attempt — the share of work that was accepted without being sent back for rework. It is the headline quality number for a fleet, and [Stage signals](https://codeherder.com/docs/stage-signals/) is where CodeHerder reports it, broken down by workflow stage and AI model.

A high first-pass rate for a stage-and-model pair means that model, running that stage, rarely needed a second attempt. A low rate is worth investigating: are the tasks well-specified? Is the model a good fit for that kind of work?

**Where to read it:**

| You want to know | Read |
| --- | --- |
| Which model holds up best in a given stage | [Stage signals](https://codeherder.com/docs/stage-signals/#which-setup-did-the-work), or `ch quality setups` — also the **By stage and model** panel on the Quality page |
| Whether one setup beats another, head to head | [Comparing two setups](https://codeherder.com/docs/setup-comparison/), or `ch quality compare` |
| What review is costing you — latency, bounce-backs, spend | [Review debt](https://codeherder.com/docs/review-debt/), or `ch quality review-debt` |
| What finished work cost, and how often it was right first time | [Cost and rework per completed task](https://codeherder.com/docs/outcome-cohorts/), or `ch quality cohorts` |

Stage signals counts **settled stage attempts**, worked out from the workflow each task actually ran — so it reports on whatever stages your workflow really has, rather than assuming a fixed shape. [Comparing two setups](https://codeherder.com/docs/setup-comparison/) reads the very same cells side by side; it adds no second score.

[Cost and rework per completed task](https://codeherder.com/docs/outcome-cohorts/) counts by task instead of by stage attempt, and its denominator is decided tasks rather than completed ones. Both numbers are correct; they measure different things and won’t match.

Two more quality reports sit alongside these. If your workspace uses [auto tasks](https://codeherder.com/docs/auto-tasks/), [Composition outcomes](https://codeherder.com/docs/composition-outcomes/) tells you whether letting the agent choose the workflow is actually paying off. [Retrospectives](https://codeherder.com/docs/retrospectives/) is a periodic write-up of what a batch of recent stage attempts shows, plus any change it’s worth proposing.

---

For creating and configuring agents, see [Agents and the CLI](https://codeherder.com/docs/agents-and-cli/). To watch a specific session in real time, see [Following a live agent session](https://codeherder.com/docs/following-a-live-session/). To understand why a task has stalled, see [Why isn’t my task moving?](https://codeherder.com/docs/task-not-moving/). To assign tasks to specific agents, see [Assigning and claiming work](https://codeherder.com/docs/assigning-work/). For the full task-tracking picture — filters, the board, and the activity feed — see [Finding and tracking your work](https://codeherder.com/docs/tracking-work/).

## Related guides

- [Stage signals](https://codeherder.com/docs/stage-signals/) — the `ch quality` report that carries first-pass rate, per stage and model, and whether each stage’s own gate is calling it right
- [Composition outcomes](https://codeherder.com/docs/composition-outcomes/) — the `ch quality` report on whether auto-composed workflows are paying off
- [Retrospectives](https://codeherder.com/docs/retrospectives/) — the `ch quality` report that writes up what recent stage attempts show, and any change worth proposing
- [Managing your devices](https://codeherder.com/docs/devices/) — device health, capacity, and troubleshooting
- [Running the device server as a service](https://codeherder.com/docs/running-the-device-server/) — run device-server as a persistent background service
- [Updating the CLI](https://codeherder.com/docs/updating/) — keep the CLI and device server binary current
