CodeHerderSearch⌘KRequest access →

AI usage limits

What the AI limits meters show, which ones pause a device, and how CodeHerder recovers automatically.

Every device probes each of its Claude credentials for usage every few minutes and reports the readings back to CodeHerder — a device usually runs on one credential, but it can run on several; see AI credentials on a device for adding more than one. This page covers what the meters mean, which ones pause a device from taking new work, and what you can do about it.

This applies to any Claude credential — a subscription plan (Max, Pro, or Team) or an enterprise/organization API credential. Both report usage; they just report different meters, covered below. Usage from other AI providers is shown for visibility only and never affects placement — see Managing your devices for where those meters live.

What the meters mean

Open a device’s detail page and look at its AI limits section — see Viewing AI limits in Managing your devices for where that is. Each meter you might see there falls into one of two groups:

Meter What it tracks Pauses the device at 100%?
5 hour / 7 day A subscription plan’s rolling usage pools. Yes
7 day (Opus), 7 day (Sonnet), 7 day (OAuth apps) Per-model 7-day pools, shown when the plan reports them. Yes
Monthly spend cap The account’s tier-level spend cap. Yes
Configured spend limit A spend limit the account holder set below the tier cap. Yes
Overage (spend) An enterprise account’s monthly pay-as-you-go allowance. Yes
Account blocked The credential can’t spend at all right now. Yes
Requests/min, Input tokens/min, Output tokens/min (or combined Tokens/min) Per-minute throughput. No

A subscription plan shows the 5 hour and 7 day meters, plus a 7 day (Opus), 7 day (Sonnet), or 7 day (OAuth apps) meter for any per-model pool the plan reports. An enterprise or organization credential shows a different set — typically Overage (spend), and one of the spend-cap meters if the account has crossed it — plus the same per-minute meters every credential reports.

Reading a meter

Each meter’s bar fills as you use that window. It turns amber at 70% used and red at 90%. A marker on the bar shows how far through the window you are; the readout under the bar spells this out as “N% elapsed”. A bar filled further than the marker is ahead of pace for that window, and one filled less is behind it. On a capped credential, the readout also shows “(cap N%)” next to the usage, so you can compare the two at a glance. When a device has more than one Claude credential, its cards are listed in the order the pool would pick them for the next session.

What happens when a capacity meter hits 100%

CodeHerder pauses a device from taking new work only once every one of its Claude credentials that’s still reporting has a meter in the first group above — a subscription pool, a spend cap, or the account-blocked verdict — at 100%, before any of them has had a chance to reset. A device with just one credential pauses as soon as that credential hits the limit. A device with more than one keeps taking work as long as at least one credential still has headroom — see AI credentials on a device for adding a second. A credential that’s gone quiet (see When the probe stops reporting, below) is left out of the check entirely, so a device whose Claude credentials have all stopped reporting isn’t paused because of it. Once a device does pause, which meter tipped it over doesn’t matter: the pause covers every agent on that device, regardless of which coding-agent CLI (harness) each one runs. The pause is per device, not per agent or per harness — an agent using a different harness on the same device is paused too, because the device itself is the thing withholding capacity.

A per-minute meter reaching 100% never pauses the device. Requests, input tokens, output tokens, and combined tokens are throughput buckets — they refill within the minute, so a full Tokens/min meter clears on its own almost immediately. If you see one at 100%, that’s normal under load, not a sign the device is about to stop.

The spend-cap and overage meters reset at the start of the next calendar month (UTC). The subscription pools reset on their own rolling schedule, shown as resets in … next to the meter.

The device’s Online and healthy status do not change when this happens — health checks don’t reflect this condition, so a device can read Online and healthy while it’s being skipped for new work.

Capping a credential is a separate mechanism from this pause. If you’ve capped one of a device’s pooled credentials (see AI credentials on a device), CodeHerder treats it as out of headroom as soon as it hits that cap, not just at the real 100% limit — its card carries a Capped badge alongside its meters. Reaching a cap takes that credential out of the running for new sessions, the same as an exhausted one; the device-wide pause described above still depends only on the real limit.

What happens to sessions already running

A pause never stops a session that’s already running — it only holds back new placements. What happens next depends on whether the device has somewhere else to put the work:

  • If the session’s own credential runs out and the device has another pooled credential with room, CodeHerder moves the session onto it automatically and the session keeps working. See AI credentials on a device for how a device gets a second credential.
  • If no credential on the device has room, the session waits at the usage limit along with the device. CodeHerder holds any new messages you send it and delivers them once the session can pick them up again, so nothing you send gets lost. A session running a different coding-agent CLI on the same device isn’t affected and keeps getting its messages as usual.
  • Once the reset time passes (see It recovers on its own, below), CodeHerder delivers the held messages and the session picks up where it left off.

When the probe stops reporting

If a device’s usage probe hasn’t reported in about 20 minutes, its account row shows a Not reporting badge instead of the meters — the meters are hidden rather than left showing stale numbers. The row’s last line reads “Usage probe stopped reporting. Last probed …” so you can see how long it’s been. A healthy, reporting row ends with a “Probed …” timestamp instead.

This is a reporting gap, not a pause — a Not reporting account doesn’t by itself stop the device taking new work, and it doesn’t count toward the device-wide pause either way (see What happens when a capacity meter hits 100%, above). If the gap is unexpected, check that the device’s Claude credential is still valid; a probe that keeps failing to authenticate never gets a fresh reading to report.

How to tell a device is paused this way

The usage meter behind a 100% pause is still there in AI limits, but you don’t need to go looking for it. CodeHerder also shows the pause itself, in three places:

  • On the devices list, a Claude exhausted badge sits next to the device’s health badge, in the Health column.
  • On the device’s own detail page, a Claude subscription exhausted badge appears above the health checks, with a note on when the pause resets.
  • From the CLI, ch device show <deviceId> prints a staffing row naming the same thing (ch device list does not show it).

It recovers on its own

There’s nothing to do here. CodeHerder tracks the reset time shown on the meter, and the pause lifts as soon as that time passes — it doesn’t wait for the next usage reading first. Tasks that were waiting during the pause start without any manual action. Messages held for a waiting session are delivered at the same point.

Keeping work moving before the reset

If a workflow stage’s Assignee is left on Auto — best fit (the default — see Assigning and claiming work), CodeHerder already tries every eligible agent and device when placing a task. If a second device is Online, healthy, not itself paused, and linked to a workspace with an agent whose capabilities the stage requires, that task is placed there automatically — you don’t need to do anything. When more than one device is eligible, CodeHerder leans toward whichever has more Claude headroom left, so a device that’s getting close to its own limit naturally takes less new work even before it pauses outright.

There’s no separate lever tied to a pinned stage here — see The Assignee dropdown — what it does today in Assigning and claiming work for exactly when a pin does and doesn’t affect dispatch. What gets a task moving before the reset, pinned or not: register or link a second, non-paused device with a capable, eligible agent, add a second Claude credential to the device that’s running low, or simply wait for the meter to roll over.

Your options, in short: wait for the reset, add a second Claude credential to the same device so it keeps working while one credential recovers — see AI credentials on a device — or make sure a second device with headroom is registered, linked to the workspace, and staffed with a capable agent — see Managing your devices for adding and sharing devices.

What won’t help

Reordering, disabling, or adding an agent’s launch configs does not clear this pause — it only changes which config runs next on that device, and selection never looks at live AI usage. A paused device only starts taking work again once at least one of its Claude credentials has headroom — either the same credential resets, a fresh one gets added to that device, or the work moves to a different device — see above for all three.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close