# AI provider outages

Source: https://codeherder.com/docs/ai-provider-outages/

See what a rate limit, quota or outage at your AI provider looks like, how CodeHerder limits the retries, what it costs, and what to do on each route.

An AI provider can fail while your agents work. It can rate-limit a request, run out of quota, return a server error, or reject a credential. CodeHerder detects these failures, retries the task a limited number of times, and shows you why the task waits.

This page covers Claude Code runs on all three routes: Anthropic direct, Amazon Bedrock, and a gateway. Codex and other harnesses are not covered. They keep their earlier behaviour.

## The three routes

A route is the path a run takes to the AI provider.

| Route | How a run takes it | Set by |
| --- | --- | --- |
| **Anthropic direct** | Neither of the settings below is present. | The default. |
| **Bedrock** | `CLAUDE_CODE_USE_BEDROCK=1` is in the run’s environment. | An environment variable on the device. |
| **Gateway** | `ANTHROPIC_BASE_URL` names a host other than `api.anthropic.com`. | The agent’s launch config. |

The device and the agent together decide the route. CodeHerder tracks each device and agent pair on its own.

## What a failure looks like

Claude Code retries a failed request by itself. When its retries end, it writes one error line to its transcript and waits at the prompt. The run does not exit. The device reads that line and reports it to the server. CodeHerder sorts the failure into one class:

| Class | Means |
| --- | --- |
| **Rate limited** | The provider throttled the request (HTTP 429). |
| **Quota exhausted** | A 429 that names a spent allowance, such as a usage or spend limit. |
| **Overloaded** | The provider is overloaded (HTTP 529). |
| **Server error** | Any other 5xx error or a gateway failure. |
| **Authentication failed** | The provider rejected the credential (HTTP 401 or 403). |

A request that the provider rejects as invalid, such as a prompt that is too long, is not an outage. It follows the normal rules.

You see the failure in three places:

- The activity feed shows a `session.provider_error` event. It names the class, the HTTP status and the route. It also names the action taken: `killed`, `recorded` or `left_to_recovery`.
- The task shows a start refusal. Its message reads, for example: “AI provider error on the Bedrock route: overloaded.” Run `ch task show <taskId>` to read it.
- When a route keeps failing, the task waits with the reason **AI provider unavailable** (`ai_provider_unavailable`).

## How the retries are bounded

CodeHerder stops the failed run and retries the task on a backoff. The count for provider errors is separate from the count for runs that end without progress. An outage cannot use up those attempts.

- Each failed run waits 2 minutes, then 4, 8, 16 and 30 minutes.
- The task parks as `blocked` on the sixth failed run. That is about one hour at the earliest.
- A blocked task carries a blocker note that names the provider error.

After two failed runs in a row on one device and agent pair, CodeHerder stops sending new work to that route. The task waits as **AI provider unavailable**. The wait ends when the backoff ends. The next run is a test. A clean run closes the route. A failed run opens it for longer.

A quota error is a special case. If the device’s Claude subscription is exhausted, CodeHerder does not stop the run. The usage recovery already owns it. It resumes the run when the limit resets. See [AI usage limits](https://codeherder.com/docs/ai-usage-limits/).

The route wait clears on the backoff alone. A usage probe cannot clear it early.

## What cost reporting shows

A failed request has no billable turn. The provider does not bill a rate limit or a server error. Claude Code writes the error line as a synthetic turn with no usage. The cost sweeper skips it.

During an outage you see this:

- No cost rows appear for the failed requests. A test with real Claude Code error lines on all three routes proves it.
- A session keeps only the turns that ran before the failure.
- Task and workspace totals stay flat while the route is down.
- A spend cap or budget does not move during the wait.

## What to do

1. Run `ch task show <taskId>`. Read the start refusal. Note the route and the class.
2. Check the provider.
  - **Anthropic direct:** check the Anthropic status page and your organization’s limits.
  - **Bedrock:** check the AWS service health page, your Bedrock model quotas and throttling, and the device’s `CH_CLAUDE_AWS_*` settings.
  - **Gateway:** check the gateway’s own health, its logs and its upstream limits.
3. For **Authentication failed**, fix the credential. See [AI credentials on a device](https://codeherder.com/docs/device-ai-credentials/).
4. Wait. The task retries on its own when the backoff ends.
5. For a parked task, fix the cause first. Then run `ch task blockers <taskId> --active` and resolve the blocker with `ch task unblock <blockerId>`. The task returns to its stage. See [Blockers](https://codeherder.com/docs/blockers/).

If the same route fails often, look at the provider quota for that route. Raise it, or spread the work across more devices and agents.
