# Monitor background workers

Source: https://codeherder.com/docs/self-host-monitoring/

Poll one endpoint to catch a stuck or always-failing background worker on a self-hosted server, and page when workers are overdue.

A self-hosted CodeHerder server runs background workers. They expire old data, deliver notifications, reconcile sessions, and more. A worker can get stuck or fail on every tick while the server still answers health checks. This page shows how to catch that.

## What to poll

Call this endpoint once every 60 seconds:

```
curl -H "Authorization: Bearer $CH_TOKEN" {{api_base}}/v1/instance/metrics/workers
```

Use the API key of an instance operator. A person who is not an instance operator gets `403`.

You can read the same data with the CLI:

```
ch instance workers
ch instance workers --json
```

The response `data` object has these fields:

| Field | Meaning |
| --- | --- |
| `overdueCount` | How many workers are overdue. This is the number to alert on. |
| `generatedAt` | When the server built the report. |
| `processStartedAt` | When the server process started. |
| `startupDeadlineSeconds` | How long a worker may stay in `starting` before it is overdue. |
| `workers` | One entry for each worker, sorted by `name`. |

Each entry in `workers` has these fields:

| Field | Meaning |
| --- | --- |
| `name` | The worker name. |
| `state` | `starting`, `running`, `stopped`, or `crashed`. |
| `intervalSeconds` | How often the worker ticks. |
| `maxAgeSeconds` | How old its last success may get before it is overdue. |
| `lastSuccessAt` | When a tick last finished without an error or a panic. It is `null` if none has. |
| `ageSeconds` | The time since `lastSuccessAt`. |
| `overdue` | `true` when the worker needs attention. |

## When to page

Page in these two cases:

1. `overdueCount` is above `0` on two polls in a row. The second poll filters out a worker that is late by one tick.
2. The poll itself fails: a non-200 status, or no answer within 10 seconds. A server that cannot answer also cannot run its workers.

Read `workers[].name` in the page text. It names the stuck worker. Then read the server log for that worker name.

## How the age rule works

A tick counts as a success only when it finishes without a panic and without a logged failure. A tick that fails on every run never counts, so the worker ages until it is overdue.

- A worker is overdue when its last success is older than three times its interval. The limit never goes below two minutes. A worker that runs once a day has a limit of three days.
- A `starting` worker has not yet reached its loop. It is overdue when the server has run for more than `startupDeadlineSeconds`.
- A `stopped` worker chose not to run, for example because its setting is off. It is never overdue.
- A `crashed` worker failed while it started. It is always overdue. Restart the server and read the log.

## Limits

- The report describes the one server process that answered. It resets when the server restarts. A restart also resets every age, so a worker that keeps failing shows as overdue again only after its limit passes.
- The report says nothing about the data the workers process. It says that each worker finishes ticks. Use the [logs](https://codeherder.com/docs/self-host-logs/) to see why a tick fails.
