Monitor background workers
Poll one endpoint to catch a stuck or always-failing background worker on a self-hosted server, and page when workers are overdue.
A self-hosted CodeHerder server runs background workers. They expire old data, deliver notifications, reconcile sessions, and more. A worker can get stuck or fail on every tick while the server still answers health checks. This page shows how to catch that.
What to poll
Call this endpoint once every 60 seconds:
curl -H "Authorization: Bearer $CH_TOKEN" {{api_base}}/v1/instance/metrics/workers
Use the API key of an instance operator. A person who is not an instance operator gets 403.
You can read the same data with the CLI:
ch instance workers
ch instance workers --json
The response data object has these fields:
| Field | Meaning |
|---|---|
overdueCount |
How many workers are overdue. This is the number to alert on. |
generatedAt |
When the server built the report. |
processStartedAt |
When the server process started. |
startupDeadlineSeconds |
How long a worker may stay in starting before it is overdue. |
workers |
One entry for each worker, sorted by name. |
Each entry in workers has these fields:
| Field | Meaning |
|---|---|
name |
The worker name. |
state |
starting, running, stopped, or crashed. |
intervalSeconds |
How often the worker ticks. |
maxAgeSeconds |
How old its last success may get before it is overdue. |
lastSuccessAt |
When a tick last finished without an error or a panic. It is null if none has. |
ageSeconds |
The time since lastSuccessAt. |
overdue |
true when the worker needs attention. |
When to page
Page in these two cases:
overdueCountis above0on two polls in a row. The second poll filters out a worker that is late by one tick.- The poll itself fails: a non-200 status, or no answer within 10 seconds. A server that cannot answer also cannot run its workers.
Read workers[].name in the page text. It names the stuck worker. Then read the server log for that worker name.
How the age rule works
A tick counts as a success only when it finishes without a panic and without a logged failure. A tick that fails on every run never counts, so the worker ages until it is overdue.
- A worker is overdue when its last success is older than three times its interval. The limit never goes below two minutes. A worker that runs once a day has a limit of three days.
- A
startingworker has not yet reached its loop. It is overdue when the server has run for more thanstartupDeadlineSeconds. - A
stoppedworker chose not to run, for example because its setting is off. It is never overdue. - A
crashedworker failed while it started. It is always overdue. Restart the server and read the log.
Limits
- The report describes the one server process that answered. It resets when the server restarts. A restart also resets every age, so a worker that keeps failing shows as overdue again only after its limit passes.
- The report says nothing about the data the workers process. It says that each worker finishes ticks. Use the logs to see why a tick fails.
Last updated