CodeHerderSearch⌘KRequest access →

Monitor background workers

Poll one endpoint to catch a stuck or always-failing background worker on a self-hosted server, and page when workers are overdue.

A self-hosted CodeHerder server runs background workers. They expire old data, deliver notifications, reconcile sessions, and more. A worker can get stuck or fail on every tick while the server still answers health checks. This page shows how to catch that.

What to poll

Call this endpoint once every 60 seconds:

curl -H "Authorization: Bearer $CH_TOKEN" {{api_base}}/v1/instance/metrics/workers

Use the API key of an instance operator. A person who is not an instance operator gets 403.

You can read the same data with the CLI:

ch instance workers
ch instance workers --json

The response data object has these fields:

Field Meaning
overdueCount How many workers are overdue. This is the number to alert on.
generatedAt When the server built the report.
processStartedAt When the server process started.
startupDeadlineSeconds How long a worker may stay in starting before it is overdue.
workers One entry for each worker, sorted by name.

Each entry in workers has these fields:

Field Meaning
name The worker name.
state starting, running, stopped, or crashed.
intervalSeconds How often the worker ticks.
maxAgeSeconds How old its last success may get before it is overdue.
lastSuccessAt When a tick last finished without an error or a panic. It is null if none has.
ageSeconds The time since lastSuccessAt.
overdue true when the worker needs attention.

When to page

Page in these two cases:

  1. overdueCount is above 0 on two polls in a row. The second poll filters out a worker that is late by one tick.
  2. The poll itself fails: a non-200 status, or no answer within 10 seconds. A server that cannot answer also cannot run its workers.

Read workers[].name in the page text. It names the stuck worker. Then read the server log for that worker name.

How the age rule works

A tick counts as a success only when it finishes without a panic and without a logged failure. A tick that fails on every run never counts, so the worker ages until it is overdue.

  • A worker is overdue when its last success is older than three times its interval. The limit never goes below two minutes. A worker that runs once a day has a limit of three days.
  • A starting worker has not yet reached its loop. It is overdue when the server has run for more than startupDeadlineSeconds.
  • A stopped worker chose not to run, for example because its setting is off. It is never overdue.
  • A crashed worker failed while it started. It is always overdue. Restart the server and read the log.

Limits

  • The report describes the one server process that answered. It resets when the server restarts. A restart also resets every age, so a worker that keeps failing shows as overdue again only after its limit passes.
  • The report says nothing about the data the workers process. It says that each worker finishes ticks. Use the logs to see why a tick fails.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close