CodeHerderSearch⌘KRequest access →

Alert rules

Watch a metric on a rolling window, open an incident when it breaches, and route the alert to your activity feed or a webhook.

An alert rule watches one metric over a rolling window and opens an incident when the metric stays above its threshold for long enough. It’s opt-in: nothing in your workspace generates an incident until you create a rule for it.

Where to find it

To read your rules and their incidents, use the command line. Any workspace member can run these, and they’re the quickest way in:

ch observability alerts
ch observability alerts incidents --state open

See From the CLI below for the full set.

To create or change a rule, you need the app page. It lives at your workspace’s own address, {{app_base}}/<slug path>/-/observability/alerts — see Using the ch CLI for what a slug path is. Nothing links to that page yet, so don’t go hunting for it in the sidebar or on the Observability page; go there by address and bookmark it.

From that page you can see every rule in the workspace, open one to read its own evaluation timeline and the incidents it raised, and create a new one.

What a rule watches

Today, a rule can watch exactly one metric: Task queue age (p95) — the same p95 figure the task flow report shows for how long queued work has been waiting. There is no catalog of metrics to choose from yet; the Metric field on the create form has one option.

A rule breaches when the observed value rises above its threshold — there’s no “falls below” direction. On each evaluation tick, CodeHerder samples the metric over the rule’s window and compares it to the threshold. If the breach holds for the rule’s sustain period, an incident opens.

Creating a rule

Only an owner or admin can create, edit, or delete a rule. The create form asks for:

  • Name — a label for the rule.
  • Metric — which metric to watch (currently just Task queue age (p95)).
  • Threshold — the value that counts as a breach. For this metric, threshold is in milliseconds.
  • Window (seconds) — how far back each evaluation looks.
  • Sustain (seconds) — how long a breach must hold before an incident opens. The same hold applies to recovery, so a brief dip back under threshold doesn’t resolve an incident early.
  • Cooldown (seconds) — the quiet period after a resolution before the rule can fire again.
  • Minimum volume — below this many open queued tasks sampled, the rule reports insufficient data instead of judging the metric.
  • Scope — see below.
  • Destination — see below.

Before you save, click Preview this rule to replay it against your workspace’s recent history and see what it would have done. Preview writes nothing — no incident, no event, no notification.

Scope

A rule’s scope decides which workspaces it evaluates:

  • Self — only the workspace that owns the rule.
  • Subtree — the owning workspace and every workspace nested beneath it.

Destination, and the webhook precondition

A rule’s destination decides where its fired and resolved events land, but both destinations emit the same underlying event — destination isn’t a second delivery path, it’s a precondition:

  • Activity — the event reaches your workspace’s activity feed and anything already subscribed to it. No setup needed.
  • Webhook — the event also needs a subscribed webhook. Before you can create or enable a rule with destination Webhook, your workspace must already have an enabled webhook subscription carrying the workspace.alert_fired event type. If it doesn’t, the save is refused.

This is the easiest thing to get wrong when setting up a rule: set up your webhook subscription and its workspace.alert_fired event type first, then create the alert rule.

Evaluation verdicts

Each evaluation tick reaches one of five verdicts:

  • OK — measured, and under the threshold.
  • Breaching — measured, and over the threshold.
  • Insufficient data — too few samples to judge at all.
  • Stale data — the evidence behind the metric doesn’t reach back far enough to trust.
  • Read failed — the metric couldn’t be read this tick.

Insufficient data and stale data are not the same as OK, and neither one resolves an open incident. If a rule can’t get a trustworthy reading, an incident it already opened stays open — a data outage should never look like recovery.

Incidents

An incident is open or resolved. While it’s open, CodeHerder tracks the first, peak, and most recent value it observed. An incident resolves once the metric holds back under threshold for the rule’s own sustain period.

Silence

Disabling a rule — the Disable action on its row — silences it. It stops evaluating and can open no new incidents, but it stays configured with all its settings intact. Editing an unrelated field on a silenced rule doesn’t re-enable it; you have to enable it explicitly. If it already had an incident open when you silenced it, that incident stays open until you re-enable the rule and it resolves normally.

Silencing and re-enabling a rule need an owner or admin too. Any workspace member can view the rule list, a rule’s own detail page, and its incidents.

From the CLI

Reading rules and incidents from the command line is open to any workspace member. Rule writes — create, edit, delete — are app-only; there’s no ch command for them.

ch observability alerts
ch observability alerts show <ruleId>
ch observability alerts incidents --state open
ch observability alerts timeline <ruleId>
  • ch observability alerts lists the workspace’s rules. Add --enabled true or --enabled false to narrow to enabled or disabled ones; leave it off to see both.
  • ch observability alerts show <ruleId> shows one rule’s full configuration and its last evaluation.
  • ch observability alerts incidents lists incidents, optionally filtered by rule (--rule <uuid>) or state (--state open/--state resolved).
  • ch observability alerts timeline <ruleId> replays a rule’s evaluation history over a window, alongside the real incidents it raised.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close