Alarm on trace export failures
Watch audit and execution trace export on a self-hosted server. Learn how long an undelivered event survives and set alarms.
A self-hosted CodeHerder server can export audit and execution spans to your collector over OTLP/HTTP. The export runs in the background. The server keeps answering requests when the collector is down. This page shows how to see a failed or late export, and how long you have before events are lost.
The server refuses to start under CH_ENV=production when the export settings are unsafe. It refuses a collector URL that is not https, the CH_OTEL_TRACES_INSECURE flag, and a malformed CH_OTEL_TRACES_HEADERS entry. Read the server log for the reason. The message names the setting and the entry position. It never prints a header value.
What to watch
| Signal | Where | Meaning |
|---|---|---|
audittrace: sdk reported an export error |
Server log, level WARN, one line for each failed export |
The collector refused or did not answer an export. |
exportFailures |
audittrace: export rollup log line, level INFO, every 5 minutes |
The number of failed exports in the last 5 minutes. |
cursorLagSec |
The same rollup line | The age of the newest event the export has read. A large value means the export is behind. |
exportFailures, cursorLagSeconds, lastSweepAt |
ch instance trace-export, or GET /v1/instance/metrics/trace-export |
The same counters for the execution trace export. |
runAuditTraceSweeper, runExecTraceSweeper |
ch instance workers |
Whether each export worker still finishes its ticks. See Monitor background workers. |
The ch instance trace-export command reports only when CH_OTEL_EXECUTION_TRACES is on. With it off, the command shows enabled: false and zero failures, even when the audit export fails. Use the log lines for the audit export.
The rollup line is itself a heartbeat. The export writes it every 5 minutes, even when nothing failed.
How long an undelivered event survives
CodeHerder keeps a failed export in memory for a short time only. Then it drops it.
- Retry. The exporter retries a failed batch with back-off for at most 1 minute. After that, the exporter drops the batch. It does not send the batch again.
- Queue. The exporter holds at most 2048 spans in memory. When the queue is full, it drops new spans.
- Cursor. The export advances its read position when it hands spans to the exporter. It does not wait for the collector to accept them. A dropped span is therefore not read again.
- Source rows. The events that the spans came from stay in the database. Retention removes them later:
CH_EVENTS_MAX_AGE_DAYS(default 180 days) for the audit trace, andCH_EXECUTION_OBSERVATION_RETENTION_DAYS(default 90 days) for the execution trace. - Stopped export. When the export worker stops or fails before it reads, its position stays in the database. It resumes from that position. Events that are still inside the retention window are then exported. Events that retention already removed are lost.
- Open spans. A stage or session span that never gets its closing event is closed after
CH_OTEL_SPAN_MAX_AGE(default 48 hours). It is markedcodeherder.span.unterminated.
So the loss limit has two parts. A collector that is down for more than 1 minute loses the spans in flight at that time. An export that stops reading loses events only when they pass the retention age. Alarm well before that age.
Alarm recipe for CloudWatch
This recipe uses the log shipping in Ship and redact logs. Vector writes each server log line to the CloudWatch group in sink-cloudwatch.yaml. Vector parses the line into an app object. The log text is in message.
Create three metric filters on that log group. Give each one a metric in your own namespace. Then create the alarms.
| Filter name | Filter pattern | Metric value |
|---|---|---|
TraceExportError |
{ $.app.msg = "audittrace: sdk reported an export error" } |
1 |
TraceExportRollup |
{ $.app.msg = "audittrace: export rollup" } |
1 |
TraceExportLag |
{ $.app.msg = "audittrace: export rollup" && $.app.cursorLagSec > 900 } |
1 |
Test each pattern before you rely on it. Run aws logs test-metric-filter with a sample rollup line. Vector keeps values from the server log as text. If TraceExportLag does not match the sample, do not rely on that alarm. Tell the support contact. Watch the cursorLagSec value in the log by hand until it is fixed.
Set these alarms. Send each one to an SNS topic.
| Alarm | Condition | Action |
|---|---|---|
| Export failing | TraceExportError sum is above 0 in 2 consecutive 5-minute periods |
Page your on-call responder. |
| Export behind | TraceExportLag sum is above 0 in 2 consecutive 5-minute periods |
Page your on-call responder. |
| Export stopped | TraceExportRollup sum is 0 for 3 consecutive 5-minute periods, with missing data treated as breaching |
Page your on-call responder. Run ch instance workers. Check runAuditTraceSweeper. |
Only alarm on the last row when the export is on. A server with no CH_OTEL_TRACES_ENDPOINT writes no rollup line.
Escalate to the support contact
- Your on-call responder checks the collector first. Is it up? Does its certificate still verify? Did its ingest key change?
- If the collector is fine and the alarm lasts longer than your own threshold (______), contact the support contact through the channel in your support agreement. Put the alarm name, the time it started, and the last rollup line in the message.
- If the lag is above 86400 seconds (24 hours), treat it as SEV2 and contact the support contact at once. This is well before the 48-hour span limit and the 90-day and 180-day retention limits. See Self-hosted support and incident response.
- Do not paste header values, tokens or the collector URL with its key into the message.
- After the fix, read the next rollup line.
exportFailuresmust be0andcursorLagSecmust fall.
Your support agreement sets the SEV2 answer time. CodeHerder runs no on-call service for self-hosted installs. Your responder owns the first response.
Last updated