# Alarm on trace export failures

Source: https://codeherder.com/docs/self-host-trace-export/

Watch audit and execution trace export on a self-hosted server. Learn how long an undelivered event survives and set alarms.

A self-hosted CodeHerder server can export audit and execution spans to your collector over OTLP/HTTP. The export runs in the background. The server keeps answering requests when the collector is down. This page shows how to see a failed or late export, and how long you have before events are lost.

The server refuses to start under `CH_ENV=production` when the export settings are unsafe. It refuses a collector URL that is not `https`, the `CH_OTEL_TRACES_INSECURE` flag, and a malformed `CH_OTEL_TRACES_HEADERS` entry. Read the server log for the reason. The message names the setting and the entry position. It never prints a header value.

## What to watch

| Signal | Where | Meaning |
| --- | --- | --- |
| `audittrace: sdk reported an export error` | Server log, level `WARN`, one line for each failed export | The collector refused or did not answer an export. |
| `exportFailures` | `audittrace: export rollup` log line, level `INFO`, every 5 minutes | The number of failed exports in the last 5 minutes. |
| `cursorLagSec` | The same rollup line | The age of the newest event the export has read. A large value means the export is behind. |
| `exportFailures`, `cursorLagSeconds`, `lastSweepAt` | `ch instance trace-export`, or `GET /v1/instance/metrics/trace-export` | The same counters for the execution trace export. |
| `runAuditTraceSweeper`, `runExecTraceSweeper` | `ch instance workers` | Whether each export worker still finishes its ticks. See [Monitor background workers](https://codeherder.com/docs/self-host-monitoring/). |

The `ch instance trace-export` command reports only when `CH_OTEL_EXECUTION_TRACES` is on. With it off, the command shows `enabled: false` and zero failures, even when the audit export fails. Use the log lines for the audit export.

The rollup line is itself a heartbeat. The export writes it every 5 minutes, even when nothing failed.

## How long an undelivered event survives

CodeHerder keeps a failed export in memory for a short time only. Then it drops it.

- **Retry.** The exporter retries a failed batch with back-off for at most 1 minute. After that, the exporter drops the batch. It does not send the batch again.
- **Queue.** The exporter holds at most 2048 spans in memory. When the queue is full, it drops new spans.
- **Cursor.** The export advances its read position when it hands spans to the exporter. It does not wait for the collector to accept them. A dropped span is therefore not read again.
- **Source rows.** The events that the spans came from stay in the database. Retention removes them later: `CH_EVENTS_MAX_AGE_DAYS` (default 180 days) for the audit trace, and `CH_EXECUTION_OBSERVATION_RETENTION_DAYS` (default 90 days) for the execution trace.
- **Stopped export.** When the export worker stops or fails before it reads, its position stays in the database. It resumes from that position. Events that are still inside the retention window are then exported. Events that retention already removed are lost.
- **Open spans.** A stage or session span that never gets its closing event is closed after `CH_OTEL_SPAN_MAX_AGE` (default 48 hours). It is marked `codeherder.span.unterminated`.

So the loss limit has two parts. A collector that is down for more than 1 minute loses the spans in flight at that time. An export that stops reading loses events only when they pass the retention age. Alarm well before that age.

## Alarm recipe for CloudWatch

This recipe uses the log shipping in [Ship and redact logs](https://codeherder.com/docs/self-host-logs/). Vector writes each server log line to the CloudWatch group in `sink-cloudwatch.yaml`. Vector parses the line into an `app` object. The log text is in `message`.

Create three metric filters on that log group. Give each one a metric in your own namespace. Then create the alarms.

| Filter name | Filter pattern | Metric value |
| --- | --- | --- |
| `TraceExportError` | `{ $.app.msg = "audittrace: sdk reported an export error" }` | `1` |
| `TraceExportRollup` | `{ $.app.msg = "audittrace: export rollup" }` | `1` |
| `TraceExportLag` | `{ $.app.msg = "audittrace: export rollup" && $.app.cursorLagSec > 900 }` | `1` |

Test each pattern before you rely on it. Run `aws logs test-metric-filter` with a sample rollup line. Vector keeps values from the server log as text. If `TraceExportLag` does not match the sample, do not rely on that alarm. Tell the support contact. Watch the `cursorLagSec` value in the log by hand until it is fixed.

Set these alarms. Send each one to an SNS topic.

| Alarm | Condition | Action |
| --- | --- | --- |
| Export failing | `TraceExportError` sum is above 0 in 2 consecutive 5-minute periods | Page your on-call responder. |
| Export behind | `TraceExportLag` sum is above 0 in 2 consecutive 5-minute periods | Page your on-call responder. |
| Export stopped | `TraceExportRollup` sum is 0 for 3 consecutive 5-minute periods, with missing data treated as breaching | Page your on-call responder. Run `ch instance workers`. Check `runAuditTraceSweeper`. |

Only alarm on the last row when the export is on. A server with no `CH_OTEL_TRACES_ENDPOINT` writes no rollup line.

## Escalate to the support contact

1. Your on-call responder checks the collector first. Is it up? Does its certificate still verify? Did its ingest key change?
2. If the collector is fine and the alarm lasts longer than your own threshold (______), contact the support contact through the channel in your support agreement. Put the alarm name, the time it started, and the last rollup line in the message.
3. If the lag is above 86400 seconds (24 hours), treat it as SEV2 and contact the support contact at once. This is well before the 48-hour span limit and the 90-day and 180-day retention limits. See [Self-hosted support and incident response](https://codeherder.com/docs/self-host-support/).
4. Do not paste header values, tokens or the collector URL with its key into the message.
5. After the fix, read the next rollup line. `exportFailures` must be `0` and `cursorLagSec` must fall.

Your support agreement sets the SEV2 answer time. CodeHerder runs no on-call service for self-hosted installs. Your responder owns the first response.
