Self-hosted service objectives
Measure each customer journey and feed on a self-hosted server, see the operating ranges CodeHerder states, and set your own targets from them.
This page helps you set service objectives for a self-hosted CodeHerder server. It states the timings the product produces. It shows how to measure each journey. It names the dependency failures that affect each one.
CodeHerder commits only to the supported ranges on this page. You set the target. Every target cell says “You set the target”.
A range comes from a setting or a fixed value in the server. It is not a promise about your network, database, or identity provider. Where no fixed value exists, this page says so. It does not invent a number.
Journeys
Each journey has an indicator to record. Record it from your proxy log, from a scheduled check, or from the worker board. Do not use the health endpoint alone. GET /v1/health proves the process and the database answer. It does not prove a journey works.
API request
- How to measure: Read
statusanddurationfor each request in your proxy access log. See Self-hosted logs. The server also sends aServer-Timingheader on each response. SetCH_SERVER_TIMING=falseto turn it off. - Indicator to record: The share of requests that return a status below 500. The 95th percentile of
duration, grouped by route. - Supported range: CodeHerder states no latency range yet. Use the load test results for the launch workload profile when they ship. Until then, measure your own baseline.
- Failure classes: Database (slow queries, pool wait,
GET /v1/instance/metrics/db). Identity provider (sign-in requests only). - Target: You set the target.
SPA load
- How to measure: Time a page request through your proxy. Run a synthetic browser check against the sign-in page from outside your network.
- Indicator to record: The time until the page returns
200. The share of checks that pass. - Supported range: CodeHerder states no latency range yet. Time it from outside your network, through your proxy.
- Failure classes: Proxy or app host. The page itself needs no AI provider, git host, or device.
- Target: You set the target.
Sign-in
- How to measure: Run
ch instance checkon a schedule. See Self-hosted synthetic checks. Thesign_in,mailandsecret_decryptionitems each report pass or fail. - Indicator to record: The result of each item, and the time the command takes.
- Supported range: CodeHerder states no latency range yet. Sign-in time depends mostly on your identity provider.
- Failure classes: Identity provider (
sign_infails, see Self-hosted Cognito sign-in). Mail (mailfails, so one-time codes do not arrive). Database. - Target: You set the target.
Session attach
Attach is the live terminal view of a running session. It uses a WebSocket.
- How to measure: Open a session from a scheduled check and record whether the socket connects. Count WebSocket closes in your proxy log.
- Indicator to record: The share of attach attempts that connect. The number of unplanned closes per hour.
- Supported range: The server pings each socket every 25 seconds. It gives a write 10 seconds to finish. It checks the caller’s access again every 50 seconds. A revoked credential stops working within 60 seconds. The browser reconnects after 1 second, then doubles the wait up to 30 seconds. It resets the wait after a connection holds for 5 seconds.
- Failure classes: Device (the session runs on a device, so a lost device ends the attach). Proxy (an idle timeout below 25 seconds cuts the socket). Database (the access check fails).
- Target: You set the target.
Task queue to session start
- How to measure: Compare the time a task became ready with the time its session started. Both stamps are on the task and session records.
ch task showprints the wait reason for a task that is not running. See Why isn’t my task moving?. - Indicator to record: The wait from ready to session start, per workflow. The count of tasks with a wait reason for longer than your target.
- Supported range: Task creation and stage change staff work at once when a device has room. The backstop reconcile pass runs every 30 seconds for staffing and capacity. A slower pass runs every 5 minutes. So the backstop retries a missed task about every 30 seconds. A device runs at most 4 sessions by default (
CH_MAX_DEVICE_SESSIONS). A failed spawn retries after 30 seconds, then backs off to a cap of 15 minutes. - Failure classes: Device (none online, or none with the needed capability). AI provider (the wait reason names it, or spawns fail and retry). Git host (clone fails, so the spawn retries). Database.
- Target: You set the target.
Device-loss recovery
- How to measure: Stop a test device and time how long its session takes to read
lost. PollGET /v1/instance/metrics/workersfor the sweeper that marks sessions lost. See Monitor background workers. - Indicator to record: The time from the last device frame to the
loststate. The time until the task has a new session on another device. - Supported range: A device sends a heartbeat every 30 seconds. The server drops a tunnel silent for 75 seconds. The sweeper marks a session
lostwhen its device tunnel stays offline for 90 seconds. The sweeper runs every 30 seconds. So a session readslost90 to 120 seconds after the last frame. If the tunnel state is unclear, a session with no device contact for 900 seconds (CH_SESSION_STALE_CONTACT_SEC) readslost. A session lease lasts 300 seconds (CH_SESSION_LEASE_TTL_SECONDS). Work prefers its old device for 5 minutes (CH_DEVICE_AFFINITY_GRACE). After that any device can take it. A device reconnects after 1 second and doubles the wait up to 60 seconds. - Failure classes: Device (power, network, or process). Proxy (a proxy that drops idle sockets looks like device loss). Database (sweepers stall, so no session is marked
lost). - Target: You set the target.
Feeds
A feed is data that flows out of the server after an event. Record its delay, not only its availability.
| Feed | Delay the product produces | Cause | Where data can be lost | How to detect |
|---|---|---|---|---|
| Activity | Under a second on a connected browser. After a dropped socket, up to 30 seconds of reconnect wait, then a full refetch. | The server pushes each event after its database commit. An event written inside a larger transaction is pushed only once that transaction commits. The browser refetches every live list on reconnect. | None. The database row is the record. A missed push shows on the next refetch. | Compare an event’s stored time with the time it shows in the SPA. |
| Audit export: OpenTelemetry traces | About 5 seconds in normal use. Up to about 65 seconds while the receiver is down. | The exporter batches spans and sends every 5 seconds. Each send has a 10-second timeout (CH_OTEL_TRACES_TIMEOUT). |
Yes. The queue holds 2048 spans. A full queue drops spans. The exporter retries a failed send for up to 1 minute. Then it drops the batch and counts the failure. | Poll GET /v1/instance/metrics/trace-export. Alert when the failure count rises. |
| Audit export: webhooks | Seconds for a delivery that succeeds first time. | A delivery is tried up to 6 times. The wait starts at 30 seconds and doubles up to 1 hour. A subscription turns itself off after 15 failed deliveries in a row. | Yes. A delivery that fails 6 times does not reach your receiver unless you replay it. | Check the digest chain for a gap. See Verify the audit export. Digests arrive about every 15 minutes and cover deliveries older than 5 minutes. |
| Device status | Offline shows within 75 seconds. Load metrics are at most 30 seconds old while the device is healthy. | The server drops a silent tunnel after 75 seconds. A device sends load metrics every 30 seconds (--metrics-interval). |
None. Status is current state, not history. | The device list marks metrics stale after 5 minutes. It marks capabilities stale after 45 minutes. |
The audit export can lose data. Set your freshness target for the receiver, not for the server. Alert on a rising trace failure count. Check the webhook digest chain on a schedule.
You set the target for each feed.
Dependency failure classes
Classify each incident by the dependency that failed first. The class tells you which check to read.
| Dependency | Journeys affected | Where to look |
|---|---|---|
| Database | All journeys. Sweepers stall, so device loss goes unnoticed. | GET /v1/instance/metrics/db, /v1/instance/metrics/slow-queries, and the worker board. |
| Identity provider | Sign-in. | ch instance check, item sign_in. |
| AI provider | Task queue to session start. Running sessions stall. | The task wait reason in ch task show. |
| Git host | Task queue to session start. Merge steps. | The task wait reason. |
| Device | Session attach. Session start. Device-loss recovery. | ch instance check, item device_connection. The device list. |
Related guides
Last updated