CodeHerderSearch⌘KRequest access →

Self-hosted service objectives

Measure each customer journey and feed on a self-hosted server, see the operating ranges CodeHerder states, and set your own targets from them.

This page helps you set service objectives for a self-hosted CodeHerder server. It states the timings the product produces. It shows how to measure each journey. It names the dependency failures that affect each one.

CodeHerder commits only to the supported ranges on this page. You set the target. Every target cell says “You set the target”.

A range comes from a setting or a fixed value in the server. It is not a promise about your network, database, or identity provider. Where no fixed value exists, this page says so. It does not invent a number.

Journeys

Each journey has an indicator to record. Record it from your proxy log, from a scheduled check, or from the worker board. Do not use the health endpoint alone. GET /v1/health proves the process and the database answer. It does not prove a journey works.

API request

  • How to measure: Read status and duration for each request in your proxy access log. See Self-hosted logs. The server also sends a Server-Timing header on each response. Set CH_SERVER_TIMING=false to turn it off.
  • Indicator to record: The share of requests that return a status below 500. The 95th percentile of duration, grouped by route.
  • Supported range: CodeHerder states no latency range yet. Use the load test results for the launch workload profile when they ship. Until then, measure your own baseline.
  • Failure classes: Database (slow queries, pool wait, GET /v1/instance/metrics/db). Identity provider (sign-in requests only).
  • Target: You set the target.

SPA load

  • How to measure: Time a page request through your proxy. Run a synthetic browser check against the sign-in page from outside your network.
  • Indicator to record: The time until the page returns 200. The share of checks that pass.
  • Supported range: CodeHerder states no latency range yet. Time it from outside your network, through your proxy.
  • Failure classes: Proxy or app host. The page itself needs no AI provider, git host, or device.
  • Target: You set the target.

Sign-in

  • How to measure: Run ch instance check on a schedule. See Self-hosted synthetic checks. The sign_in, mail and secret_decryption items each report pass or fail.
  • Indicator to record: The result of each item, and the time the command takes.
  • Supported range: CodeHerder states no latency range yet. Sign-in time depends mostly on your identity provider.
  • Failure classes: Identity provider (sign_in fails, see Self-hosted Cognito sign-in). Mail (mail fails, so one-time codes do not arrive). Database.
  • Target: You set the target.

Session attach

Attach is the live terminal view of a running session. It uses a WebSocket.

  • How to measure: Open a session from a scheduled check and record whether the socket connects. Count WebSocket closes in your proxy log.
  • Indicator to record: The share of attach attempts that connect. The number of unplanned closes per hour.
  • Supported range: The server pings each socket every 25 seconds. It gives a write 10 seconds to finish. It checks the caller’s access again every 50 seconds. A revoked credential stops working within 60 seconds. The browser reconnects after 1 second, then doubles the wait up to 30 seconds. It resets the wait after a connection holds for 5 seconds.
  • Failure classes: Device (the session runs on a device, so a lost device ends the attach). Proxy (an idle timeout below 25 seconds cuts the socket). Database (the access check fails).
  • Target: You set the target.

Task queue to session start

  • How to measure: Compare the time a task became ready with the time its session started. Both stamps are on the task and session records. ch task show prints the wait reason for a task that is not running. See Why isn’t my task moving?.
  • Indicator to record: The wait from ready to session start, per workflow. The count of tasks with a wait reason for longer than your target.
  • Supported range: Task creation and stage change staff work at once when a device has room. The backstop reconcile pass runs every 30 seconds for staffing and capacity. A slower pass runs every 5 minutes. So the backstop retries a missed task about every 30 seconds. A device runs at most 4 sessions by default (CH_MAX_DEVICE_SESSIONS). A failed spawn retries after 30 seconds, then backs off to a cap of 15 minutes.
  • Failure classes: Device (none online, or none with the needed capability). AI provider (the wait reason names it, or spawns fail and retry). Git host (clone fails, so the spawn retries). Database.
  • Target: You set the target.

Device-loss recovery

  • How to measure: Stop a test device and time how long its session takes to read lost. Poll GET /v1/instance/metrics/workers for the sweeper that marks sessions lost. See Monitor background workers.
  • Indicator to record: The time from the last device frame to the lost state. The time until the task has a new session on another device.
  • Supported range: A device sends a heartbeat every 30 seconds. The server drops a tunnel silent for 75 seconds. The sweeper marks a session lost when its device tunnel stays offline for 90 seconds. The sweeper runs every 30 seconds. So a session reads lost 90 to 120 seconds after the last frame. If the tunnel state is unclear, a session with no device contact for 900 seconds (CH_SESSION_STALE_CONTACT_SEC) reads lost. A session lease lasts 300 seconds (CH_SESSION_LEASE_TTL_SECONDS). Work prefers its old device for 5 minutes (CH_DEVICE_AFFINITY_GRACE). After that any device can take it. A device reconnects after 1 second and doubles the wait up to 60 seconds.
  • Failure classes: Device (power, network, or process). Proxy (a proxy that drops idle sockets looks like device loss). Database (sweepers stall, so no session is marked lost).
  • Target: You set the target.

Feeds

A feed is data that flows out of the server after an event. Record its delay, not only its availability.

Feed Delay the product produces Cause Where data can be lost How to detect
Activity Under a second on a connected browser. After a dropped socket, up to 30 seconds of reconnect wait, then a full refetch. The server pushes each event after its database commit. An event written inside a larger transaction is pushed only once that transaction commits. The browser refetches every live list on reconnect. None. The database row is the record. A missed push shows on the next refetch. Compare an event’s stored time with the time it shows in the SPA.
Audit export: OpenTelemetry traces About 5 seconds in normal use. Up to about 65 seconds while the receiver is down. The exporter batches spans and sends every 5 seconds. Each send has a 10-second timeout (CH_OTEL_TRACES_TIMEOUT). Yes. The queue holds 2048 spans. A full queue drops spans. The exporter retries a failed send for up to 1 minute. Then it drops the batch and counts the failure. Poll GET /v1/instance/metrics/trace-export. Alert when the failure count rises.
Audit export: webhooks Seconds for a delivery that succeeds first time. A delivery is tried up to 6 times. The wait starts at 30 seconds and doubles up to 1 hour. A subscription turns itself off after 15 failed deliveries in a row. Yes. A delivery that fails 6 times does not reach your receiver unless you replay it. Check the digest chain for a gap. See Verify the audit export. Digests arrive about every 15 minutes and cover deliveries older than 5 minutes.
Device status Offline shows within 75 seconds. Load metrics are at most 30 seconds old while the device is healthy. The server drops a silent tunnel after 75 seconds. A device sends load metrics every 30 seconds (--metrics-interval). None. Status is current state, not history. The device list marks metrics stale after 5 minutes. It marks capabilities stale after 45 minutes.

The audit export can lose data. Set your freshness target for the receiver, not for the server. Alert on a rising trace failure count. Check the webhook digest chain on a schedule.

You set the target for each feed.

Dependency failure classes

Classify each incident by the dependency that failed first. The class tells you which check to read.

Dependency Journeys affected Where to look
Database All journeys. Sweepers stall, so device loss goes unnoticed. GET /v1/instance/metrics/db, /v1/instance/metrics/slow-queries, and the worker board.
Identity provider Sign-in. ch instance check, item sign_in.
AI provider Task queue to session start. Running sessions stall. The task wait reason in ch task show.
Git host Task queue to session start. Merge steps. The task wait reason.
Device Session attach. Session start. Device-loss recovery. ch instance check, item device_connection. The device list.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close