CodeHerderSearch⌘KRequest access →

Self-hosted support and incident response

Who runs a self-host incident, what device owners do, severity levels, a notice-time template, the post-incident review, and tests for paging and key leaks.

This page describes how incidents on your self-hosted CodeHerder server are run. Two roles take part:

  • Your instance operator runs the server. The operator holds the host access and the instance operator login. Your organization can be its own operator.
  • Your support contact gives vendor-side support. Your support agreement names this contact.

Your support agreement can assign both roles to one party. If it does, that party does both sets of duties on this page. Your support agreement also says whether an SLA or service credits apply.

Support channel

Send support requests through the channel in your support agreement: ______. Record the support contact’s address in the launch acceptance record, row “Incident contacts”. See Self-hosted responsibilities and launch acceptance.

Send diagnostics only as a support bundle. See Self-hosted support bundle. CodeHerder staff have no access to your install. See Self-hosted operator access and changes.

Support hours

The support contact answers during the hours in your support agreement: ______. Your support agreement also says what happens to a SEV1 outside these hours: ______.

Response targets

Your support agreement sets these targets. Copy them into this table. Leave a cell blank until the agreement sets it.

Level First answer from the support contact Notice and updates
SEV1 ______ See “Severity levels”.
SEV2 ______ See “Severity levels”.
SEV3 ______ See “Severity levels”.
A question with no incident ______ ______

Security escalation

Your support agreement names the security escalation contact: ______. Put “SECURITY” in the subject of the message. A security report is always SEV1. See “Security incidents” below.

Who runs an incident

Your organization names the incident commander. Alarms from your server go to your instance operator. The instance operator brings in the support contact when a product defect may be the cause. Your support agreement can make one party the instance operator, the incident commander and the support contact.

Name an alternate for each role, or write “none”. If a role has no alternate, plan for the case where that person does not answer.

While you wait, you can act alone. Follow the containment steps in Replacing a compromised self-hosted server. They need access to your cloud account. They do not need access from CodeHerder.

Name your own communication owner and any alternate in the evidence records on this page. The communication owner is the person who receives incident notices.

Device-owner duties

A device owner is the person who runs a device server for the install. During an incident, the device owner does these things:

  1. Report a suspect device or credential to the incident commander at once.
  2. Do not wipe or reinstall the device.
  3. Stop the device server when the incident commander tells you to.
  4. Keep the local logs of the device server.
  5. Rotate the git-host and AI credentials held on the device.
  6. Re-enrol the device only when the incident commander says so.

The instance operator or an admin can cut off a device with ch device disable <device>. Rotate its token with ch device rotate-token <device>.

Severity levels

The incident commander sets the level. Your communication owner can raise it. Your support agreement sets the notice times and the update cadence. Copy them into this table.

Level Criterion Examples The communication owner is told within Update cadence
SEV1 Customer data or a credential may be exposed, or nobody can use the server. Suspected or confirmed compromise. A leaked credential. Customer data exposed. The server is down for everyone. ______ ______
SEV2 One key function fails for many users, and no data is exposed. A key function is degraded. Sign-in fails. Devices cannot connect. Audit export fails. Backups fail. ______ ______
SEV3 One user or a minor function fails, and a workaround exists. A minor issue, or an issue for one user with a workaround. ______ ______

Customer-impacting incidents

An incident is customer-impacting when it stops or degrades a function your users rely on. It is also customer-impacting when it exposes or changes your data outside the product’s rules. The severity table above sets the level.

A defect that affects nobody on your install is not an incident. The release notes describe it.

Security incidents

A security incident is a confirmed or suspected event where someone reads, changes or deletes your data or credentials without permission. It is also a security incident when a product boundary fails: the workspace, device or credential boundary.

A security incident is always SEV1. For a compromise of the host, see Replacing a compromised self-hosted server.

Vendor notification clock

Your organization is the data controller. The vendor clock is the time it takes the support contact to tell you about a product defect that can expose your data. Your support agreement sets the times. Copy them here.

The clock starts when the support contact first has reason to believe that a defect or incident can expose your data. That includes its own finding, a report and an alarm.

  • The support contact tells your communication owner within ______.
  • The notice says what is known, which versions are affected, which data classes may be exposed, and what to do now.
  • The support contact sends an update every ______ until the incident is contained.
  • For a vulnerability with no sign of use, the support contact tells you with the fixed release, and within ______ of confirming it at the latest.

Your organization owns its own regulatory clock and notices. See the section “What your organization owns” below.

Security signals

These events and log lines can signal a security incident. Watch for them in your logs. See Self-hosted logs.

Signal Where it appears What it means
device.cross_device_frame_dropped Event log (SQL) A device sent a frame for another device. The server dropped it.
device.cross_device_provision_dropped Event log (SQL) A device sent a provision request for another device. The server dropped it.
device.cross_workspace_frame_dropped Event log (SQL) A device sent a frame for a workspace it is not linked to.
device.cross_workspace_provision_dropped Event log (SQL) A device sent a provision request for a workspace it is not linked to.
inbound_webhook.delivery_rejected Event log (SQL) An inbound webhook delivery failed its checks.
sso.connection_disabled Event log (SQL) Someone disabled a SAML connection.
sso.domain_unverified Event log (SQL) A domain no longer routes to sign-in.
sso.scim_token_minted Event log (SQL) Someone minted a SCIM credential.
member.deprovisioned Event log (SQL) A member lost their grants, keys and devices.
device.operator_capabilities_changed Event log (SQL) An operator changed what a device can do.
device.secrets_acknowledged Event log and webhook export A device owner changed the secrets a device may hold.
inbound_webhook.secret_rotated Event log and webhook export Someone rotated an inbound webhook secret.
webhook.disabled Event log and webhook export A webhook stopped. If it fed your SIEM, the audit feed stopped.
instance.human_revoke_all codeherder.service journald An operator revoked every credential of one person.
at-rest key mismatch codeherder.service journald The at-rest key cannot open stored secrets. The server booted because this is not production, or because the reseal override is set.
at-rest key proven, but some stored secrets cannot be decrypted codeherder.service journald The key opens some stored secrets and not others. GET /v1/health shows atRestKeys as unreadable_rows.
accepting an at-rest key mismatch because CH_AT_REST_KEY_CANARY_RESEAL is set codeherder.service journald An operator accepted a key change. Old secrets stay unreadable.
auth.failed codeherder.service journald The server rejected a credential. One line per class, address and reason each minute, with an attempts count.
cognito.jwt rejected codeherder.service journald The server refused a sign-in token. Many refusals can mean an attack.

GET /v1/health also carries an atRestKeys field: ok, mismatch, unreadable_rows or not_checked. Alert on any value other than ok. The status code stays 200.

Repeated failed authentication shows as many status=401 request lines from one address. It also shows as auth_audit_log rows with outcome = 'failed' and a high attempts count. See “Sign-in outcomes” in Self-hosted logs.

The OCSF export marks a listed event with severity_id 3.

CodeHerder keeps this list in the product code. A test fails if this table and the code differ.

What your organization owns

Your organization owns its own breach assessment. It owns notices to regulators, to lawyers and to its own people. CodeHerder does not send these notices.

The instance operator and the support contact give facts on request: the timeline, the affected versions, and what data the server holds. See Self-hosted support bundle and Self-hosted logs.

Post-incident review

For a SEV1 or SEV2 incident, the incident commander writes a short review. Your support agreement sets the deadline and says whether the support contact contributes: ______. The review goes to your communication owner. It has these fields:

  • Timeline
  • Impact
  • Cause
  • What worked
  • What failed
  • Actions, each with an owner and a date

Error budget and release response

Your organization sets its own error budget. An error budget is the amount of unreliability you accept in a period. Your organization also decides what to do when it is spent, for example to hold upgrades until service is stable again. CodeHerder runs no release freeze on your server. You decide when to upgrade.

CodeHerder gives you what you need to decide:

Two limits apply. A release you hold past 90 days leaves support. A security upgrade-by date overrides a hold. See When a security fix ships.

Rule text for approval

This text is a template. Adapt it to your policy.

Incident rule.

  1. The incident commander is ______. Alarms go to the instance operator, ______. The vendor contact is the support contact, ______.
  2. The incident commander sets the severity level. The customer communication owner can raise it.
  3. For SEV1, the communication owner is told within ______. For SEV2, within ______. For SEV3, within ______.
  4. The customer owns its breach assessment and its legal and regulatory notices. The instance operator and the support contact give facts on request.
  5. For SEV1 and SEV2, the incident commander sends a written review to the communication owner within ______.
  6. A security incident is a confirmed or suspected event where someone reads, changes or deletes customer data or credentials without permission, or where a product boundary fails. It is always SEV1.
  7. The vendor clock starts when the support contact first has reason to believe a defect or incident can expose customer data. The support contact tells the communication owner within ______. For a vulnerability with no sign of use, the support contact tells the customer with the fixed release, and within ______ of confirming it at the latest.

Support rule.

  1. The support contact is ______. The channel is ______.
  2. Support hours are ______. Outside these hours, a SEV1 gets ______.
  3. The response targets are ______. Service credits: ______.
  4. The security escalation contact is ______. A message about a security event carries “SECURITY” in the subject.
  5. The support bundle is the only diagnostic channel. CodeHerder staff have no access to the customer’s install.

Error budget rule.

  1. The customer sets its own error budget and its own release response. CodeHerder sets neither and runs no release freeze on a self-hosted server.
  2. CodeHerder supplies the service objectives page, a migration notice in each release manifest, the rollback procedure and a 90-day support window for each release.
  3. A release held past 90 days is outside support. A security upgrade-by date overrides a hold.

Test the paging path

Run this test before launch. Your instance operator runs it with your communication owner. It proves that an alarm reaches the person on call and that this person acknowledges it.

  1. Pick the alarm source. For CloudWatch, run aws cloudwatch set-alarm-state --alarm-name <name> --state-value ALARM --state-reason "paging test". For the in-app route, fire an alert rule. See Alerts.
  2. Write down the time the alarm fired.
  3. Write down the time the alert reached the person on call.
  4. The person on call acknowledges the alert. Write down that time.
  5. Test the no-answer case. The person on call does not acknowledge for the agreed window. Record what your routing does next.
  6. Reset the alarm.

Fill in this record for each test.

Field Value
Date
Who ran it
Alarm source
Local routing from alarm to the person on call
Alternate named, or “none”
Fired at
Delivered at
Acknowledged at
No-answer result
Pass or fail
Notes

Rehearse a leaked API key

Run this scenario before launch. Your instance operator, your communication owner and one device owner take part. The incident commander leads. Bring in the support contact if your support agreement includes rehearsals. Act in time order.

Step Your organization The instance operator Device owner
Detect The communication owner receives a report, or a scan finds a key in a public repository. Acknowledges the report. The incident commander sets SEV1. Reports a key seen in a log or a repository.
Contain Keeps the option to isolate the host. See Replacing a compromised self-hosted server. Lists the owner’s keys with ch member keys list <member>. Deletes the key with ch member keys delete <hash>. If the owner is a person and the scope is unclear, runs ch human revoke-all <human> --yes. If misuse continues, pauses spawns with POST https://<your-server>/v1/instance/spawn-pauses/fleet/pause. A pause stops new work only. Stops the device server when told. Keeps local logs.
Investigate Reads its own cloud and proxy logs. Exports the instance audit log. See Self-hosted logs. Sends a support bundle to the support contact if a product defect may be the cause. Gives the device logs to the instance operator.
Tell Receives the notice within the time in your support agreement. Starts its own breach assessment. Tells the communication owner. Gives facts on request.
Recover Accepts the recovery. Mints a new key. Lifts the pause with POST https://<your-server>/v1/instance/spawn-pauses/fleet/unpause. Rotates git-host and AI credentials. Re-enrols the device when told.
Review Reads the review and accepts it. The incident commander writes the post-incident review. Adds notes to the review.

Fill in this record for each rehearsal.

Field Value
Date
Instance operator
Incident commander
Communication owner
Device owner
Start
End
Time to contain
Time the communication owner was told
Support contact involved (yes or no)
What went wrong
Actions, with owner
Accepted by (your organization)
Pass or fail

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close