# Replacing a compromised self-hosted server

Source: https://codeherder.com/docs/self-host-compromise/

Isolate a suspect self-hosted server, preserve evidence and its chain of custody, build a clean replacement, rotate credentials, and plan emergency host access.

This page helps you respond when you suspect your self-hosted CodeHerder server is compromised. Your organization owns the host and does the work. CodeHerder supplies this guidance only. CodeHerder has no access to your install. To ask for help, use the channels on the [support bundle](https://codeherder.com/docs/support-bundle/) page.

The steps run in this order: contain, preserve evidence, build a clean host, rotate credentials. Read the whole page before you start. Plan the emergency access path in [Plan emergency host access](https://codeherder.com/docs/self-host-compromise/#plan-emergency-host-access) before you need it.

## Contain the host

1. Keep the instance running. Do not terminate it. Do not reboot it. Memory and disk are evidence.
2. Stop new work first, while the API still answers. Pause spawns for the whole fleet over HTTP: `POST https://<your-server>/v1/instance/spawn-pauses/fleet/pause`. Devices then get no new work.
3. Isolate the network. Replace the instance security group with one that has no inbound rules. Limit outbound traffic to your evidence destination only. The reference Terraform module (`aws_security_group.app`) opens no SSH port. Use the equivalent step in your own infrastructure.
4. Detach the instance role, or revoke its active sessions. The role can call KMS Decrypt, S3 and SES.
5. Capture volatile state (see the next section, step 1). Then stop the services: `systemctl stop codeherder` and `systemctl stop caddy`.

A running server can keep sending commands to devices. If you doubt the order, isolate the network first, capture second, and stop the services last.

## Preserve evidence before replacement

Collect the items below in this order. The first items are lost fastest. Record the sha256 of every item in the custody log when you collect it. See [Chain of custody](https://codeherder.com/docs/self-host-compromise/#chain-of-custody).

1. **Volatile state.** The process list, open network connections, `systemctl status codeherder`, and the sha256 of the running binary (`sha256sum /proc/<pid>/exe`). Do this only if your policy allows live response.
2. **Disk.** An EBS snapshot of the root volume and of any data volume, tagged with the case ID. Or a disk image on other infrastructure.
3. **journald.** Export `journalctl -u codeherder.service`, `-u caddy.service`, the auth and authpriv facilities, and the audit transport. Use `-o export` or `-o json`. The local journal is capped at 500 MB and two weeks, so capture early. Add any off-host copies you already hold. See [Self-hosted logs](https://codeherder.com/docs/self-host-logs/).
4. **Proxy access logs.** The Caddy access log, from journald or from your log shipper.
5. **Instance audit export.** The `admin_audit_log` rows with `plane = 'instance'`, exported as CSV with SQL. See [Self-hosted logs](https://codeherder.com/docs/self-host-logs/). Add any audit webhook deliveries you receive.
6. **Release evidence.** The manifest checksum is for the release archive, not for the binary. Do not compare the binary hash with the manifest.
  1. Get `manifest.json`, `manifest.json.sig` and the archive for the running version. The download route serves only the current release. If a newer release superseded yours, use the copies you kept at install time. Keep each manifest, signature and archive you install.
  2. Verify the signature and the archive checksum. See [Verify the signature, then the checksum](https://codeherder.com/docs/self-hosting/#verify-the-signature-then-the-checksum). Save the `openssl dgst -verify` output.
  3. Extract the archive on a clean machine. Record the sha256 of the extracted `codeherder` binary.
  4. Record the sha256 of the running binary (`sha256sum /proc/<pid>/exe`) and of the binary on disk. Use the path your unit file runs. In the reference unit it is `/usr/local/bin/codeherder`.
  5. Compare both hashes with the extracted binary. A mismatch is evidence of tampering.
  6. Save the output of `https://<your-server>/v1/version` if the API still answers.
7. **Device state.** The output of `ch device list` and `ch device show <id>` for each connected device. Also the local logs of each device server, collected by the device operator. A compromised server can run code on devices. Treat each device that connected in the window as suspect and re-enrol it.
8. **Cloud control plane.** CloudTrail events for the instance role. SSM session records, if you use SSM. KMS Decrypt calls in the window.

## Chain of custody

This rule text is a proposal for your security owner to approve. Adapt it to your policy.

> **Custody rule.**
>
> 1. One named custodian, the security owner, holds all evidence. A named alternate may act for them.
> 2. Store evidence in write-once or access-logged storage. Keep it apart from the compromised account and its roles. Encrypt it at rest.
> 3. Grant read access by name only: the custodian, the incident commander, and named investigators. Log every access.
> 4. Keep a custody log. For each item, record its ID, description, sha256, who collected it, when, from where, and where it is stored. For each hand-over, record who gave it, who received it, when, and why.
> 5. Never analyse the original. Work on a copy. Check the copy’s hash first.
> 6. Sending evidence to CodeHerder support is your choice and is never required.
> 7. Before any transfer, remove or redact secrets (key files, database connection strings, `.env` files) and any customer content CodeHerder does not need. Send the smallest set that answers the question.
> 8. Use a support bundle by default. Send raw evidence only when the security owner approves it in writing. Use an encrypted channel that you pick. Record the hand-over in the custody log.
> 9. CodeHerder keeps transferred evidence only for the length of the incident. CodeHerder deletes it when the incident closes, or when you ask, and confirms the deletion in writing.
> 10. Keep the original evidence for the period that your legal or retention policy sets.

## Build a clean replacement

1. Build a new host from a known-good base image. Do not clean the old host.
2. Download the release. Verify the signature, then the checksum. See [Verify the signature, then the checksum](https://codeherder.com/docs/self-hosting/#verify-the-signature-then-the-checksum). If CodeHerder tells you a release is compromised, follow the note on that page.
3. Restore configuration from your source of truth, not from the old disk. This covers the unit files, the proxy configuration and the environment drop-ins.
4. Restore the four key values from your escrow, not from the suspect host: `CH_WEBHOOK_SECRET_KEY`, `CH_INTEGRATIONS_SECRET_KEY`, `CH_OTP_HMAC_KEY` and `CH_INTEGRATIONS_OAUTH_STATE_KEY`. A fresh value for any of them makes the server refuse to start under `CH_ENV=production`, unless `CH_AT_REST_KEY_CANARY_RESEAL=1` is set.
5. Get a new TLS certificate. Do not copy `/var/lib/caddy` from the compromised host. This differs from the advice for a normal replacement.
6. Start the server. Confirm that `https://<your-server>/v1/version` answers and that you can sign in.
7. Rotate the credentials in the next section before you lift the fleet spawn pause.

## Rotate credentials

The attacker may have held every key and secret on the old host. Rotate all of them.

| Secret | How | After-effect |
| --- | --- | --- |
| Database app role password (in the connection string) | Change it in PostgreSQL. Update the server environment file. Restart. | None once the server restarts. |
| `CH_INTEGRATIONS_SECRET_KEY` and `CH_WEBHOOK_SECRET_KEY` | No bulk re-seal exists. Set the new key. Boot once with `CH_AT_REST_KEY_CANARY_RESEAL=1`. Re-enter each integration token and each webhook and inbound endpoint secret. Remove the knob. Restart. Then rotate the upstream tokens (GitHub, GitLab, Slack, PagerDuty) at each provider. | Secrets sealed under the old key are unreadable until you re-enter them. The upstream rotation is needed because the attacker held both the key and the ciphertext. |
| `CH_OTP_HMAC_KEY` and `CH_INTEGRATIONS_OAUTH_STATE_KEY` | Generate new values with `openssl rand -hex 32`. Set them. Boot once with `CH_AT_REST_KEY_CANARY_RESEAL=1`. Remove the knob. Restart. | Sign-in codes in flight and OAuth flows in flight fail once. Users start them again. |
| `CH_SLACK_CLIENT_SECRET`, `CH_CLOUDFRONT_EDGE_SECRET` if set, and your identity provider app client secret if one is configured | Rotate at the provider. Update the environment file. Restart. | Slack installs and sign-in through the identity provider stop until the new value is in place. |
| `CH_LICENSE` | It is a signed token, not a key. If it may have leaked, ask CodeHerder support to reissue it. | The old token stays valid until you replace it. |
| Workspace secrets and variables | The instance role could call KMS Decrypt. Rotate each secret value at its source. Then set it again with `ch variable set`. | Anything that used the old value fails until you set the new one. |
| People and devices | For each human, run `ch human revoke-all <human>`. It revokes API keys, MCP grants, device tokens and CLI installations. Rotate each device token with `ch device rotate-token <device>`. Re-enrol each suspect device. | Each person and device must sign in again. |
| Host and cloud | The new host makes its own SSH host keys. Replace the instance role. Revoke the old role sessions. Rotate any IAM access keys that were on disk. | Old sessions stop working. |

After the rotation, check that a device list and a token mint work.

## Plan emergency host access

Normal host access can depend on SSM and on the identity service that gates it. If either is down, you need a second path. Your organization picks that path and tests it. CodeHerder names only the minimum.

The path must not depend on SSM or on your identity provider (for example Cognito or your IdP). Examples, not endorsements: the EC2 serial console, a bastion with hardware-key SSH, or console access under a break-glass IAM user with MFA.

The path must reach all of these:

1. A PostgreSQL role that can connect to the CodeHerder database. The app role is enough to run revocation SQL. A restore needs an owner role. This role is your own break-glass role, and CodeHerder does not audit it.
2. `systemctl stop`, `start` and `restart` for `codeherder` and the proxy.
3. The key files and environment files, to read and to replace.
4. `journalctl`, to fetch logs.

Apply these controls:

- Seal the credentials and store them off the host.
- Require two people to use them.
- Log and alert every use.
- Rotate the credentials after each use.

Test the path before go-live and after any identity or network change. Treat SSM and the identity provider as down during the test. The test passes when a named person reaches items 1 to 4. Record the date, who ran it, and the result.

CodeHerder has no staff path into your install.

## Related guides

- [Self-hosted deployment](https://codeherder.com/docs/self-hosting/) — download, verify, run and upgrade the server
- [Self-hosted logs](https://codeherder.com/docs/self-host-logs/) — log formats and how to ship them off the host
- [Self-hosted support bundle](https://codeherder.com/docs/support-bundle/) — what to send when you ask for help
