# Self-hosted backup and recovery

Source: https://codeherder.com/docs/self-host-backups/

Escrow the four at-rest keys, back up and copy the database and attachments, watch the recoverable point, rehearse a restore, and promote it safely.

A self-hosted CodeHerder server keeps state in three places: the PostgreSQL database, four at-rest keys, and an attachments bucket. A backup that misses one of them cannot restore a working server. Your organization owns the backups, their retention, and their access control. CodeHerder does not run them for you.

## What to back up

| What | Where it lives | What you lose without it |
| --- | --- | --- |
| PostgreSQL database | The server named in `--db` | Every task, session, member, and setting |
| The four at-rest keys | Environment variables on the app host | Sealed webhook and integration secrets, and sign-in codes and OAuth flows that are in flight |
| Attachments bucket | The S3 bucket named in `CH_S3_BUCKET` | Every uploaded file and embedded image |
| KMS key for sealed variables | The AWS KMS key named in `CH_KMS_KEY_ID` | Every value stored with `ch variable` as a secret |
| Server configuration | Your environment file and unit files | The settings you chose, such as identity mode and pool sizes |

The four keys are environment variables. They are not files. Each key protects a different thing:

- `CH_WEBHOOK_SECRET_KEY` seals the secrets of outbound webhook subscriptions and inbound webhook endpoints.
- `CH_INTEGRATIONS_SECRET_KEY` seals the secrets of integration connections.
- `CH_OTP_HMAC_KEY` signs sign-in codes. If you lose it, codes that are outstanding stop working. Nothing stored is lost.
- `CH_INTEGRATIONS_OAUTH_STATE_KEY` signs OAuth state. If you lose it, OAuth flows that are in flight fail. Nothing stored is lost.

The first two keys are the ones that cannot be replaced. Treat all four the same way, so that a restore never has to guess.

The KMS key is a fifth dependency. It is not one of the four keys. A restore needs that KMS key to exist, and the app host role must be able to use it. Do not schedule the KMS key for deletion while any backup still needs it.

## Escrow the four keys before first boot

Escrow means you keep a safe copy of each key somewhere other than the app host. Do this before the server starts for the first time. After the server seals its first secret, a lost key cannot be rebuilt.

1. Generate the four keys. Use `openssl rand -hex 32` for each one, as in [Start the server](https://codeherder.com/docs/self-hosting/#start-the-server).
2. Store all four in a secret manager outside the app host, for example AWS Secrets Manager or a vault. Store one more copy offline, under the control of two people.
3. Read each key back from the secret manager. Compare its `sha256sum` with the value you generated. Do the same for the offline copy.
4. Only now export the keys on the app host and start the server.
5. Write down where each copy lives in your operator runbook. Name the people who can open the offline copy.

Give the app host role read access to the secret manager entries only. Do not let the same role delete them.

## Never mint a new key on an existing database

Reuse the escrowed values on every new host, every rebuild, and every restore. Never run `openssl rand` again for a database that already holds data.

The server checks all four keys at boot: `CH_WEBHOOK_SECRET_KEY`, `CH_INTEGRATIONS_SECRET_KEY`, `CH_OTP_HMAC_KEY` and `CH_INTEGRATIONS_OAUTH_STATE_KEY`. This is what the check does:

- **Fresh database:** the server records a small test value for each key. The next boot checks it.
- **The database holds sealed secrets, and the key opens at least one:** the server starts. It reads every sealed row. If some rows do not open, it logs an `at-rest key proven, but some stored secrets cannot be decrypted` warning.
- **The database holds sealed secrets, and the key opens none of them:** with `CH_ENV=production`, the server refuses to start. Outside production, the server logs an `at-rest key mismatch` error and starts anyway.
- **The OTP or OAuth-state key does not match its recorded test value:** the server refuses to start in every environment.
- **`CH_AT_REST_KEY_CANARY_RESEAL=1`:** the server starts and adopts the current key. Secrets sealed under the old key stay unreadable. When only a test value failed, one boot re-records it. Remove the setting after that boot.
- **The key opens none of the sealed secrets, and you set `CH_AT_REST_KEY_CANARY_RESEAL=1`:** the setting does not re-record anything for these rows. Keep the setting until you re-enter at least one secret, or restore the original key. Then remove it. If you remove it too early, the next production boot refuses to start again.

Under `CH_ENV=production` the server never invents a key. An unset or short `CH_OTP_HMAC_KEY` or `CH_INTEGRATIONS_OAUTH_STATE_KEY` stops the start. The release job also never creates a key. It stops with `FATAL` when a key file is missing on the host.

A nonfatal problem shows on `GET /v1/health` as the `atRestKeys` field. Its value is `ok`, `mismatch`, `unreadable_rows` or `not_checked`. Alert on any value other than `ok`. The status code stays 200, so the server keeps running while you fix the key.

If the log shows the error, put the original key back and restart.

To replace a key on purpose, re-enter every secret that key protected. Create each webhook secret and each integration connection again. Do this while the server runs on the new key.

## Back up the database

Use one of these methods:

- **Managed snapshots with point-in-time recovery.** On Amazon RDS, turn on automated backups and set a retention period. Seven days is a common start. Recovery can then reach any second inside that period.
- **`pg_dump`.** Run `pg_dump --format=custom` on a schedule. Copy each dump off the database host.

Keep the backups in your own account. Copy them to a second account or region if your risk rules ask for it. A backup in the same account as the database does not survive the loss of that account.

Take a manual snapshot before every [upgrade](https://codeherder.com/docs/self-hosting/#upgrade-to-a-new-release). CodeHerder applies schema migrations at start, so a snapshot is your way back.

## Back up attachments

Attachments live in one S3 bucket, `CH_S3_BUCKET`. Uploaded files sit under `attachments/<workspace_id>/<id>`. Embedded images sit under `images/<workspace_id>/<id>`. Rows in the database point at these objects.

Protect the bucket with one of these:

- S3 versioning, with a lifecycle rule that expires old versions after your retention period.
- Replication to a bucket in another account or region.
- AWS Backup for S3.

Pair the bucket’s restore point with the database’s restore point. A database restored to Monday and a bucket restored to Friday leave links that point at missing or newer objects. Restore both to the same time.

## Keep the revocation journal

A point-in-time restore brings back a credential you revoked after the recovery point. It also brings back a person you erased after that point. CodeHerder therefore writes each API key, device token, CLI installation and MCP grant revocation, and each person erasure, to a journal outside the database. A revoke-all writes one entry per credential. A member deprovision writes one entry too, so a restore disables the member’s login and removes the grants again. An entry holds an id, a kind, the id of the credential or person, and a time. A deprovision entry also holds the id of the workspace root it ran in. It holds no token, name, email or reason.

Choose where the journal lives:

- **A local directory.** Set `CH_REVOCATION_JOURNAL_DIR`. The directory must sit off the database host. Back it up on its own schedule. The server adds files and never rewrites or deletes one.
- **The attachments bucket.** Leave `CH_REVOCATION_JOURNAL_DIR` unset and set `CH_S3_BUCKET`. The journal goes under the `revocation-journal/` prefix. The app host role needs `s3:PutObject`, `s3:GetObject` and `s3:ListBucket` on that prefix. It needs no delete grant.

If you set neither, nothing ships and a restored clone cannot be promoted.

Keep the journal at least as long as your longest database backup. Protect it as you protect the bucket: turn on S3 versioning or Object Lock for the prefix. A change in the last ten seconds before a total loss of the database host may not have shipped yet.

The journal does not cover a device or credential row that someone hard-deleted after the recovery point. It does not delete the person’s sign-in identity in Cognito: the user pool is not part of the database restore. If you also rolled the pool back, delete that identity by hand.

## Recover

Restore in this order. Each step needs the one before it.

1. Fetch the four keys from escrow. Confirm each `sha256sum`.
2. Restore the database from a snapshot or a dump.
3. Fence the restored database before any server connects to it. Run `INSERT INTO restore_fences (reason) VALUES ('restore <database id>')`. Check that `SELECT count(*) FROM restore_fences WHERE promoted_at IS NULL` is at least 1. The fence is not automatic. Without this row, the server sends email and webhooks and accepts devices before the replay.
4. Restore the attachments bucket to the same point in time as the database.
5. Confirm the KMS key is active and the app host role can use it.
6. Set the four keys, `--db`, `CH_S3_BUCKET`, `CH_KMS_KEY_ID`, `CH_REVOCATION_JOURNAL_DIR` if you use it, and the rest of your configuration. Start the server. Run `ch instance restore-fence show`. It must read `fenced: yes`.
7. Read the server log. Look for an `at-rest key mismatch` error. If you see one, stop and go back to step 1.
8. Run `ch instance restore-fence replay --yes`. It re-applies the revocations and erasures made after the recovery point. Run `ch instance restore-fence show` and confirm the journal has no entry left. Then run `ch instance restore-fence promote --yes`. The server refuses to promote until the replay has run. Until the replay has run, a fenced clone answers only `/health` and the restore-fence routes. Every other request, including sign-in, gets 503 `restore_replay_required`. Run `ch instance restore-fence show` again. It must read `fenced: no`. See [Promote a restored database to production](https://codeherder.com/docs/self-host-backups/#promote-a-restored-database-to-production).
9. Sign in. Open a task that has an attachment and confirm the file opens.

## Erased data in backups

An erasure deletes rows from the live database. It does not change a backup. Erased rows stay in each snapshot and dump until that backup expires. A restore brings them back. After you restore, run each erasure again. Set a backup retention period that fits your erasure obligations. See [Self-hosted log retention and access](https://codeherder.com/docs/self-host-log-retention/) for the logs. See [Erased data in backups](https://codeherder.com/docs/self-host-retention-and-dsar/#erased-data-in-backups) for the window of each store.

## Keep recovery copies in a second region

Production snapshots stay in your production AWS account. This is the CodeHerder rule. A copy in another region of the same account is allowed. A copy in another account is not.

Pick one method for the database:

- **Cross-Region automated backup replication.** Run `aws rds start-db-instance-automated-backups-replication --source-db-instance-arn <arn> --kms-key-id <destination-key> --backup-retention-period <days> --region <recovery-region>`. You get point-in-time recovery in the second region.
- **Scheduled snapshot copy.** Run `aws rds copy-db-snapshot --source-region <region> --kms-key-id <destination-key>` from a schedule. You recover to the snapshot time only.
- **AWS Backup copy rule.** Copy to a vault in the recovery region. The vault needs its own key.

Copy the other dependencies the same way:

- **Attachments.** Use S3 Cross-Region Replication to a bucket in the same account. Give the replica its own KMS key.
- **The four keys.** Use the Secrets Manager `replicate-secret-to-regions` option. Keep the offline copy as well.
- **The sealed-variables key.** Make `CH_KMS_KEY_ID` a multi-Region key with a replica in the recovery region. Sealed variables cannot open there without it.

Each copy needs these key grants. Reuse the key policy shape from [Self-hosted database](https://codeherder.com/docs/self-host-database/).

| Who | Key | Grants |
| --- | --- | --- |
| The RDS service | Destination key | Use through `kms:ViaService` set to `rds.<recovery-region>.amazonaws.com` |
| The copy operator role | Source key | `kms:DescribeKey`, `kms:Decrypt` |
| The copy operator role | Destination key | `kms:CreateGrant`, `kms:DescribeKey`, `kms:Encrypt`, `kms:GenerateDataKey*`, `kms:ReEncrypt*` |
| The S3 replication role | Source bucket key | `kms:Decrypt` |
| The S3 replication role | Destination bucket key | `kms:Encrypt`, `kms:GenerateDataKey` |
| The app host role in the recovery region | Sealed-variables key replica | `kms:Decrypt`, `kms:DescribeKey` |

Never schedule deletion of a key that a copy still needs. Your organization approves the regions and the key grants.

## Restore dependencies alongside the database

A database restore alone does not give a working server. Restore each item below to the same point in time.

| Item | Where it lives | Back it up | Check after restore |
| --- | --- | --- | --- |
| The four keys | Secrets Manager and the offline copy | [Escrow the four keys](https://codeherder.com/docs/self-host-backups/#escrow-the-four-keys-before-first-boot) | `sha256sum` matches. `ch instance check` passes `secret_decryption`. |
| Sealed-variables key | The KMS key in `CH_KMS_KEY_ID` | Multi-Region replica | The app host role can decrypt with it. |
| Cognito pool settings | Your user pool | Save the output of `describe-user-pool`, `describe-user-pool-client` for each app client, and `list-identity-providers` and `describe-identity-provider`. Keep it in version control. | `ch instance check` passes `sign_in`. |
| Attachments bucket | `CH_S3_BUCKET` | Versioning or replication | An attachment on a restored task opens. |
| Revocation journal | A directory or the bucket prefix | [Keep the revocation journal](https://codeherder.com/docs/self-host-backups/#keep-the-revocation-journal) | `ch instance restore-fence show` lists the pending entries. |
| Application configuration | Environment file and unit files | Keep them in version control | The server boots with the same settings. |

For Cognito, record the pool id, the region, the domain, and the three app client ids with their callback and logout URLs. Record MFA, threat protection and each identity provider. These map to `CH_COGNITO_USER_POOL_ID`, `CH_COGNITO_REGION`, `CH_COGNITO_DOMAIN`, `CH_COGNITO_SPA_CLIENT_ID`, `CH_COGNITO_CLI_CLIENT_ID`, `CH_COGNITO_MCP_CLIENT_ID` and `CH_COGNITO_PROVIDERS`. See [Self-hosted Cognito](https://codeherder.com/docs/self-host-cognito/) for the server role permissions.

Never delete the user pool. A pool cannot export passwords. A new pool gives each person a new subject id. A person’s row keeps the subject id of the old pool. The server then refuses that person’s sign-in. So keep the original pool.

For the application configuration, record every `CH_*` value you set, `CH_BASE_URL`, the `--db` value, and the release version. Read the version from `/v1/version`. Restore with the same release or a newer one.

## Watch the recoverable point

The recoverable point is the newest time you can restore to. Its lag is the current time minus that point. RDS has no metric for it. Read `LatestRestorableTime` from `aws rds describe-db-instances` instead.

CodeHerder recommends these thresholds. Your organization approves them.

| Signal | Normal | Warn | Page |
| --- | --- | --- | --- |
| Database lag | About 5 minutes | Above 15 minutes | Above 30 minutes |
| Second-region replicated backup lag | About 5 minutes | Above 15 minutes | Above 30 minutes |
| Newest copied snapshot age | Under 24 hours | Above 24 hours | Above 26 hours |
| Attachments `ReplicationLatency` | Under 15 minutes | Above 15 minutes | Any `OperationsFailedReplication` |

Run this every 5 minutes from a scheduler:

```
latest=$(aws rds describe-db-instances --db-instance-identifier "$DB_ID" \
  --query 'DBInstances[0].LatestRestorableTime' --output text)
lag=$(( $(date +%s) - $(date -d "$latest" +%s) ))
aws cloudwatch put-metric-data --namespace CodeHerder/Recovery \
  --metric-name RecoverableLagSeconds --dimensions DBInstance="$DB_ID" \
  --value "$lag" --unit Seconds
```

Then create the alarm. The hosted service’s own RDS alarms use this row shape. Use it for the lag alarm too.

| Name | Metric | Stat | Comparison | Threshold | Periods | Datapoints | Missing data |
| --- | --- | --- | --- | --- | --- | --- | --- |
| recovery-lag-warn | `RecoverableLagSeconds` | Maximum | Greater than | 900 | 3 | 3 | Breaching |
| recovery-lag-page | `RecoverableLagSeconds` | Maximum | Greater than | 1800 | 2 | 2 | Breaching |

A missing datapoint pages. A dead check must not look healthy.

To prove the latest point is usable, run the rehearsal below each month. Restore with `--use-latest-restorable-time`. Record the lag it reached. Keep each evidence record as long as you keep the backups.

## Recovery targets

Two numbers define recovery. The recovery-point objective (RPO) is how much recent data you accept to lose. The recovery-time objective (RTO) is how long you accept to be without service. Each target needs a place where you measure it. Without that place, two people read the same number two ways.

**Where to measure RPO.** Measure from the last recoverable point. RPO is the incident time minus the recovery point you restored. For the database, the recovery point is `LatestRestorableTime` at the moment you restore. For attachments, the measure is the replication latency of the copy. The four escrowed keys need no loss at all. A lost key makes the data it protects unrecoverable.

**Where to measure RTO.** Start the clock when the outage starts. That is the first failed customer journey or the first page alarm, whichever is earlier. Stop the clock when a signed-in person completes a task. That is T1 in [Rehearse a full restore](https://codeherder.com/docs/self-host-backups/#rehearse-a-full-restore). The drill’s T1 − T0 is only the restore part. It leaves out detection and the decision to restore. Add both.

These defaults come from the measured drill. Your organization approves them.

| Scenario | RPO | RTO | Basis |
| --- | --- | --- | --- |
| Database loss in the primary region | 15 minutes | 4 hours | The measured recovery-point lag is about 5 minutes. The warn threshold is 15 minutes. |
| Loss of the primary region, with replicated automated backups | 15 minutes | 8 hours | The replicated backup lag has the same thresholds. |
| Loss of the primary region, with a scheduled snapshot copy | 26 hours | 8 hours | The newest copied snapshot is at most 26 hours old. |
| Attachments | 15 minutes | With the database | The `ReplicationLatency` warn threshold. |
| Escrowed keys | Zero | With the database | Escrow them before first boot. |

The in-region RTO of 4 hours has this budget:

| Step | Budget |
| --- | --- |
| Detect the outage | 30 minutes |
| Decide to restore | 30 minutes |
| Database reaches `available` | 25 minutes (measured: 875 to 1402 seconds) |
| Keys, bucket, fence, replay, promote and verify | 90 minutes |
| Margin | The rest |

The region-loss RTO adds the time to build the application host and the identity provider connection in the second region. Measure it in your first region-loss drill.

Replace a default with your own first full-drill measurement. A target below the measured complete recovery time plus detection and decision time is not valid. Raise the target, or shorten the steps and drill again.

### Rule text for approval

This text is a template. Adapt it to your policy.

> **Recovery target rule.**
>
> 1. The recovery targets are the table above, unless the customer replaces a row with its own measured value.
> 2. RPO is measured from the last recoverable point. RTO is measured from the start of the outage to a signed-in person completing a task.
> 3. The customer runs the full-restore drill each month and records the outage start and the result against the target in the drill record.
> 4. A drill that misses a target opens an action with an owner and a date. The target stays until the action closes or the customer approves a new target.
> 5. The customer approves its targets. CodeHerder does not set them for the customer.

## Rehearse a full restore

A backup you have never restored is a guess. Run this drill in the production account on a separate drill host. A drill never serves production traffic.

The fence blocks email, invitations, webhooks, device connections and agent spawns on the clone. It does not block Cognito admin calls, S3 writes or git-host writes. So the drill uses its own bucket and its own base URL.

1. Record the retention and `LatestRestorableTime` from `aws rds describe-db-instances`.
2. Restore with `aws rds restore-db-instance-to-point-in-time`. Set all six flags: `--use-latest-restorable-time`, `--db-subnet-group-name`, `--vpc-security-group-ids`, `--no-publicly-accessible`, `--db-parameter-group-name` and `--no-deletion-protection`. Name the clone with a `drill-` prefix. Start the clock (T0) when the call returns.
3. When the clone is `available`, fence it before any server connects: `INSERT INTO restore_fences (reason) VALUES ('drill <clone id>')`. Check that `SELECT count(*) FROM restore_fences WHERE promoted_at IS NULL` is at least 1.
4. Restore the attachments to a drill bucket at the same point in time. Do not use the production bucket.
5. Start a drill server from the same release. Use the escrowed keys, `--db` set to the clone, `CH_S3_BUCKET` set to the drill bucket, and a drill `CH_BASE_URL`. Copy the production revocation journal into a drill directory, and set `CH_REVOCATION_JOURNAL_DIR` to that directory. If production keeps the journal in its bucket, copy the `revocation-journal/` prefix of that bucket into the drill directory with `aws s3 sync`. Replay only reads the journal, and a fenced clone never ships to it, so the copy stays unchanged. Add the drill sign-in callback URL to the SPA app client.
6. Run `ch instance restore-fence show`. It must read `fenced: yes`. The log must show no `at-rest key mismatch`.
7. Run `ch instance restore-fence replay --yes`. Until it runs, the clone answers only `/health` and the restore-fence routes, so sign-in fails. Run `ch instance restore-fence show` and confirm the journal has no entry left. The clone holds real customer data, so this keeps every later revocation in force on it.
8. After replay and before sign-in, confirm that a key revoked after the recovery point gets 401, and that `ch instance restore-fence show` still reads `fenced: yes`. Run `ch instance check --skip mail --skip device_connection`. It must pass.
9. A person signs in through the web app. They open a task that has an attachment and confirm the file opens. They move a task to `done`. Stop the clock (T1) when the task reads `done`.

Do not do these on a drill clone: promote it, point its journal at the production journal store itself (use the copy), deprovision or erase a person, merge a request, or enrol a device. Cognito calls are not fenced.

Record the result.

| Field | Value |
| --- | --- |
| Date |  |
| Outage start (a real outage, or the drill’s simulated start) |  |
| Recovery point |  |
| Recovery-point lag |  |
| Time to `available` |  |
| Complete recovery time (T1 − T0) |  |
| Time from outage start to T1 |  |
| Met the RPO and RTO targets (yes or no) |  |
| Release |  |
| Person who ran it |  |

## Clean up a drill

This procedure deletes only the drill clone, the drill bucket and the drill host. Never touch the source database.

1. Read the identifier back from `describe-db-instances`. Do not use a variable you did not read back.
2. Confirm it starts with `drill-`.
3. Confirm `restore_fences` on that clone holds a row with `promoted_at IS NULL`.
4. Ask a second person to confirm the identifier.
5. Run `aws rds delete-db-instance --skip-final-snapshot` on that identifier. Then delete the drill bucket and the drill host.

## Promote a restored database to production

This procedure replaces production. It is separate from the drill cleanup. Never run the drill cleanup on a promoted target.

1. Restore the database. Follow [Recover](https://codeherder.com/docs/self-host-backups/#recover) for the order of keys, bucket and configuration.
2. Fence the restored database before any server connects to it. Run `INSERT INTO restore_fences (reason) VALUES ('restore <database id>')`. Check that `SELECT count(*) FROM restore_fences WHERE promoted_at IS NULL` is at least 1.
3. Start the server. Run `ch instance restore-fence show`. It must read `fenced: yes`. Verify the data.
4. Run `ch instance restore-fence replay --yes`. Run `ch instance restore-fence show` and confirm no entry is pending.
5. Run `ch instance restore-fence promote --yes`. Run `ch instance restore-fence show`. It must read `fenced: no`.
6. Turn on deletion protection on the recovered target: `aws rds modify-db-instance --deletion-protection`. This step is required. A restore made with the drill flags has `--no-deletion-protection`.
7. Only now point DNS and configuration at the recovered database.

Never delete the recovered target. Never delete the old primary until the new one has run a full backup cycle and an owner signs off. One person runs the procedure. A second person confirms the identifier before any delete.

## Related guides

- [Self-hosted deployment](https://codeherder.com/docs/self-hosting/) — download, start, and upgrade the server
- [Sizing a self-hosted deployment](https://codeherder.com/docs/self-host-sizing/) — app host, database, and device fleet sizes
- [Self-hosted logs](https://codeherder.com/docs/self-host-logs/) — where to find the server log that reports a key mismatch
- [Self-hosted key custody](https://codeherder.com/docs/self-host-keys/) — custody record, production key settings, KMS key recovery
- [Secrets and variables](https://codeherder.com/docs/secrets/) — the values that use the KMS key
- [Attachments](https://codeherder.com/docs/attachments/) — what the attachments bucket holds
