Self-hosted backup and recovery
Escrow the four at-rest keys, back up and copy the database and attachments, watch the recoverable point, rehearse a restore, and promote it safely.
A self-hosted CodeHerder server keeps state in three places: the PostgreSQL database, four at-rest keys, and an attachments bucket. A backup that misses one of them cannot restore a working server. Your organization owns the backups, their retention, and their access control. CodeHerder does not run them for you.
What to back up
| What | Where it lives | What you lose without it |
|---|---|---|
| PostgreSQL database | The server named in --db |
Every task, session, member, and setting |
| The four at-rest keys | Environment variables on the app host | Sealed webhook and integration secrets, and sign-in codes and OAuth flows that are in flight |
| Attachments bucket | The S3 bucket named in CH_S3_BUCKET |
Every uploaded file and embedded image |
| KMS key for sealed variables | The AWS KMS key named in CH_KMS_KEY_ID |
Every value stored with ch variable as a secret |
| Server configuration | Your environment file and unit files | The settings you chose, such as identity mode and pool sizes |
The four keys are environment variables. They are not files. Each key protects a different thing:
CH_WEBHOOK_SECRET_KEYseals the secrets of outbound webhook subscriptions and inbound webhook endpoints.CH_INTEGRATIONS_SECRET_KEYseals the secrets of integration connections.CH_OTP_HMAC_KEYsigns sign-in codes. If you lose it, codes that are outstanding stop working. Nothing stored is lost.CH_INTEGRATIONS_OAUTH_STATE_KEYsigns OAuth state. If you lose it, OAuth flows that are in flight fail. Nothing stored is lost.
The first two keys are the ones that cannot be replaced. Treat all four the same way, so that a restore never has to guess.
The KMS key is a fifth dependency. It is not one of the four keys. A restore needs that KMS key to exist, and the app host role must be able to use it. Do not schedule the KMS key for deletion while any backup still needs it.
Escrow the four keys before first boot
Escrow means you keep a safe copy of each key somewhere other than the app host. Do this before the server starts for the first time. After the server seals its first secret, a lost key cannot be rebuilt.
- Generate the four keys. Use
openssl rand -hex 32for each one, as in Start the server. - Store all four in a secret manager outside the app host, for example AWS Secrets Manager or a vault. Store one more copy offline, under the control of two people.
- Read each key back from the secret manager. Compare its
sha256sumwith the value you generated. Do the same for the offline copy. - Only now export the keys on the app host and start the server.
- Write down where each copy lives in your operator runbook. Name the people who can open the offline copy.
Give the app host role read access to the secret manager entries only. Do not let the same role delete them.
Never mint a new key on an existing database
Reuse the escrowed values on every new host, every rebuild, and every restore. Never run openssl rand again for a database that already holds data.
The server checks all four keys at boot: CH_WEBHOOK_SECRET_KEY, CH_INTEGRATIONS_SECRET_KEY, CH_OTP_HMAC_KEY and CH_INTEGRATIONS_OAUTH_STATE_KEY. This is what the check does:
- Fresh database: the server records a small test value for each key. The next boot checks it.
- The database holds sealed secrets, and the key opens at least one: the server starts. It reads every sealed row. If some rows do not open, it logs an
at-rest key proven, but some stored secrets cannot be decryptedwarning. - The database holds sealed secrets, and the key opens none of them: with
CH_ENV=production, the server refuses to start. Outside production, the server logs anat-rest key mismatcherror and starts anyway. - The OTP or OAuth-state key does not match its recorded test value: the server refuses to start in every environment.
CH_AT_REST_KEY_CANARY_RESEAL=1: the server starts and adopts the current key. Secrets sealed under the old key stay unreadable. When only a test value failed, one boot re-records it. Remove the setting after that boot.- The key opens none of the sealed secrets, and you set
CH_AT_REST_KEY_CANARY_RESEAL=1: the setting does not re-record anything for these rows. Keep the setting until you re-enter at least one secret, or restore the original key. Then remove it. If you remove it too early, the next production boot refuses to start again.
Under CH_ENV=production the server never invents a key. An unset or short CH_OTP_HMAC_KEY or CH_INTEGRATIONS_OAUTH_STATE_KEY stops the start. The release job also never creates a key. It stops with FATAL when a key file is missing on the host.
A nonfatal problem shows on GET /v1/health as the atRestKeys field. Its value is ok, mismatch, unreadable_rows or not_checked. Alert on any value other than ok. The status code stays 200, so the server keeps running while you fix the key.
If the log shows the error, put the original key back and restart.
To replace a key on purpose, re-enter every secret that key protected. Create each webhook secret and each integration connection again. Do this while the server runs on the new key.
Back up the database
Use one of these methods:
- Managed snapshots with point-in-time recovery. On Amazon RDS, turn on automated backups and set a retention period. Seven days is a common start. Recovery can then reach any second inside that period.
pg_dump. Runpg_dump --format=customon a schedule. Copy each dump off the database host.
Keep the backups in your own account. Copy them to a second account or region if your risk rules ask for it. A backup in the same account as the database does not survive the loss of that account.
Take a manual snapshot before every upgrade. CodeHerder applies schema migrations at start, so a snapshot is your way back.
Back up attachments
Attachments live in one S3 bucket, CH_S3_BUCKET. Uploaded files sit under attachments/<workspace_id>/<id>. Embedded images sit under images/<workspace_id>/<id>. Rows in the database point at these objects.
Protect the bucket with one of these:
- S3 versioning, with a lifecycle rule that expires old versions after your retention period.
- Replication to a bucket in another account or region.
- AWS Backup for S3.
Pair the bucket’s restore point with the database’s restore point. A database restored to Monday and a bucket restored to Friday leave links that point at missing or newer objects. Restore both to the same time.
Keep the revocation journal
A point-in-time restore brings back a credential you revoked after the recovery point. It also brings back a person you erased after that point. CodeHerder therefore writes each API key, device token, CLI installation and MCP grant revocation, and each person erasure, to a journal outside the database. A revoke-all writes one entry per credential. A member deprovision writes one entry too, so a restore disables the member’s login and removes the grants again. An entry holds an id, a kind, the id of the credential or person, and a time. A deprovision entry also holds the id of the workspace root it ran in. It holds no token, name, email or reason.
Choose where the journal lives:
- A local directory. Set
CH_REVOCATION_JOURNAL_DIR. The directory must sit off the database host. Back it up on its own schedule. The server adds files and never rewrites or deletes one. - The attachments bucket. Leave
CH_REVOCATION_JOURNAL_DIRunset and setCH_S3_BUCKET. The journal goes under therevocation-journal/prefix. The app host role needss3:PutObject,s3:GetObjectands3:ListBucketon that prefix. It needs no delete grant.
If you set neither, nothing ships and a restored clone cannot be promoted.
Keep the journal at least as long as your longest database backup. Protect it as you protect the bucket: turn on S3 versioning or Object Lock for the prefix. A change in the last ten seconds before a total loss of the database host may not have shipped yet.
The journal does not cover a device or credential row that someone hard-deleted after the recovery point. It does not delete the person’s sign-in identity in Cognito: the user pool is not part of the database restore. If you also rolled the pool back, delete that identity by hand.
Recover
Restore in this order. Each step needs the one before it.
- Fetch the four keys from escrow. Confirm each
sha256sum. - Restore the database from a snapshot or a dump.
- Fence the restored database before any server connects to it. Run
INSERT INTO restore_fences (reason) VALUES ('restore <database id>'). Check thatSELECT count(*) FROM restore_fences WHERE promoted_at IS NULLis at least 1. The fence is not automatic. Without this row, the server sends email and webhooks and accepts devices before the replay. - Restore the attachments bucket to the same point in time as the database.
- Confirm the KMS key is active and the app host role can use it.
- Set the four keys,
--db,CH_S3_BUCKET,CH_KMS_KEY_ID,CH_REVOCATION_JOURNAL_DIRif you use it, and the rest of your configuration. Start the server. Runch instance restore-fence show. It must readfenced: yes. - Read the server log. Look for an
at-rest key mismatcherror. If you see one, stop and go back to step 1. - Run
ch instance restore-fence replay --yes. It re-applies the revocations and erasures made after the recovery point. Runch instance restore-fence showand confirm the journal has no entry left. Then runch instance restore-fence promote --yes. The server refuses to promote until the replay has run. Until the replay has run, a fenced clone answers only/healthand the restore-fence routes. Every other request, including sign-in, gets 503restore_replay_required. Runch instance restore-fence showagain. It must readfenced: no. See Promote a restored database to production. - Sign in. Open a task that has an attachment and confirm the file opens.
Erased data in backups
An erasure deletes rows from the live database. It does not change a backup. Erased rows stay in each snapshot and dump until that backup expires. A restore brings them back. After you restore, run each erasure again. Set a backup retention period that fits your erasure obligations. See Self-hosted log retention and access for the logs. See Erased data in backups for the window of each store.
Keep recovery copies in a second region
Production snapshots stay in your production AWS account. This is the CodeHerder rule. A copy in another region of the same account is allowed. A copy in another account is not.
Pick one method for the database:
- Cross-Region automated backup replication. Run
aws rds start-db-instance-automated-backups-replication --source-db-instance-arn <arn> --kms-key-id <destination-key> --backup-retention-period <days> --region <recovery-region>. You get point-in-time recovery in the second region. - Scheduled snapshot copy. Run
aws rds copy-db-snapshot --source-region <region> --kms-key-id <destination-key>from a schedule. You recover to the snapshot time only. - AWS Backup copy rule. Copy to a vault in the recovery region. The vault needs its own key.
Copy the other dependencies the same way:
- Attachments. Use S3 Cross-Region Replication to a bucket in the same account. Give the replica its own KMS key.
- The four keys. Use the Secrets Manager
replicate-secret-to-regionsoption. Keep the offline copy as well. - The sealed-variables key. Make
CH_KMS_KEY_IDa multi-Region key with a replica in the recovery region. Sealed variables cannot open there without it.
Each copy needs these key grants. Reuse the key policy shape from Self-hosted database.
| Who | Key | Grants |
|---|---|---|
| The RDS service | Destination key | Use through kms:ViaService set to rds.<recovery-region>.amazonaws.com |
| The copy operator role | Source key | kms:DescribeKey, kms:Decrypt |
| The copy operator role | Destination key | kms:CreateGrant, kms:DescribeKey, kms:Encrypt, kms:GenerateDataKey*, kms:ReEncrypt* |
| The S3 replication role | Source bucket key | kms:Decrypt |
| The S3 replication role | Destination bucket key | kms:Encrypt, kms:GenerateDataKey |
| The app host role in the recovery region | Sealed-variables key replica | kms:Decrypt, kms:DescribeKey |
Never schedule deletion of a key that a copy still needs. Your organization approves the regions and the key grants.
Restore dependencies alongside the database
A database restore alone does not give a working server. Restore each item below to the same point in time.
| Item | Where it lives | Back it up | Check after restore |
|---|---|---|---|
| The four keys | Secrets Manager and the offline copy | Escrow the four keys | sha256sum matches. ch instance check passes secret_decryption. |
| Sealed-variables key | The KMS key in CH_KMS_KEY_ID |
Multi-Region replica | The app host role can decrypt with it. |
| Cognito pool settings | Your user pool | Save the output of describe-user-pool, describe-user-pool-client for each app client, and list-identity-providers and describe-identity-provider. Keep it in version control. |
ch instance check passes sign_in. |
| Attachments bucket | CH_S3_BUCKET |
Versioning or replication | An attachment on a restored task opens. |
| Revocation journal | A directory or the bucket prefix | Keep the revocation journal | ch instance restore-fence show lists the pending entries. |
| Application configuration | Environment file and unit files | Keep them in version control | The server boots with the same settings. |
For Cognito, record the pool id, the region, the domain, and the three app client ids with their callback and logout URLs. Record MFA, threat protection and each identity provider. These map to CH_COGNITO_USER_POOL_ID, CH_COGNITO_REGION, CH_COGNITO_DOMAIN, CH_COGNITO_SPA_CLIENT_ID, CH_COGNITO_CLI_CLIENT_ID, CH_COGNITO_MCP_CLIENT_ID and CH_COGNITO_PROVIDERS. See Self-hosted Cognito for the server role permissions.
Never delete the user pool. A pool cannot export passwords. A new pool gives each person a new subject id. A person’s row keeps the subject id of the old pool. The server then refuses that person’s sign-in. So keep the original pool.
For the application configuration, record every CH_* value you set, CH_BASE_URL, the --db value, and the release version. Read the version from /v1/version. Restore with the same release or a newer one.
Watch the recoverable point
The recoverable point is the newest time you can restore to. Its lag is the current time minus that point. RDS has no metric for it. Read LatestRestorableTime from aws rds describe-db-instances instead.
CodeHerder recommends these thresholds. Your organization approves them.
| Signal | Normal | Warn | Page |
|---|---|---|---|
| Database lag | About 5 minutes | Above 15 minutes | Above 30 minutes |
| Second-region replicated backup lag | About 5 minutes | Above 15 minutes | Above 30 minutes |
| Newest copied snapshot age | Under 24 hours | Above 24 hours | Above 26 hours |
Attachments ReplicationLatency |
Under 15 minutes | Above 15 minutes | Any OperationsFailedReplication |
Run this every 5 minutes from a scheduler:
latest=$(aws rds describe-db-instances --db-instance-identifier "$DB_ID" \
--query 'DBInstances[0].LatestRestorableTime' --output text)
lag=$(( $(date +%s) - $(date -d "$latest" +%s) ))
aws cloudwatch put-metric-data --namespace CodeHerder/Recovery \
--metric-name RecoverableLagSeconds --dimensions DBInstance="$DB_ID" \
--value "$lag" --unit Seconds
Then create the alarm. The hosted service’s own RDS alarms use this row shape. Use it for the lag alarm too.
| Name | Metric | Stat | Comparison | Threshold | Periods | Datapoints | Missing data |
|---|---|---|---|---|---|---|---|
| recovery-lag-warn | RecoverableLagSeconds |
Maximum | Greater than | 900 | 3 | 3 | Breaching |
| recovery-lag-page | RecoverableLagSeconds |
Maximum | Greater than | 1800 | 2 | 2 | Breaching |
A missing datapoint pages. A dead check must not look healthy.
To prove the latest point is usable, run the rehearsal below each month. Restore with --use-latest-restorable-time. Record the lag it reached. Keep each evidence record as long as you keep the backups.
Recovery targets
Two numbers define recovery. The recovery-point objective (RPO) is how much recent data you accept to lose. The recovery-time objective (RTO) is how long you accept to be without service. Each target needs a place where you measure it. Without that place, two people read the same number two ways.
Where to measure RPO. Measure from the last recoverable point. RPO is the incident time minus the recovery point you restored. For the database, the recovery point is LatestRestorableTime at the moment you restore. For attachments, the measure is the replication latency of the copy. The four escrowed keys need no loss at all. A lost key makes the data it protects unrecoverable.
Where to measure RTO. Start the clock when the outage starts. That is the first failed customer journey or the first page alarm, whichever is earlier. Stop the clock when a signed-in person completes a task. That is T1 in Rehearse a full restore. The drill’s T1 − T0 is only the restore part. It leaves out detection and the decision to restore. Add both.
These defaults come from the measured drill. Your organization approves them.
| Scenario | RPO | RTO | Basis |
|---|---|---|---|
| Database loss in the primary region | 15 minutes | 4 hours | The measured recovery-point lag is about 5 minutes. The warn threshold is 15 minutes. |
| Loss of the primary region, with replicated automated backups | 15 minutes | 8 hours | The replicated backup lag has the same thresholds. |
| Loss of the primary region, with a scheduled snapshot copy | 26 hours | 8 hours | The newest copied snapshot is at most 26 hours old. |
| Attachments | 15 minutes | With the database | The ReplicationLatency warn threshold. |
| Escrowed keys | Zero | With the database | Escrow them before first boot. |
The in-region RTO of 4 hours has this budget:
| Step | Budget |
|---|---|
| Detect the outage | 30 minutes |
| Decide to restore | 30 minutes |
Database reaches available |
25 minutes (measured: 875 to 1402 seconds) |
| Keys, bucket, fence, replay, promote and verify | 90 minutes |
| Margin | The rest |
The region-loss RTO adds the time to build the application host and the identity provider connection in the second region. Measure it in your first region-loss drill.
Replace a default with your own first full-drill measurement. A target below the measured complete recovery time plus detection and decision time is not valid. Raise the target, or shorten the steps and drill again.
Rule text for approval
This text is a template. Adapt it to your policy.
Recovery target rule.
- The recovery targets are the table above, unless the customer replaces a row with its own measured value.
- RPO is measured from the last recoverable point. RTO is measured from the start of the outage to a signed-in person completing a task.
- The customer runs the full-restore drill each month and records the outage start and the result against the target in the drill record.
- A drill that misses a target opens an action with an owner and a date. The target stays until the action closes or the customer approves a new target.
- The customer approves its targets. CodeHerder does not set them for the customer.
Rehearse a full restore
A backup you have never restored is a guess. Run this drill in the production account on a separate drill host. A drill never serves production traffic.
The fence blocks email, invitations, webhooks, device connections and agent spawns on the clone. It does not block Cognito admin calls, S3 writes or git-host writes. So the drill uses its own bucket and its own base URL.
- Record the retention and
LatestRestorableTimefromaws rds describe-db-instances. - Restore with
aws rds restore-db-instance-to-point-in-time. Set all six flags:--use-latest-restorable-time,--db-subnet-group-name,--vpc-security-group-ids,--no-publicly-accessible,--db-parameter-group-nameand--no-deletion-protection. Name the clone with adrill-prefix. Start the clock (T0) when the call returns. - When the clone is
available, fence it before any server connects:INSERT INTO restore_fences (reason) VALUES ('drill <clone id>'). Check thatSELECT count(*) FROM restore_fences WHERE promoted_at IS NULLis at least 1. - Restore the attachments to a drill bucket at the same point in time. Do not use the production bucket.
- Start a drill server from the same release. Use the escrowed keys,
--dbset to the clone,CH_S3_BUCKETset to the drill bucket, and a drillCH_BASE_URL. Copy the production revocation journal into a drill directory, and setCH_REVOCATION_JOURNAL_DIRto that directory. If production keeps the journal in its bucket, copy therevocation-journal/prefix of that bucket into the drill directory withaws s3 sync. Replay only reads the journal, and a fenced clone never ships to it, so the copy stays unchanged. Add the drill sign-in callback URL to the SPA app client. - Run
ch instance restore-fence show. It must readfenced: yes. The log must show noat-rest key mismatch. - Run
ch instance restore-fence replay --yes. Until it runs, the clone answers only/healthand the restore-fence routes, so sign-in fails. Runch instance restore-fence showand confirm the journal has no entry left. The clone holds real customer data, so this keeps every later revocation in force on it. - After replay and before sign-in, confirm that a key revoked after the recovery point gets 401, and that
ch instance restore-fence showstill readsfenced: yes. Runch instance check --skip mail --skip device_connection. It must pass. - A person signs in through the web app. They open a task that has an attachment and confirm the file opens. They move a task to
done. Stop the clock (T1) when the task readsdone.
Do not do these on a drill clone: promote it, point its journal at the production journal store itself (use the copy), deprovision or erase a person, merge a request, or enrol a device. Cognito calls are not fenced.
Record the result.
| Field | Value |
|---|---|
| Date | |
| Outage start (a real outage, or the drill’s simulated start) | |
| Recovery point | |
| Recovery-point lag | |
Time to available |
|
| Complete recovery time (T1 − T0) | |
| Time from outage start to T1 | |
| Met the RPO and RTO targets (yes or no) | |
| Release | |
| Person who ran it |
Clean up a drill
This procedure deletes only the drill clone, the drill bucket and the drill host. Never touch the source database.
- Read the identifier back from
describe-db-instances. Do not use a variable you did not read back. - Confirm it starts with
drill-. - Confirm
restore_fenceson that clone holds a row withpromoted_at IS NULL. - Ask a second person to confirm the identifier.
- Run
aws rds delete-db-instance --skip-final-snapshoton that identifier. Then delete the drill bucket and the drill host.
Promote a restored database to production
This procedure replaces production. It is separate from the drill cleanup. Never run the drill cleanup on a promoted target.
- Restore the database. Follow Recover for the order of keys, bucket and configuration.
- Fence the restored database before any server connects to it. Run
INSERT INTO restore_fences (reason) VALUES ('restore <database id>'). Check thatSELECT count(*) FROM restore_fences WHERE promoted_at IS NULLis at least 1. - Start the server. Run
ch instance restore-fence show. It must readfenced: yes. Verify the data. - Run
ch instance restore-fence replay --yes. Runch instance restore-fence showand confirm no entry is pending. - Run
ch instance restore-fence promote --yes. Runch instance restore-fence show. It must readfenced: no. - Turn on deletion protection on the recovered target:
aws rds modify-db-instance --deletion-protection. This step is required. A restore made with the drill flags has--no-deletion-protection. - Only now point DNS and configuration at the recovered database.
Never delete the recovered target. Never delete the old primary until the new one has run a full backup cycle and an owner signs off. One person runs the procedure. A second person confirms the identifier before any delete.
Related guides
- Self-hosted deployment — download, start, and upgrade the server
- Sizing a self-hosted deployment — app host, database, and device fleet sizes
- Self-hosted logs — where to find the server log that reports a key mismatch
- Self-hosted key custody — custody record, production key settings, KMS key recovery
- Secrets and variables — the values that use the KMS key
- Attachments — what the attachments bucket holds
Last updated