CodeHerderSearch⌘KRequest access →

Self-hosted backup and recovery

Escrow the four at-rest keys, back up and copy the database and attachments, watch the recoverable point, rehearse a restore, and promote it safely.

A self-hosted CodeHerder server keeps state in three places: the PostgreSQL database, four at-rest keys, and an attachments bucket. A backup that misses one of them cannot restore a working server. Your organization owns the backups, their retention, and their access control. CodeHerder does not run them for you.

What to back up

What Where it lives What you lose without it
PostgreSQL database The server named in --db Every task, session, member, and setting
The four at-rest keys Environment variables on the app host Sealed webhook and integration secrets, and sign-in codes and OAuth flows that are in flight
Attachments bucket The S3 bucket named in CH_S3_BUCKET Every uploaded file and embedded image
KMS key for sealed variables The AWS KMS key named in CH_KMS_KEY_ID Every value stored with ch variable as a secret
Server configuration Your environment file and unit files The settings you chose, such as identity mode and pool sizes

The four keys are environment variables. They are not files. Each key protects a different thing:

  • CH_WEBHOOK_SECRET_KEY seals the secrets of outbound webhook subscriptions and inbound webhook endpoints.
  • CH_INTEGRATIONS_SECRET_KEY seals the secrets of integration connections.
  • CH_OTP_HMAC_KEY signs sign-in codes. If you lose it, codes that are outstanding stop working. Nothing stored is lost.
  • CH_INTEGRATIONS_OAUTH_STATE_KEY signs OAuth state. If you lose it, OAuth flows that are in flight fail. Nothing stored is lost.

The first two keys are the ones that cannot be replaced. Treat all four the same way, so that a restore never has to guess.

The KMS key is a fifth dependency. It is not one of the four keys. A restore needs that KMS key to exist, and the app host role must be able to use it. Do not schedule the KMS key for deletion while any backup still needs it.

Escrow the four keys before first boot

Escrow means you keep a safe copy of each key somewhere other than the app host. Do this before the server starts for the first time. After the server seals its first secret, a lost key cannot be rebuilt.

  1. Generate the four keys. Use openssl rand -hex 32 for each one, as in Start the server.
  2. Store all four in a secret manager outside the app host, for example AWS Secrets Manager or a vault. Store one more copy offline, under the control of two people.
  3. Read each key back from the secret manager. Compare its sha256sum with the value you generated. Do the same for the offline copy.
  4. Only now export the keys on the app host and start the server.
  5. Write down where each copy lives in your operator runbook. Name the people who can open the offline copy.

Give the app host role read access to the secret manager entries only. Do not let the same role delete them.

Never mint a new key on an existing database

Reuse the escrowed values on every new host, every rebuild, and every restore. Never run openssl rand again for a database that already holds data.

The server checks all four keys at boot: CH_WEBHOOK_SECRET_KEY, CH_INTEGRATIONS_SECRET_KEY, CH_OTP_HMAC_KEY and CH_INTEGRATIONS_OAUTH_STATE_KEY. This is what the check does:

  • Fresh database: the server records a small test value for each key. The next boot checks it.
  • The database holds sealed secrets, and the key opens at least one: the server starts. It reads every sealed row. If some rows do not open, it logs an at-rest key proven, but some stored secrets cannot be decrypted warning.
  • The database holds sealed secrets, and the key opens none of them: with CH_ENV=production, the server refuses to start. Outside production, the server logs an at-rest key mismatch error and starts anyway.
  • The OTP or OAuth-state key does not match its recorded test value: the server refuses to start in every environment.
  • CH_AT_REST_KEY_CANARY_RESEAL=1: the server starts and adopts the current key. Secrets sealed under the old key stay unreadable. When only a test value failed, one boot re-records it. Remove the setting after that boot.
  • The key opens none of the sealed secrets, and you set CH_AT_REST_KEY_CANARY_RESEAL=1: the setting does not re-record anything for these rows. Keep the setting until you re-enter at least one secret, or restore the original key. Then remove it. If you remove it too early, the next production boot refuses to start again.

Under CH_ENV=production the server never invents a key. An unset or short CH_OTP_HMAC_KEY or CH_INTEGRATIONS_OAUTH_STATE_KEY stops the start. The release job also never creates a key. It stops with FATAL when a key file is missing on the host.

A nonfatal problem shows on GET /v1/health as the atRestKeys field. Its value is ok, mismatch, unreadable_rows or not_checked. Alert on any value other than ok. The status code stays 200, so the server keeps running while you fix the key.

If the log shows the error, put the original key back and restart.

To replace a key on purpose, re-enter every secret that key protected. Create each webhook secret and each integration connection again. Do this while the server runs on the new key.

Back up the database

Use one of these methods:

  • Managed snapshots with point-in-time recovery. On Amazon RDS, turn on automated backups and set a retention period. Seven days is a common start. Recovery can then reach any second inside that period.
  • pg_dump. Run pg_dump --format=custom on a schedule. Copy each dump off the database host.

Keep the backups in your own account. Copy them to a second account or region if your risk rules ask for it. A backup in the same account as the database does not survive the loss of that account.

Take a manual snapshot before every upgrade. CodeHerder applies schema migrations at start, so a snapshot is your way back.

Back up attachments

Attachments live in one S3 bucket, CH_S3_BUCKET. Uploaded files sit under attachments/<workspace_id>/<id>. Embedded images sit under images/<workspace_id>/<id>. Rows in the database point at these objects.

Protect the bucket with one of these:

  • S3 versioning, with a lifecycle rule that expires old versions after your retention period.
  • Replication to a bucket in another account or region.
  • AWS Backup for S3.

Pair the bucket’s restore point with the database’s restore point. A database restored to Monday and a bucket restored to Friday leave links that point at missing or newer objects. Restore both to the same time.

Keep the revocation journal

A point-in-time restore brings back a credential you revoked after the recovery point. It also brings back a person you erased after that point. CodeHerder therefore writes each API key, device token, CLI installation and MCP grant revocation, and each person erasure, to a journal outside the database. A revoke-all writes one entry per credential. A member deprovision writes one entry too, so a restore disables the member’s login and removes the grants again. An entry holds an id, a kind, the id of the credential or person, and a time. A deprovision entry also holds the id of the workspace root it ran in. It holds no token, name, email or reason.

Choose where the journal lives:

  • A local directory. Set CH_REVOCATION_JOURNAL_DIR. The directory must sit off the database host. Back it up on its own schedule. The server adds files and never rewrites or deletes one.
  • The attachments bucket. Leave CH_REVOCATION_JOURNAL_DIR unset and set CH_S3_BUCKET. The journal goes under the revocation-journal/ prefix. The app host role needs s3:PutObject, s3:GetObject and s3:ListBucket on that prefix. It needs no delete grant.

If you set neither, nothing ships and a restored clone cannot be promoted.

Keep the journal at least as long as your longest database backup. Protect it as you protect the bucket: turn on S3 versioning or Object Lock for the prefix. A change in the last ten seconds before a total loss of the database host may not have shipped yet.

The journal does not cover a device or credential row that someone hard-deleted after the recovery point. It does not delete the person’s sign-in identity in Cognito: the user pool is not part of the database restore. If you also rolled the pool back, delete that identity by hand.

Recover

Restore in this order. Each step needs the one before it.

  1. Fetch the four keys from escrow. Confirm each sha256sum.
  2. Restore the database from a snapshot or a dump.
  3. Fence the restored database before any server connects to it. Run INSERT INTO restore_fences (reason) VALUES ('restore <database id>'). Check that SELECT count(*) FROM restore_fences WHERE promoted_at IS NULL is at least 1. The fence is not automatic. Without this row, the server sends email and webhooks and accepts devices before the replay.
  4. Restore the attachments bucket to the same point in time as the database.
  5. Confirm the KMS key is active and the app host role can use it.
  6. Set the four keys, --db, CH_S3_BUCKET, CH_KMS_KEY_ID, CH_REVOCATION_JOURNAL_DIR if you use it, and the rest of your configuration. Start the server. Run ch instance restore-fence show. It must read fenced: yes.
  7. Read the server log. Look for an at-rest key mismatch error. If you see one, stop and go back to step 1.
  8. Run ch instance restore-fence replay --yes. It re-applies the revocations and erasures made after the recovery point. Run ch instance restore-fence show and confirm the journal has no entry left. Then run ch instance restore-fence promote --yes. The server refuses to promote until the replay has run. Until the replay has run, a fenced clone answers only /health and the restore-fence routes. Every other request, including sign-in, gets 503 restore_replay_required. Run ch instance restore-fence show again. It must read fenced: no. See Promote a restored database to production.
  9. Sign in. Open a task that has an attachment and confirm the file opens.

Erased data in backups

An erasure deletes rows from the live database. It does not change a backup. Erased rows stay in each snapshot and dump until that backup expires. A restore brings them back. After you restore, run each erasure again. Set a backup retention period that fits your erasure obligations. See Self-hosted log retention and access for the logs. See Erased data in backups for the window of each store.

Keep recovery copies in a second region

Production snapshots stay in your production AWS account. This is the CodeHerder rule. A copy in another region of the same account is allowed. A copy in another account is not.

Pick one method for the database:

  • Cross-Region automated backup replication. Run aws rds start-db-instance-automated-backups-replication --source-db-instance-arn <arn> --kms-key-id <destination-key> --backup-retention-period <days> --region <recovery-region>. You get point-in-time recovery in the second region.
  • Scheduled snapshot copy. Run aws rds copy-db-snapshot --source-region <region> --kms-key-id <destination-key> from a schedule. You recover to the snapshot time only.
  • AWS Backup copy rule. Copy to a vault in the recovery region. The vault needs its own key.

Copy the other dependencies the same way:

  • Attachments. Use S3 Cross-Region Replication to a bucket in the same account. Give the replica its own KMS key.
  • The four keys. Use the Secrets Manager replicate-secret-to-regions option. Keep the offline copy as well.
  • The sealed-variables key. Make CH_KMS_KEY_ID a multi-Region key with a replica in the recovery region. Sealed variables cannot open there without it.

Each copy needs these key grants. Reuse the key policy shape from Self-hosted database.

Who Key Grants
The RDS service Destination key Use through kms:ViaService set to rds.<recovery-region>.amazonaws.com
The copy operator role Source key kms:DescribeKey, kms:Decrypt
The copy operator role Destination key kms:CreateGrant, kms:DescribeKey, kms:Encrypt, kms:GenerateDataKey*, kms:ReEncrypt*
The S3 replication role Source bucket key kms:Decrypt
The S3 replication role Destination bucket key kms:Encrypt, kms:GenerateDataKey
The app host role in the recovery region Sealed-variables key replica kms:Decrypt, kms:DescribeKey

Never schedule deletion of a key that a copy still needs. Your organization approves the regions and the key grants.

Restore dependencies alongside the database

A database restore alone does not give a working server. Restore each item below to the same point in time.

Item Where it lives Back it up Check after restore
The four keys Secrets Manager and the offline copy Escrow the four keys sha256sum matches. ch instance check passes secret_decryption.
Sealed-variables key The KMS key in CH_KMS_KEY_ID Multi-Region replica The app host role can decrypt with it.
Cognito pool settings Your user pool Save the output of describe-user-pool, describe-user-pool-client for each app client, and list-identity-providers and describe-identity-provider. Keep it in version control. ch instance check passes sign_in.
Attachments bucket CH_S3_BUCKET Versioning or replication An attachment on a restored task opens.
Revocation journal A directory or the bucket prefix Keep the revocation journal ch instance restore-fence show lists the pending entries.
Application configuration Environment file and unit files Keep them in version control The server boots with the same settings.

For Cognito, record the pool id, the region, the domain, and the three app client ids with their callback and logout URLs. Record MFA, threat protection and each identity provider. These map to CH_COGNITO_USER_POOL_ID, CH_COGNITO_REGION, CH_COGNITO_DOMAIN, CH_COGNITO_SPA_CLIENT_ID, CH_COGNITO_CLI_CLIENT_ID, CH_COGNITO_MCP_CLIENT_ID and CH_COGNITO_PROVIDERS. See Self-hosted Cognito for the server role permissions.

Never delete the user pool. A pool cannot export passwords. A new pool gives each person a new subject id. A person’s row keeps the subject id of the old pool. The server then refuses that person’s sign-in. So keep the original pool.

For the application configuration, record every CH_* value you set, CH_BASE_URL, the --db value, and the release version. Read the version from /v1/version. Restore with the same release or a newer one.

Watch the recoverable point

The recoverable point is the newest time you can restore to. Its lag is the current time minus that point. RDS has no metric for it. Read LatestRestorableTime from aws rds describe-db-instances instead.

CodeHerder recommends these thresholds. Your organization approves them.

Signal Normal Warn Page
Database lag About 5 minutes Above 15 minutes Above 30 minutes
Second-region replicated backup lag About 5 minutes Above 15 minutes Above 30 minutes
Newest copied snapshot age Under 24 hours Above 24 hours Above 26 hours
Attachments ReplicationLatency Under 15 minutes Above 15 minutes Any OperationsFailedReplication

Run this every 5 minutes from a scheduler:

latest=$(aws rds describe-db-instances --db-instance-identifier "$DB_ID" \
  --query 'DBInstances[0].LatestRestorableTime' --output text)
lag=$(( $(date +%s) - $(date -d "$latest" +%s) ))
aws cloudwatch put-metric-data --namespace CodeHerder/Recovery \
  --metric-name RecoverableLagSeconds --dimensions DBInstance="$DB_ID" \
  --value "$lag" --unit Seconds

Then create the alarm. The hosted service’s own RDS alarms use this row shape. Use it for the lag alarm too.

Name Metric Stat Comparison Threshold Periods Datapoints Missing data
recovery-lag-warn RecoverableLagSeconds Maximum Greater than 900 3 3 Breaching
recovery-lag-page RecoverableLagSeconds Maximum Greater than 1800 2 2 Breaching

A missing datapoint pages. A dead check must not look healthy.

To prove the latest point is usable, run the rehearsal below each month. Restore with --use-latest-restorable-time. Record the lag it reached. Keep each evidence record as long as you keep the backups.

Recovery targets

Two numbers define recovery. The recovery-point objective (RPO) is how much recent data you accept to lose. The recovery-time objective (RTO) is how long you accept to be without service. Each target needs a place where you measure it. Without that place, two people read the same number two ways.

Where to measure RPO. Measure from the last recoverable point. RPO is the incident time minus the recovery point you restored. For the database, the recovery point is LatestRestorableTime at the moment you restore. For attachments, the measure is the replication latency of the copy. The four escrowed keys need no loss at all. A lost key makes the data it protects unrecoverable.

Where to measure RTO. Start the clock when the outage starts. That is the first failed customer journey or the first page alarm, whichever is earlier. Stop the clock when a signed-in person completes a task. That is T1 in Rehearse a full restore. The drill’s T1 − T0 is only the restore part. It leaves out detection and the decision to restore. Add both.

These defaults come from the measured drill. Your organization approves them.

Scenario RPO RTO Basis
Database loss in the primary region 15 minutes 4 hours The measured recovery-point lag is about 5 minutes. The warn threshold is 15 minutes.
Loss of the primary region, with replicated automated backups 15 minutes 8 hours The replicated backup lag has the same thresholds.
Loss of the primary region, with a scheduled snapshot copy 26 hours 8 hours The newest copied snapshot is at most 26 hours old.
Attachments 15 minutes With the database The ReplicationLatency warn threshold.
Escrowed keys Zero With the database Escrow them before first boot.

The in-region RTO of 4 hours has this budget:

Step Budget
Detect the outage 30 minutes
Decide to restore 30 minutes
Database reaches available 25 minutes (measured: 875 to 1402 seconds)
Keys, bucket, fence, replay, promote and verify 90 minutes
Margin The rest

The region-loss RTO adds the time to build the application host and the identity provider connection in the second region. Measure it in your first region-loss drill.

Replace a default with your own first full-drill measurement. A target below the measured complete recovery time plus detection and decision time is not valid. Raise the target, or shorten the steps and drill again.

Rule text for approval

This text is a template. Adapt it to your policy.

Recovery target rule.

  1. The recovery targets are the table above, unless the customer replaces a row with its own measured value.
  2. RPO is measured from the last recoverable point. RTO is measured from the start of the outage to a signed-in person completing a task.
  3. The customer runs the full-restore drill each month and records the outage start and the result against the target in the drill record.
  4. A drill that misses a target opens an action with an owner and a date. The target stays until the action closes or the customer approves a new target.
  5. The customer approves its targets. CodeHerder does not set them for the customer.

Rehearse a full restore

A backup you have never restored is a guess. Run this drill in the production account on a separate drill host. A drill never serves production traffic.

The fence blocks email, invitations, webhooks, device connections and agent spawns on the clone. It does not block Cognito admin calls, S3 writes or git-host writes. So the drill uses its own bucket and its own base URL.

  1. Record the retention and LatestRestorableTime from aws rds describe-db-instances.
  2. Restore with aws rds restore-db-instance-to-point-in-time. Set all six flags: --use-latest-restorable-time, --db-subnet-group-name, --vpc-security-group-ids, --no-publicly-accessible, --db-parameter-group-name and --no-deletion-protection. Name the clone with a drill- prefix. Start the clock (T0) when the call returns.
  3. When the clone is available, fence it before any server connects: INSERT INTO restore_fences (reason) VALUES ('drill <clone id>'). Check that SELECT count(*) FROM restore_fences WHERE promoted_at IS NULL is at least 1.
  4. Restore the attachments to a drill bucket at the same point in time. Do not use the production bucket.
  5. Start a drill server from the same release. Use the escrowed keys, --db set to the clone, CH_S3_BUCKET set to the drill bucket, and a drill CH_BASE_URL. Copy the production revocation journal into a drill directory, and set CH_REVOCATION_JOURNAL_DIR to that directory. If production keeps the journal in its bucket, copy the revocation-journal/ prefix of that bucket into the drill directory with aws s3 sync. Replay only reads the journal, and a fenced clone never ships to it, so the copy stays unchanged. Add the drill sign-in callback URL to the SPA app client.
  6. Run ch instance restore-fence show. It must read fenced: yes. The log must show no at-rest key mismatch.
  7. Run ch instance restore-fence replay --yes. Until it runs, the clone answers only /health and the restore-fence routes, so sign-in fails. Run ch instance restore-fence show and confirm the journal has no entry left. The clone holds real customer data, so this keeps every later revocation in force on it.
  8. After replay and before sign-in, confirm that a key revoked after the recovery point gets 401, and that ch instance restore-fence show still reads fenced: yes. Run ch instance check --skip mail --skip device_connection. It must pass.
  9. A person signs in through the web app. They open a task that has an attachment and confirm the file opens. They move a task to done. Stop the clock (T1) when the task reads done.

Do not do these on a drill clone: promote it, point its journal at the production journal store itself (use the copy), deprovision or erase a person, merge a request, or enrol a device. Cognito calls are not fenced.

Record the result.

Field Value
Date
Outage start (a real outage, or the drill’s simulated start)
Recovery point
Recovery-point lag
Time to available
Complete recovery time (T1 − T0)
Time from outage start to T1
Met the RPO and RTO targets (yes or no)
Release
Person who ran it

Clean up a drill

This procedure deletes only the drill clone, the drill bucket and the drill host. Never touch the source database.

  1. Read the identifier back from describe-db-instances. Do not use a variable you did not read back.
  2. Confirm it starts with drill-.
  3. Confirm restore_fences on that clone holds a row with promoted_at IS NULL.
  4. Ask a second person to confirm the identifier.
  5. Run aws rds delete-db-instance --skip-final-snapshot on that identifier. Then delete the drill bucket and the drill host.

Promote a restored database to production

This procedure replaces production. It is separate from the drill cleanup. Never run the drill cleanup on a promoted target.

  1. Restore the database. Follow Recover for the order of keys, bucket and configuration.
  2. Fence the restored database before any server connects to it. Run INSERT INTO restore_fences (reason) VALUES ('restore <database id>'). Check that SELECT count(*) FROM restore_fences WHERE promoted_at IS NULL is at least 1.
  3. Start the server. Run ch instance restore-fence show. It must read fenced: yes. Verify the data.
  4. Run ch instance restore-fence replay --yes. Run ch instance restore-fence show and confirm no entry is pending.
  5. Run ch instance restore-fence promote --yes. Run ch instance restore-fence show. It must read fenced: no.
  6. Turn on deletion protection on the recovered target: aws rds modify-db-instance --deletion-protection. This step is required. A restore made with the drill flags has --no-deletion-protection.
  7. Only now point DNS and configuration at the recovered database.

Never delete the recovered target. Never delete the old primary until the new one has run a full backup cycle and an owner signs off. One person runs the procedure. A second person confirms the identifier before any delete.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close