CodeHerderSearch⌘KRequest access →

Self-hosted device disk growth

Measure device disk use, see which directories grow and which setting bounds each, set a disk alert, and clean the stores that have no bound.

This page tells you how to watch and bound the disk of a device that runs agent sessions. It covers devices in process mode and in Docker mode. Read it with Sizing a self-hosted deployment.

Set a disk alert

Alert on the percent of disk in use. Do not rely on the Disk headroom readiness check. That check warns only when less than 512 MiB is free, and it blocks work at 128 MiB. That is too late to act.

Level Disk in use Action
Warn 70 percent Find the largest store with the commands below. Clean it the same week.
Act 85 percent Clean now. Lower Max concurrent sessions if a clean-up does not free enough.

Each device reports diskPct for the volume that holds its session root. Read it in either place:

ch device metrics <device>        # min, average, max and current DISK % over the last 30 minutes
ch device show <device> --json    # the current diskPct field

CodeHerder does not alert on diskPct. Poll one of these commands from your own monitoring and raise the two levels above.

Why 70 and 85. The stores that CodeHerder bounds settle at a known size. Sessions in use add one checkout of the repo each, up to 8 by default. The build cache adds up to 20 GiB (CH_CACHE_BUDGET_MB). The stores that CodeHerder does not bound grew by about 210 MB a day on a busy host that we sampled. On the 50 GB default disk, 70 percent leaves 15 GB before the Act level, which is about 70 days of that growth. The Act level leaves 7.5 GB, about 35 days. A burst of trial runs can use more. Replace the levels with your own once you have a month of diskPct history.

Measure what is using the disk

Run these on the device. <root> is the session root, ~/codeherder-sessions unless you set --session-root. In Docker mode the session root is a Docker named volume. Find its path with docker volume inspect.

du -sh <root>/.mirrors <root>/.pool <root>/.transcripts <root>/.reclaimed
du -sh <root>/*/ | sort -h | tail        # live and stranded sandboxes
du -sh ~/.cache ~/go ~/.npm              # build cache
du -sh ~/.claude/projects                # host harness transcripts
du -sh ~/.codeherder/stage-runs          # Docker mode only
docker system df                         # Docker mode only: images, containers, volumes

Which directories grow

Store Where Bound Setting
Live sandboxes <root>/<sandboxId>/ One checkout of the repo each; the device removes a sandbox when its session ends --max-session-worktrees or CH_MAX_SESSION_WORKTREES (8); --session-disk-budget-mb (off)
Git mirrors <root>/.mirrors/ One bare clone for each repo, plus per-task branches. The device removes a delivered branch 14 days after its last commit and then runs git gc once a day. It never removes a branch that is not on your git host. CH_MIRROR_BRANCH_RETENTION_DAYS (14; 0 turns the clean-up off)
Warm pool <root>/.pool/ A few spare checkouts for each repo CH_WORKTREE_POOL_PER_REPO, CH_WORKTREE_POOL_MAX
Transcript archives <root>/.transcripts/ The device removes an archive 7 days after the sandbox ends. One task keeps at most 256 MiB. CH_TRANSCRIPT_ARCHIVE_RETENTION_DAYS (7; 0 keeps archives)
Parked bundles <root>/.reclaimed/ Work that could not be pushed is saved as a bundle for 30 days. Each trial run saves one, and it holds the full repo history. No setting. Clean by hand, see below.
Build cache ~/.cache, ~/go, ~/.npm The device trims it over the budget CH_CACHE_BUDGET_MB (20480)
Device logs the device log directory Each file is cut to the budget, with one old copy CH_LOCAL_LOG_BUDGET_MB (8)
Docker containers, dangling images, volumes the Docker daemon The device removes an exited container after 4 hours CH_DOCKER_REAP_MAX_AGE_HOURS (4)
Docker stage images the Docker daemon None Clean by hand, see below.
Docker stage-run files ~/.codeherder/stage-runs/<runId>/ None for the transcript directory Clean by hand, see below.
Host harness transcripts ~/.claude/projects/ None Clean by hand, see below.

The measured sizes and the growth for each store are in the engineering record docs/device-disk-growth.md of the CodeHerder source. In short, a session costs one checkout while it runs and about 0.5 MB of mirror after it ends. A trial run costs about two checkouts while it runs, and a parked bundle of the repo size for 30 days.

Clean the stores that have no bound

Lower the work on the device before you delete anything. Run ch device edit <deviceId> --max-sessions 1, and wait for running sessions to end.

  1. Parked bundles. A bundle is the only copy of work that could not be pushed. List them, and open any you may need with git clone <bundle>. Delete a bundle you no longer need. If trial runs are the cause, delete the bundles older than 7 days:
    find <root>/.reclaimed -type f -mtime +7 -print     # check the list first
    find <root>/.reclaimed -type f -mtime +7 -delete
  2. Host harness transcripts. Delete files older than the period you keep for resume and audit:
    find ~/.claude/projects -name '*.jsonl' -mtime +14 -delete
  3. Docker stage-run files. Delete directories for runs that ended: find ~/.codeherder/stage-runs -mindepth 1 -maxdepth 1 -mtime +1 -exec rm -rf {} +.
  4. Docker stage images. Remove images no stage uses: docker image prune -a --filter "until=720h". The device pulls an image again when it needs it.
  5. Stranded sandboxes. A directory under <root>/ for a session that no longer runs is removed by the device on its next clean-up cycle. If one stays, follow “Reclaiming stuck worktree slots” in Managing your devices.

Run ch device metrics <device> again. Restore Max concurrent sessions with ch device edit.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close