# Running a stage in your own container image

Source: https://codeherder.com/docs/stage-images/

Give a workflow stage its own container image and a setup script to prepare it, for projects that need a toolchain the default image doesn't have.

If your project needs PHP, a specific Python version, or a system library the default stage container doesn’t carry, you don’t have to work around it — a workflow stage can name its own container image, plus a setup script that prepares that image before the agent starts. Both are per-stage content settings — see *Setting an image and a setup script* below for exactly where you set them.

## Where each setting takes effect

`image` and `before_script` do NOT share one execution mode — read this before you set either one.

- **`image`** takes effect on one execution mode only: a device started with `--docker-executor` (or `CH_DOCKER_EXECUTOR=1`), which runs every stage in its own fresh container. See **Option 2: a fresh container per stage** in [Isolating agent runs on a device](https://codeherder.com/docs/agent-isolation/) for what that mode is and how to turn it on. On a device running stages as host processes — the default — a custom image is simply ignored. Nothing warns you.
- **`before_script`** depends on the mode: On the host path there’s no per-stage container around the script — it runs as its own process wherever the device server itself runs, under the same OS identity the agent itself runs as (the account that started `ch device-server`, or the dedicated agent-uid account from [Option 1](https://codeherder.com/docs/agent-isolation/) when that’s configured). Treat a host `before_script` like a command you’d type on that machine yourself: it can read and write anything that identity can, with none of the container hardening described under **Isolation with a custom image** below.
  - `--docker-executor` (a container per stage): runs by default, inside the stage container.
  - `--docker` (the device server in its own container, as on every Launch-on-AWS device): runs by default, inside that container, not on the machine.
  - Direct mode (the device server runs on the machine with no container): refused unless the operator sets `CH_HOST_BEFORE_SCRIPT=1` on the device, or a workspace admin turns it on in the device’s **Settings** section (see [Changing a device’s settings](https://codeherder.com/docs/device-settings/)). The stage setup event shows outcome `unsupported`.

If a stage’s custom image doesn’t seem to be picking up, check which mode the device running your work is in first. If a stage’s setup script doesn’t seem to be running, check for direct mode first, then see **The setup script** below for where else to look, including the host execution boundary and identity it runs under.

## Setting an image and a setup script

Both settings live alongside **Writable** and the other stage-content settings described in [Customising workflows](https://codeherder.com/docs/task-types/#per-stage-settings) — and, like those, where you set them depends on how the type’s pipeline is built. Neither appears in the workflow editor itself either way.

**For a type composed from the shared stage library** — which every built-in type is — both settings live on the library stage itself, not on the type. Set them with `ch stage edit <key>` (or **Customize** / **Edit** on the stage’s page under **Settings → Stages**) — see [The stage library](https://codeherder.com/docs/stage-library/) for the full round-trip. A change to the library stage carries to every workflow instance that references it.

**For a type written out by hand**, they’re two more fields on the stage’s own schema. You set them through the same CLI round-trip that page documents for any other schema change: read the type’s schema with `ch workspace workflow show --json`, edit it, then write it back with `ch workspace workflow edit`. Add `image` and `before_script` to the stage you want to change:

```
{
  "stage_specs": {
    "code": {
      "writable": true,
      "image": "php:8.3-cli",
      "before_script": "apt-get update && apt-get install -y git unzip"
    }
  }
}
```

Setting either one is workspace owner or admin only, the same as any other workflow edit.

`image` is a standard image reference — a name, optionally with a registry, a tag, or a digest. It’s checked when you save: it has to look like a real image reference rather than something that could be mistaken for a command-line option, so a value like `-privileged` is refused with an error naming the stage. `before_script` is a shell script, run as-is; keep it to what your image needs prepared. The example above shows where the two settings go, not a complete image — it still needs the coding-agent CLI installed before it’s usable; see the next section.

`ch workspace workflow show` ’s plain-text output doesn’t print either setting — add `--json` to see them on a type that has them configured.

Remember that `edit` replaces a type’s whole schema, not just the stage you’re touching — include every other stage, gate, and field you want to keep. See **Manage workflows from the CLI** in [Customising workflows](https://codeherder.com/docs/task-types/#manage-workflows-from-the-cli) for the full round-trip, including how to move tasks off a stage you rename and what `--migrate` does.

## What your image needs to provide

A custom image is otherwise on its own: CodeHerder doesn’t install anything into it beyond what’s described below. At minimum it needs a POSIX shell, git, and whichever coding-agent CLI the agent’s launch config runs — see [Choosing the coding-agent CLI your agents run](https://codeherder.com/docs/harnesses/) for what that CLI is per harness. If your base image is missing any of these, use `before_script` to install them before the agent starts.

CodeHerder still adds a thin layer on top of your image so the agent can operate normally:

- Its own `ch` CLI, so the agent can talk to CodeHerder from inside your container without you needing to install it.
- A writable directory for the coding agent’s own configuration and session state.
- Git configuration so commits and pushes work the same way they do in the default image.

The task’s worktree is mounted into the container and is the working directory the agent starts in, exactly as with the default image.

## The setup script

`before_script` runs before the agent starts — think of it as a place to run an install command, prime a cache, or export something onto `PATH`. Where it runs, and what it exports reaches the agent, differ by execution mode:

- **In a stage container** (`--docker-executor`, with or without a custom `image`): the script runs as the container’s entrypoint, and the agent runs after it via `exec` in the same shell — so everything it exports (a toolchain’s `PATH`, a variable it set) is inherited by the agent unfiltered.
- **On the device host** (the default, no Docker executor): under `--docker`, or in direct mode when the operator sets `CH_HOST_BEFORE_SCRIPT`, the script runs as a separate process under the same OS identity the agent itself runs as (see **Where each setting takes effect** above), and only its exported env DELTA is merged into the agent’s environment afterward — not inherited via a shared shell. That delta is filtered: a change to `PATH`, `HOME`, `LD_PRELOAD`, `NODE_OPTIONS`, `PYTHONPATH`, a provider base-URL (`*_BASE_URL`, `*_API_BASE`), or a git credential/config variable (`GIT_SSH_COMMAND`, `GIT_ASKPASS`, `GIT_CONFIG_*`, and the rest of that family) is dropped, so a setup script can’t rewrite how the agent finds binaries or reach a different git remote or AI provider. CodeHerder’s own read tokens are dropped on every stage; its push credentials are dropped too on a non-writable stage, where the agent isn’t meant to hold them. Any other variable it exports does reach the agent — `GIT_AUTHOR_NAME` and `GIT_DIR`, for instance, are not on the filtered list. In direct mode without `CH_HOST_BEFORE_SCRIPT`, the device refuses the script, and the stage setup event shows outcome `unsupported`.

The two paths are bounded differently. A host-path script is capped by `CH_STAGE_BEFORE_SCRIPT_TIMEOUT` (default 5 minutes) and killed by process group (SIGTERM, then SIGKILL) once it’s past. A script running as a stage container’s entrypoint has no such budget — nothing caps it, so keep a container `before_script` to work that finishes on its own.

Both paths are best-effort: if a step in the script fails — or, on the host path, it runs past that timeout — the stage doesn’t stop, and the agent still starts either way. The outcome is reported as a `session.setup_script` activity event on the run rather than only through whatever the script itself printed — `ok`, `failed`, `timeout`, `unsupported`, `error`, or `env_unreadable` on the host path, or `delegated` (no timing of its own) when a stage container ran it.

### The per-task tool cache

Every stage container for the same task shares a persistent tool cache mounted at `/opt/ch-tools`, so a toolchain one stage installs is still there for the next one — your build stage doesn’t have to reinstall the same dependencies your review stage already pulled down. This cache is a stage-container feature only: a host-process stage has no `/opt/ch-tools` mount, and a host `before_script` exporting a `PATH` change to it wouldn’t reach the agent anyway (see **The setup script** above). On the default image, the cache’s own directories are already on `PATH`. On a custom image, your image’s own `PATH` is left exactly as you built it, so add `/opt/ch-tools/bin` to `PATH` yourself in `before_script` (or from the agent) if you want to use it. The cache is scoped to the task, and stale caches for finished tasks are reclaimed automatically over time.

## Isolation with a custom image

A stage running in your own image gets the same hardening a default-image stage does: every extra Linux capability dropped, privilege escalation disabled, the same memory, CPU, and process caps, the same filtered outbound network, and a fresh, disposable container torn down the moment the stage ends. See [Isolating agent runs on a device](https://codeherder.com/docs/agent-isolation/#option-2-a-fresh-container-per-stage) for the full list.

The one difference: the default image runs as a fixed non-root user CodeHerder controls. A custom image runs as whatever user your image itself defaults to — CodeHerder doesn’t override it, since forcing a different user could break an image that expects to own its own files.

## Getting the image onto the device

The device only pre-fetches its own default image ahead of time. A custom image is pulled the first time a stage actually needs it, so it has to be reachable from wherever that stage runs. If it lives in a private registry, the device needs to already be authenticated to that registry on its own — CodeHerder doesn’t supply credentials for an image you name yourself.

## The workflow editor doesn’t show these settings, but saving it is safe

The Workflows editor in the web app has no fields for `image` or `before_script` — but for a hand-written type, saving the type from there no longer touches them. The editor only changes the pipeline’s shape (stages, gates, transitions); it carries every stage’s content, including `image` and `before_script`, through untouched. The one exception is a stage you remove outright: its content, image and setup script included, goes with it.

## Related guides

- [Isolating agent runs on a device](https://codeherder.com/docs/agent-isolation/) — the per-stage container mode this page depends on, and the isolation these settings get
- [Customising workflows](https://codeherder.com/docs/task-types/) — where each stage setting lives, and the CLI schema round-trip
- [Choosing the coding-agent CLI your agents run](https://codeherder.com/docs/harnesses/) — which CLI your image needs to provide
- [Managing your devices](https://codeherder.com/docs/devices/) — device readiness and the Docker checks that apply to per-stage containers
