CodeHerderSearch⌘KRequest access →

Self-hosted logs

Find each log a self-hosted server writes, its format, the headers and URL parameters to redact, and a tested Vector example to ship them.

A self-hosted CodeHerder server writes its logs to journald on the host. If the host is compromised, the local logs go with it. Ship them to storage you control. This page lists each log, its format, the values to redact, and one tested way to ship them.

Your organization owns the log storage. That includes retention, access control, and the SIEM. CodeHerder does not run them for you. For each log class, its readers and its retention control, see Self-hosted log retention and access.

Where each log lives

Every source writes to journald. Read it with journalctl -u <unit>.

Source Journald unit Format What it holds
CodeHerder server codeherder.service logfmt on stderr Application logs and one msg=http line for each request
Caddy access log caddy.service console or json, on stdout One line for each proxied request
Caddy runtime log caddy.service JSON on stderr Certificates, reloads, and errors
Sign-in and privilege facility auth (4) and authpriv (10) syslog text sshd, sudo, and PAM
Kernel audit transport audit text Audit records, when journald collects them

The instance plane is part of the codeherder process. Its requests appear in codeherder.service as msg=http lines. Its durable audit trail is the admin_audit_log table in PostgreSQL, with plane = 'instance'. That trail is not in journald. Read it with SQL.

Log formats

CodeHerder server

The server writes logfmt: key=value pairs, with a quoted value when it has spaces.

time=2026-09-29T03:13:23.062Z level=INFO msg=http method=GET path=/v1/tasks route=/v1/tasks status=200 durMs=12 reqId=req-1
  • time, level, and msg are on every line. level is DEBUG, INFO, WARN, or ERROR.
  • A request line has msg=http and the keys method, path, route, status, durMs, and reqId.
  • A request line is WARN when a request that does not stream runs slowly.
  • A request line has no header and no query string.
  • The server replaces the value of a sensitive key with [redacted].

Caddy

Use the json encoder for a log you ship. It gives a stable field set: ts, logger, msg, status, duration, request.method, request.uri, request.headers, and resp_headers. The SaaS app host’s reference Caddyfile uses console for the human-facing site. To ship that log, wrap json instead.

What to redact

The server redacts its own logs before it writes them. The ch_access_log snippet in Self-hosted deployment redacts the Caddy access log before journald sees it. If you use another proxy, or a collector, apply the same list yourself.

Delete these request headers:

  • Authorization
  • Proxy-Authorization
  • Cookie
  • Sec-WebSocket-Protocol
  • X-CH-Edge-Secret
  • X-Api-Key

Delete these response headers:

  • Set-Cookie
  • Sec-WebSocket-Protocol

Strip the whole query string from the logged URI, the Referer request header, and the Location response header. Do not keep a list of names for this. A future parameter would be missing from the list. These parameters carry secrets today:

Parameter Where it appears
code, state The OAuth callback for an integration, and the MCP OAuth callback
pendingToken The redirect after an integration sign-in
code_challenge, redirect_uri, request The MCP OAuth authorize and callback steps
X-Amz-Signature, X-Amz-Security-Token The presigned download redirect for a server release

The server also redacts these attribute names in its own logs, ignoring case, -, and _: Authorization, Proxy-Authorization, Cookie, Set-Cookie, X-Api-Key, Sec-WebSocket-Protocol, X-Ch-Edge-Secret, Api-Key, Token, Access-Token, Refresh-Token, Id-Token, Device-Token, Bearer, Password, Secret, Client-Secret, and Private-Key.

It also redacts any value that has the shape of a credential: a CodeHerder key or token (ch_, chd_, chia_, chp_, mcpa_, mcpc_, mcpr_, mcps_), a common third-party token (sk-, sk_live_, sk_test_, ghp_, gho_, ghr_, ghs_, ghu_, github_pat_, glpat-, xoxb-, xoxp-), a JWT, an AWS access key ID, a PEM private key, a user:password@ URL, and a Bearer value.

Ship logs off the host

The log-shipping bundle holds a Vector pipeline. It reads the five sources above from journald, parses the logfmt and JSON lines into fields, and redacts again before it ships. It adds a source field to each event: codeherder, caddy, auth, or audit.

  1. Install Vector 0.43.1 on the host. Run it as a systemd service that can read the journal. Add its user to the systemd-journal group.
  2. Copy vector.yaml and sink-cloudwatch.yaml to /etc/vector/.
  3. In sink-cloudwatch.yaml, set region and group_name.
  4. Give the host’s instance role these permissions: logs:CreateLogGroup, logs:CreateLogStream, logs:PutLogEvents, and logs:DescribeLogGroups.
  5. Check the files: vector validate --no-environment /etc/vector/vector.yaml /etc/vector/sink-cloudwatch.yaml.
  6. Start Vector with both files: vector --config /etc/vector/vector.yaml --config /etc/vector/sink-cloudwatch.yaml.

To ship somewhere else, replace sink-cloudwatch.yaml with another Vector sink that takes redact as its input. The redact step is the only thing you must keep.

CodeHerder tests this pipeline. The test starts a container, writes sample journal entries that carry a canary secret, runs Vector, and checks that the canary does not reach the output. It then removes the redaction step and checks that the canary does appear, so the test can see a leak. The kernel audit source is checked for syntax only, because a container cannot produce a kernel audit record.

Sign-in outcomes

The auth_audit_log table records each sign-in, token exchange and rejected credential. It holds the credential class (cognito_session, api_key, cli_installation, device_token, mcp_grant, otp, scim_token), the client address, the outcome and a short reason. It holds no token text and no email.

A success is written once per sign-in or exchange. A Cognito or OIDC session counts once per auth_time, so token refreshes add no rows and a new sign-in always adds one. Each accepted device tunnel handshake adds one row. The server also emits an auth.succeeded event into the workspace feed, so webhooks and the audit export carry it.

A failure is counted in memory and written once a minute. One row covers one class, address and reason, and attempts holds the count. A window keeps at most 100 keys. Extra keys fold into one row per credential class, with an empty address and reason overflow. The reason rate_limited means the server refused an address before it checked the credential. The address was over its failed sign-in budget, or it sent too many MCP token requests. A crash loses at most one minute of failure counts.

SELECT window_start, credential_class, client_address, reason, attempts
FROM auth_audit_log
WHERE outcome = 'failed' AND created_at > now() - interval '1 day'
ORDER BY attempts DESC;

Failures are not workspace events. Each flushed failure row also writes one auth.failed line to codeherder.service journald. The line holds credential_class, reason, client_address, attempts and window_start. It is bounded like the table, and Vector ships it to your SIEM. Set retention for this table with the other audit logs.

Alert on sign-in

Alert your SIEM on these signals. See the Audit egress guide (docs/audit-egress.md) for the webhook and OCSF export.

  • auth.failed log lines with a high attempts count from one address.
  • Any auth.failed line with reason=overflow or reason=rate_limited.
  • auth.failed with class scim_token or device_token. Your IdP and your devices should not send bad tokens.
  • auth.succeeded (webhook or OCSF export) from an address that member has never used.
  • The existing cognito.jwt rejected line.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close