Self-hosted logs
Find each log a self-hosted server writes, its format, the headers and URL parameters to redact, and a tested Vector example to ship them.
A self-hosted CodeHerder server writes its logs to journald on the host. If the host is compromised, the local logs go with it. Ship them to storage you control. This page lists each log, its format, the values to redact, and one tested way to ship them.
Your organization owns the log storage. That includes retention, access control, and the SIEM. CodeHerder does not run them for you. For each log class, its readers and its retention control, see Self-hosted log retention and access.
Where each log lives
Every source writes to journald. Read it with journalctl -u <unit>.
| Source | Journald unit | Format | What it holds |
|---|---|---|---|
| CodeHerder server | codeherder.service |
logfmt on stderr | Application logs and one msg=http line for each request |
| Caddy access log | caddy.service |
console or json, on stdout |
One line for each proxied request |
| Caddy runtime log | caddy.service |
JSON on stderr | Certificates, reloads, and errors |
| Sign-in and privilege | facility auth (4) and authpriv (10) |
syslog text | sshd, sudo, and PAM |
| Kernel audit | transport audit |
text | Audit records, when journald collects them |
The instance plane is part of the codeherder process. Its requests appear in codeherder.service as msg=http lines. Its durable audit trail is the admin_audit_log table in PostgreSQL, with plane = 'instance'. That trail is not in journald. Read it with SQL.
Log formats
CodeHerder server
The server writes logfmt: key=value pairs, with a quoted value when it has spaces.
time=2026-09-29T03:13:23.062Z level=INFO msg=http method=GET path=/v1/tasks route=/v1/tasks status=200 durMs=12 reqId=req-1
time,level, andmsgare on every line.levelisDEBUG,INFO,WARN, orERROR.- A request line has
msg=httpand the keysmethod,path,route,status,durMs, andreqId. - A request line is
WARNwhen a request that does not stream runs slowly. - A request line has no header and no query string.
- The server replaces the value of a sensitive key with
[redacted].
Caddy
Use the json encoder for a log you ship. It gives a stable field set: ts, logger, msg, status, duration, request.method, request.uri, request.headers, and resp_headers. The SaaS app host’s reference Caddyfile uses console for the human-facing site. To ship that log, wrap json instead.
What to redact
The server redacts its own logs before it writes them. The ch_access_log snippet in Self-hosted deployment redacts the Caddy access log before journald sees it. If you use another proxy, or a collector, apply the same list yourself.
Delete these request headers:
AuthorizationProxy-AuthorizationCookieSec-WebSocket-ProtocolX-CH-Edge-SecretX-Api-Key
Delete these response headers:
Set-CookieSec-WebSocket-Protocol
Strip the whole query string from the logged URI, the Referer request header, and the Location response header. Do not keep a list of names for this. A future parameter would be missing from the list. These parameters carry secrets today:
| Parameter | Where it appears |
|---|---|
code, state |
The OAuth callback for an integration, and the MCP OAuth callback |
pendingToken |
The redirect after an integration sign-in |
code_challenge, redirect_uri, request |
The MCP OAuth authorize and callback steps |
X-Amz-Signature, X-Amz-Security-Token |
The presigned download redirect for a server release |
The server also redacts these attribute names in its own logs, ignoring case, -, and _: Authorization, Proxy-Authorization, Cookie, Set-Cookie, X-Api-Key, Sec-WebSocket-Protocol, X-Ch-Edge-Secret, Api-Key, Token, Access-Token, Refresh-Token, Id-Token, Device-Token, Bearer, Password, Secret, Client-Secret, and Private-Key.
It also redacts any value that has the shape of a credential: a CodeHerder key or token (ch_, chd_, chia_, chp_, mcpa_, mcpc_, mcpr_, mcps_), a common third-party token (sk-, sk_live_, sk_test_, ghp_, gho_, ghr_, ghs_, ghu_, github_pat_, glpat-, xoxb-, xoxp-), a JWT, an AWS access key ID, a PEM private key, a user:password@ URL, and a Bearer value.
Ship logs off the host
The log-shipping bundle holds a Vector pipeline. It reads the five sources above from journald, parses the logfmt and JSON lines into fields, and redacts again before it ships. It adds a source field to each event: codeherder, caddy, auth, or audit.
- Install Vector 0.43.1 on the host. Run it as a systemd service that can read the journal. Add its user to the
systemd-journalgroup. - Copy
vector.yamlandsink-cloudwatch.yamlto/etc/vector/. - In
sink-cloudwatch.yaml, setregionandgroup_name. - Give the host’s instance role these permissions:
logs:CreateLogGroup,logs:CreateLogStream,logs:PutLogEvents, andlogs:DescribeLogGroups. - Check the files:
vector validate --no-environment /etc/vector/vector.yaml /etc/vector/sink-cloudwatch.yaml. - Start Vector with both files:
vector --config /etc/vector/vector.yaml --config /etc/vector/sink-cloudwatch.yaml.
To ship somewhere else, replace sink-cloudwatch.yaml with another Vector sink that takes redact as its input. The redact step is the only thing you must keep.
CodeHerder tests this pipeline. The test starts a container, writes sample journal entries that carry a canary secret, runs Vector, and checks that the canary does not reach the output. It then removes the redaction step and checks that the canary does appear, so the test can see a leak. The kernel audit source is checked for syntax only, because a container cannot produce a kernel audit record.
Related guides
- Self-hosted log retention and access — readers and retention for each log class, audit receiver and device
- Self-hosted deployment — download, verify, run, and put a reverse proxy in front of the server
Sign-in outcomes
The auth_audit_log table records each sign-in, token exchange and rejected credential. It holds the credential class (cognito_session, api_key, cli_installation, device_token, mcp_grant, otp, scim_token), the client address, the outcome and a short reason. It holds no token text and no email.
A success is written once per sign-in or exchange. A Cognito or OIDC session counts once per auth_time, so token refreshes add no rows and a new sign-in always adds one. Each accepted device tunnel handshake adds one row. The server also emits an auth.succeeded event into the workspace feed, so webhooks and the audit export carry it.
A failure is counted in memory and written once a minute. One row covers one class, address and reason, and attempts holds the count. A window keeps at most 100 keys. Extra keys fold into one row per credential class, with an empty address and reason overflow. The reason rate_limited means the server refused an address before it checked the credential. The address was over its failed sign-in budget, or it sent too many MCP token requests. A crash loses at most one minute of failure counts.
SELECT window_start, credential_class, client_address, reason, attempts
FROM auth_audit_log
WHERE outcome = 'failed' AND created_at > now() - interval '1 day'
ORDER BY attempts DESC;
Failures are not workspace events. Each flushed failure row also writes one auth.failed line to codeherder.service journald. The line holds credential_class, reason, client_address, attempts and window_start. It is bounded like the table, and Vector ships it to your SIEM. Set retention for this table with the other audit logs.
Alert on sign-in
Alert your SIEM on these signals. See the Audit egress guide (docs/audit-egress.md) for the webhook and OCSF export.
auth.failedlog lines with a highattemptscount from one address.- Any
auth.failedline withreason=overfloworreason=rate_limited. auth.failedwith classscim_tokenordevice_token. Your IdP and your devices should not send bad tokens.auth.succeeded(webhook or OCSF export) from an address that member has never used.- The existing
cognito.jwt rejectedline.
Last updated