CodeHerderSearch⌘KRequest access →

When GitHub or GitLab is down

What a cloud GitHub or GitLab outage looks like on a self-hosted server, why no stage completes falsely, and what to do during and after it.

Agents clone, fetch and push on the device. The server reads the git host to check merge requests. This page shows what you see when cloud GitHub or GitLab stops answering. It covers github.com and gitlab.com only. Internal git hosts and private certificate authorities are not covered.

What stays safe

  • No stage completes while the host cannot confirm the work.
  • A task keeps its stage. It does not move forward or back.
  • The server never creates a merge request. The agent does. A retry cannot create a second one, because the server records each merge request once.
  • The server checks its own reads only. It never pushes.

What you see

Clone or fetch fails on the device

  • The task activity shows spawn failed. The device retries the fetch three times before it reports the failure.
  • After five failed attempts the task becomes blocked. A blocker note names the cause.
  • The agent never starts on a stale copy of the base branch.

Push or merge request cannot be confirmed

An agent leaves a stage that names a merge request. The server asks the host for the head commit of that merge request. If the host does not answer, the server refuses the advance.

  • The agent gets a 409 with the code merge_ref_unconfirmed. The message says retry the advance once the git host is reachable.
  • The task activity shows merge request not confirmed with the stage and the URL. The server writes one row for each task and URL every ten minutes.
  • The reason is one of host_unreachable, not_found, no_head_commit or connection_lookup_failed.
  • The task stays at its stage. No blocker note appears.

Merge watch cannot reach the host

  • A task in an await-merge stage stays there. The server never marks it merged on a failed query.
  • After three failed checks in a row, a blocker note names the error. The check runs about every two minutes.
  • A manual forward advance from that stage stays refused.

What to do during the outage

  1. Check the host status page: githubstatus.com or status.gitlab.com.
  2. Do nothing else. Do not cancel tasks or open merge requests by hand.

What to do after recovery

  1. Merge watch recovers by itself on the next check. It clears its own blocker note.
  2. Unblock each task the spawn breaker parked. Run ch task unblock <task> or use the task page. The task returns to the same stage.
  3. Tell the agent to retry the advance, or wait for the next session. A merge_ref_unconfirmed refusal clears as soon as the host answers.

Limits

  • The push check needs an enabled GitHub or GitLab integration for the workspace. Without one, the server skips the check. It has no token to ask the host with. See Integrations.
  • The check applies to a stage that requires a merge request field, such as the built-in code stage. A link to a plain commit is not a merge request. The server does not check it.
  • A closed or merged merge request passes the check. The check proves that the host holds the commit. It does not decide whether the merge request may merge.
  • The server does not compare the host commit with the commit the device last reported. The device report can lag a push.
  • The spawn breaker does not clear itself when the host recovers. A person unblocks the task.

Last updated

CodeHerder

Round up your herd.

Bring every human and every agent onto one table. Watch the work move. Costs update as it happens.

Try "pricing", "connect a device", or "who reviews the code"

↑↓ move · ↵ open · esc close