Release a provisioning breaker
Find why a task stays blocked after repeated provision failures, fix the cause and release the breaker. Hand off to the support contact without a credential.
Use this page when a task stays blocked, or when a device stays “ineligible for this task” after you fixed its fault. CodeHerder staff have no access to your install. You run every step.
What you need
- A
chlogin as a workspace administrator, signed in as a person. An agent login or a session credential cannot release the breaker. - A shell on the device, held by the device owner.
- An instance operator login, to produce the support bundle.
Keep credentials out
Never paste a token, CH_TOKEN, an API key, service.env, the device token file or a raw log line into a ticket, email or chat. Read logs on the device. Send the bundle and typed answers only. Read the bundle before you send it. See Self-hosted support bundle.
How the breaker works
A device that fails to start a session for a task twice in a row is skipped for that task for 10 minutes. The server settings CH_DEVICE_PROVISION_FAILURE_THRESHOLD and CH_DEVICE_PROVISION_COOLDOWN set these numbers. A device that is only busy does not count as a failure. When every eligible device is cooling down, the task waits. The task also keeps its own count of failures, on any device. The fifth provision_failed in a row parks the task as blocked behind a note that only a person can clear. A full disk on the device parks it on the third failure.
| Count | Result |
|---|---|
| 2 failures on one device | That device is skipped for this task for 10 minutes. The feed shows task.device_provision_ineligible. |
5 provision_failed failures for the task |
The task is blocked behind a note that only a person can clear. |
3 device_disk_full failures for the task |
The task is blocked behind a note that only a person can clear. |
Step 1: confirm the fault
ch task show <task>
ch task blockers <task> --active
ch activity task <task> --type task.device_provision_ineligible,task.device_provision_breaker_cleared
A task.device_provision_ineligible event names the device and the cooldown. Note the time of the first one.
Step 2: find the cause
ch device show <device>
Each failing readiness row names a fix. Common causes are git_credentials, gh_auth and harness_auth. Fix the cause on the device. The device owner does this step. For a git credential, see Git tokens on a device. Do not edit database tables by hand. That does not release the task.
Step 3: release the breaker
Many tasks release without you. The breaker clears when a device that failed a readiness check passes it again. It also clears when you resolve the blocker.
If the task stays stuck, run this as a person with a workspace administrator role:
ch device clear-provision-breaker <device>
The reply reads cleared provision breaker for N task(s). With --json, the reply is {"data":{"cleared":N}}. For each released task, the activity feed shows task.device_provision_breaker_cleared. Blocked tasks return on the next sweep, which takes up to 5 minutes.
A count of 0 is not a failure. It means nothing was left to release. This happens when the device recovered from a failed readiness check before you ran the command: the server released the breaker then. Look for task.device_provision_breaker_cleared in ch activity task <task> to confirm.
The command refuses when the device is archived. It also refuses when you call it with an agent or session credential.
Step 4: clear other blockers
ch device clear-provision-breaker does not clear a blocker that a person filed. List active blockers with ch task blockers <task> --active. Resolve each one that no longer applies with ch task unblock <blockerId>.
Answer table
Fill this in before you hand off. Type each value. Do not paste log lines.
| Question | Answer |
|---|---|
| Device ID | |
| Task ID | |
| Time window (UTC) | |
| Task status now | |
Failing check keys from ch device show |
|
Count of task.device_provision_ineligible events |
|
| Did you fix the cause (yes or no, and what)? | |
Did ch device clear-provision-breaker run (yes or no, and the cleared count)? |
|
| Active blockers after release |
Hand off to the support contact
If the task stays stuck 10 minutes after the release, or you cannot find the cause, hand off. Send these items over the channel in Self-hosted support and incident response:
- The support bundle file from
ch instance support-bundle. - The device ID from
ch device list. - The time window, in UTC.
- The task ID.
- The filled answer table.
Do not send a token, service.env, the device token file or a raw log.
The support contact sends back:
- The severity level and the response time from Self-hosted support and incident response.
- What the support contact reads in the bundle, and the answer.
- The next step or fix for you to run.
- Whether a product defect exists, and which release fixes it.
Related guides
- Diagnose device tunnel failures — find why a device keeps going offline
- Self-hosted diagnosis exercise — rehearse both runbooks on a staged fault
- Self-hosted support bundle — what the bundle holds
Last updated