# Diagnose device tunnel failures

Source: https://codeherder.com/docs/self-host-tunnel-failures/

Find why a device keeps going offline, using only ch, the device logs and the support bundle. Then hand off to the support contact without a credential.

Use this page when a device goes offline and comes back, or sessions on it restart. CodeHerder staff have no access to your install. You run every step. See [Self-hosted operator access and changes](https://codeherder.com/docs/self-host-operator-access/).

## What you need

- A `ch` login as a workspace member. Fixes that change settings need a workspace administrator.
- An instance operator login, to produce the support bundle. See [Self-hosted support bundle](https://codeherder.com/docs/support-bundle/).
- A shell on the device, held by the device owner.

## Keep credentials out

Never paste a token, `CH_TOKEN`, an API key, `service.env`, the device token file or a raw log line into a ticket, email or chat. Read logs on the device. Send the bundle and typed answers only. Read the bundle before you send it.

## Step 1: confirm the fault

Run these commands from any machine with `ch`:

```
ch device list
ch device show <device>
ch instance check --skip mail --skip release_retrieval
```

`ch device show` prints whether the device is online, its version and one row for each readiness check. Note the `tunnel_connected` row. The `device_connection` line of `ch instance check` fails when no device is online. See [Self-hosted synthetic checks](https://codeherder.com/docs/self-host-checks/).

## Step 2: read the server timeline

The server records each connect and disconnect in the workspace activity feed. Run:

```
ch activity workspace <workspace> --subject-id <deviceId> --since 24h --limit 200 \
  --type device.tunnel_connected,device.tunnel_disconnected,device.tunnel_superseded,device.tunnel_error_reported
```

Count the disconnects for each hour. Note whether they come at a regular interval, such as every 60 or 120 seconds. A regular interval points at a proxy or load balancer timeout. Irregular drops that match sleep times point at the device’s power settings.

A `device.tunnel_superseded` event means a second device server used the same device token. Stop one of the two. The event does not say why the socket closed. The reason is only in the device log.

## Step 3: read the device log

The device owner does this step on the device. Do not copy the lines off the device.

| Where the device server runs | Log |
| --- | --- |
| Any host, as a user | `~/.codeherder/logs/device-server.log` |
| Linux systemd user service | `journalctl --user -u <your-unit> --since "1 hour ago"` |
| macOS launchd | `/tmp/codeherder-device-server.log` |
| Docker | `docker logs <container> --since 1h` |

Search for `tunnel session ended` and `tunnel: connected`. See [Self-hosted log retention and access](https://codeherder.com/docs/self-host-log-retention/#device-logs). If the same line holds `short header` or `EOF`, something between the device and the server cut the socket. Record the count of such lines. Do not paste them.

## Step 4a: check your proxy or load balancer

The device and the server exchange a heartbeat every 30 seconds. The server drops a tunnel after 75 seconds of silence. So the idle or read timeout on every proxy, load balancer and firewall between them must be longer than 30 seconds. Set it to 120 seconds or more.

Also allow the `Upgrade` header on `/v1/device-tunnel`. See [Put a reverse proxy in front](https://codeherder.com/docs/self-hosting/#put-a-reverse-proxy-in-front). If a firewall or VPN sits in front, allow long-lived WebSocket connections through it.

After you change the timeout, watch the timeline from step 2 for an hour.

## Step 4b: check the device’s power

A laptop that sleeps drops its tunnel. On macOS, run `pmset -g log` and `pmset -g assertions`. On Linux, search `journalctl` for suspend lines. Keep the laptop awake, or run the device server as a service. See [Running the device server](https://codeherder.com/docs/running-the-device-server/).

## Answer table

Fill this in before you hand off. Type each value. Do not paste log lines.

| Question | Answer |
| --- | --- |
| Device ID |  |
| Time window (UTC) |  |
| Is the device online now (yes or no)? |  |
| Failing check keys from `ch device show` |  |
| Disconnects in the window, and the interval |  |
| `short header` or `EOF` lines on the device (count) |  |
| `device.tunnel_superseded` events (count) |  |
| Proxy or load balancer idle timeout, in seconds |  |
| Power or sleep events in the window (yes or no) |  |

## Hand off to the support contact

If steps 4a and 4b do not fix the fault, or you cannot find the cause, hand off. Send these five items over the channel in [Self-hosted support and incident response](https://codeherder.com/docs/self-host-support/):

1. The support bundle file from `ch instance support-bundle`.
2. The device ID from `ch device list`.
3. The time window, in UTC.
4. The filled answer table.
5. The severity you choose from [Self-hosted support and incident response](https://codeherder.com/docs/self-host-support/).

Do not send a token, `service.env`, the device token file or a raw log. The bundle holds no device data, so the answer table carries the device facts.

The support contact sends back:

- The severity level and the response time from [Self-hosted support and incident response](https://codeherder.com/docs/self-host-support/).
- What the support contact reads in the bundle, and the answer.
- The next step or fix for you to run.
- Whether a product defect exists, and which release fixes it.

## Related guides

- [Self-hosted provisioning breaker](https://codeherder.com/docs/self-host-provision-breaker/) — release a device that stays ineligible for a task
- [Self-hosted diagnosis exercise](https://codeherder.com/docs/self-host-diagnosis-exercise/) — rehearse both runbooks on a staged fault
- [Self-hosted support bundle](https://codeherder.com/docs/support-bundle/) — what the bundle holds
- [Self-hosted support and incident response](https://codeherder.com/docs/self-host-support/) — severity levels and response times
