# "Cannot connect to the Docker daemon" in CI: docker-in-docker fix

TL;DR: In CI there is usually no daemon inside the job container, and the docker CLI defaults to a socket that does not exist there. Point `DOCKER_HOST` at a docker-in-docker service (GitLab) or mount the host socket (GitHub container jobs) and the CLI connects.

```text
Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running?
```

## Steps

1. Find out where the CLI is looking. Print `DOCKER_HOST` in the job; empty or unset means it fell back to the unix socket, which does not exist in the job container.
**Expected:** You know the intended daemon address.

2. On GitLab, declare the dind service and point the CLI at it:
```yaml
services:
  - name: docker:dind
variables:
  DOCKER_HOST: tcp://docker:2375
  DOCKER_TLS_CERTDIR: ""
```
**Expected:** `docker info` prints the server version.

3. On GitHub Actions container jobs, either run docker steps directly on the runner instead of inside the container, or mount the host socket into the container job with the appropriate container options.
**Expected:** Docker commands work from inside the container.

4. On self-hosted runners, confirm the daemon is actually running and the job user can reach it. Run `docker info` as that user.
**Expected:** The server section appears; if not, start the daemon or fix socket permissions.

5. Check for TLS mismatch. If `DOCKER_HOST` uses port 2376, certificates must be mounted and valid; mixing 2375 and 2376, or a stale cert directory, breaks the handshake.
**Expected:** Scheme, port, and certs agree.

## Use this when
- Docker CLI commands in CI fail with `Cannot connect to the Docker daemon`

## Not for this skill when
- The error is `permission denied` on the socket. The user is not in the docker group; different fix
- The error mentions TLS handshake failure. Certs are wrong, not the address
- The daemon is running but out of disk. That fails later with `no space left on device`

## Variant phrasings
### Cannot connect to the Docker daemon at tcp://docker:2375
The CLI is pointed at dind but the service is not up or not reachable; check the service declaration.

### Is the docker daemon running?
Same error, trailing question variant from older CLI versions.

## Why it happens
The docker CLI is only a client. CI job containers do not run a daemon, so the default unix socket path points at nothing. Something has to provide a daemon over TCP (dind) or share the host's socket, and the CLI has to be told where it is.

## Edge cases
- Setting `DOCKER_TLS_CERTDIR` to empty disables TLS for the dind service; that is the common simple setup, but it is plaintext on the job network
- DinD needs privileged runners; without that the service container never starts
- Kubernetes executors have their own way of declaring the dind service; the compose-style syntax does not apply

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_ANyM4A2Wf82gpX6oumbkeQ
