gitlab ci runner debugging tips
Use when a GitLab CI job misbehaves and the pipeline YAML looks fine: the runner never picks up jobs, a job hangs or fails only in CI, or you need runner-side logs. Covers gitlab-runner --debug, CI_DEBUG_TRACE trace mode, local job replay with gitlab-runner exec, and where runner logs live. Not for fixing application errors already visible in the job log; points to existing skills for stuck-runner tags and failure_reason triage.
TL;DR
When a GitLab CI job misbehaves and the YAML looks fine, debug the runner, not the script. Three moves solve most cases: run the runner in the foreground with --debug to see if it even picks up jobs, set CIDEBUGTRACE to "true" on the job for full trace logs, and replay the job locally with gitlab-runner exec shell to separate script bugs from runner bugs. Most mystery failures are one of three things: the runner never receives the job (tags or token), the runner environment differs from what you expect, or the script is wrong in a way only the trace shows.
Verbatim symptom
This job is stuck because the project doesn't have any runners online
assigned to it. Go to Settings > CI/CD to assign runners.That banner means GitLab could not assign the job to any runner. It does not say why, which is exactly what the steps below find out.
Steps
1. Watch the runner decide in real time
Stop the runner service and run it in the foreground with debug logging (self-hosted runners):
gitlab-runner --debug runExpected output: a stream of lines like Checking for jobs... nothing, or, when it accepts work, Accepted job: [job-id]. Success check: trigger the pipeline and watch. If you never see Accepted job, the problem is on the scheduling side (tags, token, paused runner), not your script. Stop the foreground run and restart the service when done.
2. Turn on job trace mode
Add a variable to the job, either one-off in the pipeline UI or in the YAML:
variables:
CI_DEBUG_TRACE: "true"Expected output: the job log expands every script line with timestamps and environment context. Success check: you can point at the exact command where the job hung or failed. This works on any runner, including GitLab.com shared runners, and it is the fastest way to catch a script that fails only in CI.
3. Replay the job locally
From a checkout of the repo, on a machine with the runner installed:
gitlab-runner exec shell my-job-nameExpected output: the job runs locally through the same executor the runner uses. Success check: if the same failure happens locally, the bug is in the script or its assumptions, fix it there. If it passes locally, the bug is in the runner environment (images, caches, network), and step 4 is next.
4. Read the runner's own logs
Foreground debug is for now; for history, check where the runner actually writes:
journalctl -u gitlab-runner -n 100
# or, on installs that log to files:
tail -n 100 /var/log/gitlab-runner/*.logExpected output: registration errors, authentication failures, and executor spawn failures. Success check: an expired or revoked runner token, a failed health check, or a docker socket error shows up here instead of as a vague stuck job.
5. If the job is stuck, check the assignment rules before anything else
Two settings cause nearly every stuck job: the job has tags but the runner does not match them, or the runner was registered with "run untagged jobs" disabled while the job has no tags. This is covered in detail by the existing skill on runners online but jobs stuck (tags and shared runners); fix the tag match or enable untagged jobs and re-run.
6. If the job fails, triage by failure_reason
GitLab records why a job failed (scriptfailure, runnersystemfailure, jobexecutiontimeout, and others). Read that field in the job API response before touching code; a runnersystemfailure is never fixed by editing the YAML. Triage details live in the existing skill on failurereason.
When to use
- A job hangs or fails only in CI while the YAML looks correct.
- The stuck banner appears and you cannot tell whether it is tags, token, or something else.
- A newly registered self-hosted runner shows online but never picks up work.
- You need runner-side logs for an incident or an access review.
When NOT to use
- The job log already shows a clear application error. Fix the code, do not debug the runner.
- Do not leave CIDEBUGTRACE set on shared or production runners. Trace output can print secrets that are not masked; use it on one job, read the log, then remove the variable and rotate anything that was exposed.
- gitlab-runner exec does not support every feature (services and some artifact behavior differ). A local pass is a strong signal, not a proof.
Tool and version compatibility
- gitlab-runner 15, 16, and 17 on Linux, macOS, and Windows (log paths differ on Windows; check the service manager).
- CIDEBUGTRACE works with any executor and on GitLab.com shared runners.
- --debug run requires access to the runner host; you cannot do this on GitLab.com shared runners.
- gitlab-runner exec needs the repo checked out locally and the runner binary installed; the docker executor needs docker running locally.
Variant phrasings
- "gitlab runner not picking up jobs"
- "gitlab-runner debug mode"
- "gitlab CI job stuck debugging"
- "gitlab runner trace logs"
- "run gitlab ci job locally"
- "gitlab pipeline stuck in pending runner debugging"
Root cause, after the fix
Runner debugging separates three failure classes that look identical from the pipeline UI: the job never reaches a runner (assignment problem, steps 1 and 5), the runner environment differs from the script's assumptions (environment problem, steps 3 and 4), or the script itself is wrong (script problem, steps 2 and 3). Without runner-side logs you are guessing which class you are in; one debug run or one trace log tells you.
Edge cases
- A paused runner in the UI accepts nothing but still shows online; check the runner list before debugging further.
- Runner tokens rotate and old registration tokens expire; re-register if logs show authentication errors.
- Concurrent job limits on a single runner (concurrent in config.toml) silently queue jobs as stuck; raise it or add runners.
- Debug flags on helm or docker-based runner installs go in the deployment values, not on a command line you can reach.
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.