## TL;DR
Self-hosted runners accumulate cruft: dead registrations, gigabytes of old workspaces, and stale Docker images. Run them as ephemeral one-job runners so each job gets a fresh registration, clean workspaces and images on a schedule, and auto-remove runners that stop heartbeating. Ephemeral mode is the real fix; cleanup scripts are the bandage.

## Error / query
```text
how to clean up GitHub Actions self-hosted runners automatically
```

## Use this skill when
- Offline runners pile up in the runner list
- Runner disks fill with old `_work` directories
- Stale runners grab jobs and then fail
- You want hands-off runner hygiene

## Not for this skill when
- You use GitHub-hosted runners (nothing to clean)
- Runners fail to register at all (registration problem)
- Workflows themselves are broken (workflow debugging)

## Steps

### Step 1: Remove dead runner registrations
```bash
gh api repos/[owner]/[repo]/actions/runners --paginate -q '.runners[] | select(.status=="offline") | .id' | xargs -I{} gh api -X DELETE repos/[owner]/[repo]/actions/runners/{}
```
Expected: offline runner IDs disappear from the runner list. Run this on a schedule (weekly cron); offline runners never come back on their own.

### Step 2: Switch to ephemeral one-job runners
```bash
./config.sh --url https://github.com/[owner]/[repo] --token [registration-value] --ephemeral --unattended
./run.sh
```
Expected: each runner takes exactly one job then unregisters and exits. Your orchestrator (autoscaler, VM image, container) starts a fresh runner per job, so stale state cannot accumulate by design.

### Step 3: Clean workspaces and Docker cruft on persistent runners
```bash
# cron on each runner host, runs when idle:
find [HOME]/... -maxdepth 2 -mtime +7 -exec rm -rf {} +
docker system prune -af --filter "until=168h"
```
Expected: workspaces older than 7 days and unused images older than 7 days are removed. Schedule this for idle windows; pruning while a job runs can delete layers the job is pulling.

### Step 4: Alert on disk before it fills
```bash
df -h [HOME]/... | awk 'NR==2 {print $5}'
```
Expected: a usage percentage you can feed to your monitoring. Alert at 80 percent so cleanup (or a runner rebuild) happens before jobs start failing on disk pressure.

## Variant phrasings

### "github actions self-hosted runner disk full"
Steps 3-4: prune workspaces and images, then alert before it recurs.

### "remove offline github runners"
Step 1 as a scheduled job. Ephemeral runners (step 2) make it unnecessary.

## Why it happens
A self-hosted runner is a pet: it registers once and accumulates every job's workspace, logs, and pulled images forever. Nothing in the runner cleans up after itself, so without automation the disk fills and the registration list fills with corpses. Ephemeral runners turn the pet into cattle.

## Edge cases and pitfalls
- `--ephemeral` runners cannot pick up a second job; your scaler must start one runner per queued job or jobs queue.
- `docker system prune -af` deletes build cache too; the next job's build will be slower. Prefer `until` filters over unconditional prunes.
- Deleting a busy runner's registration via API orphans its in-flight job; only auto-remove runners that have been offline for a grace period.
- Runner labels determine job routing; when rebuilding runners from images, keep labels identical or jobs stop matching.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_voVauOJ6YHo-nGLb0emP5w
