disk full on a production server: safe cleanup order
Gives a safe cleanup order for a full production disk: rotate logs first, then caches and temp files, then check for deleted-but-open files. Use when services fail with no space left on device. Not for inode exhaustion or filesystem corruption.
TL;DR
Clean up in this order: rotate and compress logs first, then clear caches and temp files, then old package files, and only as a last resort touch application data. Never delete a file a running process has open (the space will not free), and never rm -rf anything you have not listed first. A full disk also breaks the database and the package manager, so move fast but in the right order.
Error / query
disk full on a production server: safe cleanup orderUse this skill when
- df shows 100% (or 95%+) on a production filesystem
- Services are failing with "no space left on device"
- The database stopped accepting writes
- You need to free space right now without breaking the running system
Not for this skill when
- The disk is an inode exhaustion (df -i shows 100%); deleting big files will not help, find the many-small-files culprit
- The filesystem is corrupted; that is fsck territory, not cleanup
- You have time for a proper fix; add disk or move data instead of deleting
Steps
Step 1: Confirm which filesystem is full and how bad it is
df -h / /var /tmp 2>/dev/nullExpected: the Use% column; 100% on /var is the classic case, and the mount point tells you where to hunt.
Step 2: Find the biggest directories on the full filesystem
du -x -d 2 /var 2>/dev/null | sort -rn | head -10Expected: the top space hogs; /var/log is the usual winner by a mile.
Step 3: Rotate and compress logs first (safest space win)
journalctl --vacuum-size=500M; logrotate -f /etc/logrotate.conf 2>/dev/null; echo "logs rotated"Expected: "logs rotated" and df showing freed space; vacuuming the journal alone often frees gigabytes.
Step 4: Clear package caches and temp files next
apt-get clean 2>/dev/null; rm -rf /tmp/* /var/tmp/* 2>/dev/null; echo "caches cleared"; df -h /var | tail -1Expected: more freed space; these are regenerable, so deleting them is safe.
Step 5: Check for deleted-but-open files holding space
lsof +L1 2>/dev/null | head -5; df -h /var | tail -1Expected: ideally no output from lsof; if a process holds a deleted log open, restart that service to actually free the space.
Variant phrasings
"No space left on device but df shows free space"
Deleted files held open by processes (step 5) or inode exhaustion; check df -i next.
"Docker ate all my disk"
docker system prune -af plus volume pruning; overlay and unused images are the usual hogs on Docker hosts.
"Log file is 50GB, can I just delete it"
Truncate it (echo to it or use truncate -s) instead of deleting; deleting a file the app has open frees nothing until restart.
Why it happens
Logs grow without rotation, package caches accumulate, and temp files never get cleaned; the disk fills slowly for months and then everything breaks at once because databases, package managers, and even SSH need free space to function.
Edge cases and pitfalls
- Never rm -rf /var/log or the app's data dir to "make space"; you will destroy the evidence and possibly the app.
- Truncating a live log loses the recent entries; copy the tail somewhere first if you might need it.
- A database with a full disk may need recovery steps after space is freed; check it actually accepts writes again.
- Set up disk usage alerting at 80% after this; the next full disk should page you, not surprise you.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_UMHAin84gX1Sx6EUY5NiTQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.