how to quarantine a compromised container
A containment runbook for a suspected-compromised container: cutting network access first, preserving logs and a filesystem snapshot for forensics, stopping the workload, scanning the image and checking siblings, then rebuilding from a clean base. Use when a runtime alert or odd behavior flags a container, in Docker or Kubernetes. Triggers: 'compromised container', 'container isolation'. Not for: host-level compromise, routine image scanning.
how to quarantine a compromised container
TL;DR
Isolate first, investigate second. Cut the container off the network before you touch anything else, snapshot its logs and filesystem for forensics, then kill it and rebuild from a clean image. Never try to "clean" a compromised container in place; the image is untrusted the moment you suspect compromise.
how to quarantine a compromised containerUse this when
- A runtime alert, scanner, or teammate flags a container as possibly compromised
- You see unexpected outbound connections or processes in a container
- A container is mining crypto, spawning shells, or talking to unknown hosts
- You need a repeatable containment runbook for container incidents
Not for this skill when
- The host itself is compromised (bigger problem; isolate the whole host)
- You are doing routine vulnerability scanning (prevention, not response)
- You need to find how the attacker got in (that is the post-mortem, after containment)
- The workload runs on a managed platform with its own isolation flow (use the platform tooling)
Steps
1. Cut the network
Disconnect the container from every network so it cannot phone home, move laterally, or exfiltrate more data. Do this before you exec into it or poke at it.
docker network disconnect [network-name] [container-name]Expected: the container loses connectivity. On Kubernetes, apply a deny-all NetworkPolicy selecting the pod's labels instead of deleting the pod; you want it alive but silent for step 2.
2. Preserve evidence
Grab the logs and a filesystem snapshot before the container disappears. Logs go to your incident store; a committed image preserves the filesystem exactly as the attacker left it.
docker logs [container-name] | tee incident-[date]-container.logExpected: a log file with the container's output up to now. Then snapshot the filesystem:
docker commit [container-name] [registry]/forensics/[name]:[date]Expected: a new image tagged for forensics, pushed where analysts can pull it later. Never run this image as a workload; it is evidence, and it is still hostile.
3. Stop the container
Now kill it. The evidence is preserved, so there is no reason to keep it running.
docker stop [container-name] && docker rm [container-name]Expected: the container is gone from docker ps. On Kubernetes, delete the pod only after checking what the controller will reschedule; if the deployment still references the bad image, pause the rollout or fix the image reference first.
4. Check siblings and scan the image
Look for the same compromise in other containers built from the same image, and scan the image for the vulnerability that was likely the entry point.
trivy image [registry]/[name]:[tag]Expected: scan results showing known CVEs in the image, or a clean image (which points at an app-layer bug instead). Any sibling container showing the same odd behavior gets the same quarantine treatment. (trivy 0.x.)
5. Rebuild from a clean base and redeploy
Rebuild with no cache from a patched base image, push a new tag, and redeploy. Then rotate every secret the compromised container could reach; its environment and mounted secrets are all suspect.
docker build --no-cache -t [registry]/[name]:[new-tag] .Expected: the new container passes health checks and the odd behavior is gone. Confirm the old tag is no longer referenced by any deployment; a stale reference will happily reschedule the compromised image.
Variant: you only suspect, not confirmed
Same steps. Quarantine is cheap and reversible; a breach you waited to confirm is not. If the investigation clears the container, you lost an hour, not a network.
Variant: managed platforms (ECS, Fargate, Cloud Run)
Use the platform's stop and drain controls instead of docker commands, and pull logs from the platform's logging service before the task disappears. The order stays the same: isolate, preserve, stop, rebuild.
Variant: the container keeps coming back
A rescheduling loop means the bad image reference or a malicious admission hook is still in place. Check the deployment spec, the image pull policy, and any mutating webhooks before you redeploy again; killing the pod without fixing the source is mowing the lawn.
Why this happens
Containers share the host kernel, so a vulnerable app or a bad base image gives attackers a foothold with the container's network access and mounted secrets attached. Attackers love containers because teams treat them as disposable and hesitate to kill them for fear of losing evidence. The runbook fixes both failure modes: snapshot first so evidence survives, then kill without mercy because the image can never be trusted again.
Edge cases and pitfalls
- Host escape: check host processes and scheduled tasks; if the host looks suspect, isolate the whole host too.
- docker commit on a huge container is slow and disk-hungry; capture logs first, image second.
- Do not exec into the container to "look around" before isolating; interaction can trigger anti-forensics.
- Mounted volumes survive container deletion; treat shared volumes as potentially tainted and scan them.
- Forensics images are still hostile; pull and analyze them only in an isolated environment.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_aPUMJRLmpz9kgwWMZ2qV1Q
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.