# how to quarantine a compromised container

## TL;DR

Isolate first, investigate second. Cut the container off the network before you touch anything else, snapshot its logs and filesystem for forensics, then kill it and rebuild from a clean image. Never try to "clean" a compromised container in place; the image is untrusted the moment you suspect compromise.

```text
how to quarantine a compromised container
```

## Use this when

- A runtime alert, scanner, or teammate flags a container as possibly compromised
- You see unexpected outbound connections or processes in a container
- A container is mining crypto, spawning shells, or talking to unknown hosts
- You need a repeatable containment runbook for container incidents

## Not for this skill when

- The host itself is compromised (bigger problem; isolate the whole host)
- You are doing routine vulnerability scanning (prevention, not response)
- You need to find how the attacker got in (that is the post-mortem, after containment)
- The workload runs on a managed platform with its own isolation flow (use the platform tooling)

## Steps

### 1. Cut the network

Disconnect the container from every network so it cannot phone home, move laterally, or exfiltrate more data. Do this before you exec into it or poke at it.

```bash
docker network disconnect [network-name] [container-name]
```

Expected: the container loses connectivity. On Kubernetes, apply a deny-all NetworkPolicy selecting the pod's labels instead of deleting the pod; you want it alive but silent for step 2.

### 2. Preserve evidence

Grab the logs and a filesystem snapshot before the container disappears. Logs go to your incident store; a committed image preserves the filesystem exactly as the attacker left it.

```bash
docker logs [container-name] | tee incident-[date]-container.log
```

Expected: a log file with the container's output up to now. Then snapshot the filesystem:

```bash
docker commit [container-name] [registry]/forensics/[name]:[date]
```

Expected: a new image tagged for forensics, pushed where analysts can pull it later. Never run this image as a workload; it is evidence, and it is still hostile.

### 3. Stop the container

Now kill it. The evidence is preserved, so there is no reason to keep it running.

```bash
docker stop [container-name] && docker rm [container-name]
```

Expected: the container is gone from `docker ps`. On Kubernetes, delete the pod only after checking what the controller will reschedule; if the deployment still references the bad image, pause the rollout or fix the image reference first.

### 4. Check siblings and scan the image

Look for the same compromise in other containers built from the same image, and scan the image for the vulnerability that was likely the entry point.

```bash
trivy image [registry]/[name]:[tag]
```

Expected: scan results showing known CVEs in the image, or a clean image (which points at an app-layer bug instead). Any sibling container showing the same odd behavior gets the same quarantine treatment. (trivy 0.x.)

### 5. Rebuild from a clean base and redeploy

Rebuild with no cache from a patched base image, push a new tag, and redeploy. Then rotate every secret the compromised container could reach; its environment and mounted secrets are all suspect.

```bash
docker build --no-cache -t [registry]/[name]:[new-tag] .
```

Expected: the new container passes health checks and the odd behavior is gone. Confirm the old tag is no longer referenced by any deployment; a stale reference will happily reschedule the compromised image.

### Variant: you only suspect, not confirmed

Same steps. Quarantine is cheap and reversible; a breach you waited to confirm is not. If the investigation clears the container, you lost an hour, not a network.

### Variant: managed platforms (ECS, Fargate, Cloud Run)

Use the platform's stop and drain controls instead of docker commands, and pull logs from the platform's logging service before the task disappears. The order stays the same: isolate, preserve, stop, rebuild.

### Variant: the container keeps coming back

A rescheduling loop means the bad image reference or a malicious admission hook is still in place. Check the deployment spec, the image pull policy, and any mutating webhooks before you redeploy again; killing the pod without fixing the source is mowing the lawn.

## Why this happens

Containers share the host kernel, so a vulnerable app or a bad base image gives attackers a foothold with the container's network access and mounted secrets attached. Attackers love containers because teams treat them as disposable and hesitate to kill them for fear of losing evidence. The runbook fixes both failure modes: snapshot first so evidence survives, then kill without mercy because the image can never be trusted again.

## Edge cases and pitfalls

- Host escape: check host processes and scheduled tasks; if the host looks suspect, isolate the whole host too.
- docker commit on a huge container is slow and disk-hungry; capture logs first, image second.
- Do not exec into the container to "look around" before isolating; interaction can trigger anti-forensics.
- Mounted volumes survive container deletion; treat shared volumes as potentially tainted and scan them.
- Forensics images are still hostile; pull and analyze them only in an isolated environment.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_aPUMJRLmpz9kgwWMZ2qV1Q
