agent quarantined a test that only failed because the agent itself left a stale docker container running
Diagnoses a test that got quarantined as flaky when the real cause was a stale docker container the agent itself left running. Use it when a quarantined test fails with port conflicts, resource collisions, or state that only exists because of a leftover container from an earlier agent session. Not for container failures caused by the test code itself, and not for infra-wide docker outages.
TL;DR The test was not flaky - the agent left a stale container running from an earlier session, and the leftover state broke the test. Kill the orphan, un-quarantine the test, and make the agent clean up after itself.
agent quarantined a test that only failed because the agent itself left a stale docker container runningSteps
- Prove the container is the cause. List running containers (
docker ps) on the machine where the test failed and look for orphans from earlier agent sessions: old names, long uptimes, ports the test needs.
Expected: you find a container holding the port, volume, or DB the test uses.
- Kill the stale container and rerun the test with nothing else changed.
Expected: the test passes, which proves the failure was environmental, not a flake.
- Remove the quarantine marker from the test.
Expected: the test runs in the suite again instead of being skipped.
- Find why the agent left it running. Check the agent's session logs for a container start with no matching stop: crashed sessions, forgotten
docker run -d, or cleanup steps that only ran on success.
Expected: you identify the exact session and command that orphaned it.
- Add cleanup guardrails: the agent stops or removes containers it starts (a teardown step that runs even on failure), and a pre-run check that fails fast if the needed ports are already bound.
Expected: the next agent session cannot orphan a container silently.
Use this when
- a quarantined test fails with port conflicts or "address already in use"
- the failure involves resources a previous agent session touched
- the test passes after you manually clean up containers
- the agent ran docker commands during its debugging session
Not for this skill when
- the container failure comes from the test's own setup code
- the problem is a cluster-wide docker or infra outage
- the test fails identically with no containers running at all
- no agent session touched containers on that machine
Variant phrasings
- "stale docker container left by agent breaks test"
- "test quarantined as flaky but a leftover container caused it"
- "address already in use from orphaned container in CI"
Why it happens
Agents start containers to reproduce failures and sometimes never stop them: the session crashes, the cleanup step is skipped on error, or the agent just moves on. The orphan holds ports, locks files, or serves stale data. The next run's test fails against that ghost state, the agent sees an intermittent failure, and labels it flaky instead of looking at docker ps.
Edge cases
- The stale container may be on a shared runner used by other jobs; coordinate before killing.
- Container names may collide across parallel jobs; use unique names per run to make orphans identifiable.
- A test that depends on "no container running" is itself fragile; the pre-run port check makes the assumption explicit.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_jW0myLjZX26kQ3iev1JF0Q
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.