trivy scan timed out after 20 minutes when the agent ran it against a 4GB container image - the OOM killer took the...
A recovery and prevention playbook for trivy scans that time out or get OOM-killed on large container images: confirming the OOM kill, sizing memory, setting an explicit scan timeout, warming the vulnerability DB, and splitting the work via server mode or SBOM-first scanning. Use when an agent's trivy run dies on a multi-GB image. Not for trivy DB download failures or registry auth errors.
TL;DR
Trivy's default scan timeout is 5 minutes and large images need far more memory than a small CI runner provides, so the agent's 20-minute timeout and the OOM kill are two separate problems with two separate fixes. Confirm the OOM kill in the kernel log, give the scan real memory, set --timeout longer than the harness budget, warm the vulnerability DB once with --download-db-only, and split the heavy lifting with server mode or SBOM-first scanning.
The query
trivy scan timed out after 20 minutes when the agent ran it against a 4GB container image - the OOM killer took the scan mid-layerUse this when
- trivy dies partway through a large image scan
- The scan process vanishes with no trivy error message
- An agent's trivy step keeps hitting the harness timeout on big images
- You need to make large-image scanning reliable in a pipeline
Not for
- "FATAL failed to download vulnerability DB" (that is a network or mirror problem)
- Registry authentication failures ("unauthorized" on pull)
- trivy fs or repo scans of source directories (different memory profile)
Steps
1. Confirm the OOM killer did it
dmesg | grep -i "killed process" | tail -5Expected output: a line naming the trivy process as the OOM victim. If there is no such line, the death was the timeout, not memory - treat them as separate problems.
2. Give the scan real memory
A 4GB image with many layers needs several GB of RAM for layer extraction and analysis. Raise the container or runner memory limit to at least 8GB for large images, and do not run the scan on the same small runner as the build.
Expected output: the scan completes instead of vanishing mid-layer.
3. Set an explicit timeout longer than the harness budget
Trivy's default timeout is 5 minutes, which large images blow past immediately. Set it explicitly:
trivy image --timeout 30m your-image:tagAnd make sure the agent harness timeout is larger than the trivy timeout, so trivy fails cleanly with its own error instead of being killed silently.
Expected output: no more silent kills; if it times out, trivy says so itself.
4. Warm the vulnerability DB once, before the scan
trivy image --download-db-only --no-progressThe first run downloads the DB (~60MB plus the Java index); every scan after that reuses the cache. An agent that downloads the DB inside every scan wastes minutes and memory on every run.
Expected output: subsequent scans skip the download and start analyzing immediately.
5. Reduce the work: scan only what matters
trivy image --scanners vuln --skip-dirs /usr/share/doc your-image:tagDrop the misconfig and secret scanners when you only need vulnerabilities, and skip directories that cannot contain anything scannable.
Expected output: lower peak memory and a faster scan for the same vulnerability results.
6. Split the heavy lifting: server mode on a beefy host
Run trivy server on a machine with plenty of RAM, and point the agent's lightweight client at it. The server binds on the scan host's address; the client just needs to reach it:
trivy server
trivy client --server YOUR_SCAN_HOST:4954 --timeout 30m your-image:tagExpected output: the agent runner stays small; the memory-hungry analysis happens on the server.
7. Or go SBOM-first
Generate the SBOM once with syft, then scan the SBOM instead of the image:
syft your-image:tag -o spdx-json > sbom.json
trivy sbom sbom.jsonExpected output: the vulnerability scan runs in a fraction of the memory because layer extraction is already done.
Variant phrasings
trivy scan killed on large image
Steps 1 and 2: confirm the OOM kill, then give it memory. Most "trivy just dies" cases are memory, not bugs.
trivy timeout on big container image
Step 3: the 5-minute default is the trap. Set --timeout explicitly and keep the harness budget above it.
trivy out of memory scanning image
Steps 2, 5, and 7 in that order: more RAM, less work, or split the work.
Why it happens
Trivy extracts every layer of the image to analyze file contents, and a 4GB image with dozens of layers is a lot of filesystem to hold in memory at once. Small CI runners and agent sandboxes typically ship with 2 to 4GB, so the kernel's OOM killer picks the biggest process - the scan - and shoots it mid-layer with no error message. Meanwhile the 5-minute default timeout was designed for small images, so large scans trip it long before they finish. Two resource limits, two failures, one confused agent.
Edge cases
- The OOM kill happens during DB download, not the scan: that is a different bottleneck; warm the DB on a bigger machine and copy the cache.
- Multi-arch images: trivy analyzes each platform variant; pin --platform to the one you ship to cut the work.
- Air-gapped environments: pre-seed the DB tarball and the image tarball, then scan with --input; do not let the agent try to reach the network mid-scan.
- Server mode needs its own timeout and memory sizing: the server is where the memory goes now, not the client.
- A scan that passes with --scanners vuln but the pipeline also needs misconfigs: run the misconfig scan as a separate, cheaper step against the Dockerfile, not the image.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstpTalB882VuKuWfP5Ia7QQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.