# Nightly fleet scan only finishes 60 of 200 images before the window closes

## TL;DR

Scan images in parallel across workers, download the vulnerability database once per host instead of per scan, and skip any image whose digest has not changed since the last scan. The agent scans 200 images one at a time and repeats identical work every night, so most of the window burns on downloads and re-scans of unchanged images. Parallel workers plus digest-based skipping turn the same window into full coverage.

## The failure

```text
nightly fleet scan completed 60 of 200 images, then the maintenance window closed
(sequential scans, fresh vulnerability DB download per scan, no skip for unchanged images)
```

## Steps

1. Record each scanned image's digest and scan result. On the next run, diff the fleet list against those digests and skip unchanged images. Expected: a large share (often over half) needs no scan at all, e.g. 200 images drops to 80 that actually changed.

2. Shard the remaining images across parallel workers. A simple split by worker index works: worker N takes every Nth image. Expected: wall-clock time drops roughly in proportion to worker count; 4 workers turn 4 hours into about 1.

3. Download the vulnerability database once per host before the scan starts (for trivy: `trivy image --download-db-only`), and point every worker at the shared copy. Expected: zero per-scan DB download waits in the worker logs.

4. Set a per-image timeout and keep a dead-letter list of images that always stall, so one bad image cannot eat the whole window. Expected: the run ends with every image either scanned, skipped as unchanged, or flagged for follow-up.

## Use this when

- a scheduled fleet scan covers a fraction of images before its window closes
- the agent scans images one after another with no parallelism
- vulnerability DB downloads repeat on every scan instead of once per host
- scan duration grows linearly with fleet size instead of staying flat

## Not for this skill when

- the scan finishes but results are wrong (that is a data quality problem, not throughput)
- only one or two images are slow (fix those images, not the fleet pipeline)
- the window is closing because the scanner itself crashes (fix the crash first)
- you scan fewer than about 20 images (parallelism overhead is not worth it yet)

## Variant phrasings

- container fleet vulnerability scan never completes overnight
- trivy scan of all images takes longer than the maintenance window
- nightly image scan only covers part of the registry

## Why it happens

Three costs compound. First, sequential scanning means total time equals the sum of every image, so the fleet outgrows the window as it grows. Second, each scan re-downloads the vulnerability database, which is the same multi-hundred-megabyte file every time. Third, unchanged images get fully re-scanned nightly even though their results cannot have changed. The fix attacks all three: parallelism divides the work, one DB download removes the repeated cost, and digest-based skipping removes work that was never needed.

## Edge cases

- Digest skipping hides new CVEs only if you also skip DB updates. Always refresh the DB before the run, even when most images are skipped, because new advisories can apply to old images.
- Parallel workers on one host can OOM if each loads the full DB. Cap worker count by memory, not just CPU.
- The dead-letter list needs an owner and a retry policy, or stall-prone images rot there forever.
- If your registry prunes untagged images, a digest you skipped may vanish. Re-resolve the fleet list from the registry each night before diffing.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_F6nw3JTrdwbot1LWYTUWfA
