## TL;DR
Grype loads the whole vulnerability DB and builds the full match set in memory, so a monorepo with 4000 transitive dependencies can blow past any small runner's RAM and the OOM killer ends the scan. The fix is sharding: scan each workspace or package directory separately instead of the repo root, generate one SBOM per shard, and merge the JSON results. No single grype run should ever see the whole monorepo at once.

## The query

```text
grype burned 9GB of RAM and got OOM-killed scanning a monorepo with 4000 transitive dependencies - how to shard the scan
```

## Use this when

- grype is killed by the OOM killer on a large repo
- A monorepo scan's memory grows with every added workspace
- An agent's nightly scan never finishes on the monorepo root
- You need per-team vulnerability reports from one repo

## Not for

- "failed to fetch vulnerability data URIs context deadline exceeded" (that is a DB download problem)
- "unable to parse lockfile" (that is a lockfile format problem)
- Scanning a single small service (no sharding needed)

## Steps

### 1. Confirm the OOM kill

```bash
dmesg | grep -i "killed process" | tail -5
```

Expected output: a line naming the grype process as the OOM victim. If the process exited some other way, you are chasing the wrong problem.

### 2. List the natural shard boundaries

```bash
ls packages/
```

In a monorepo the shards are the workspaces: packages/[name], apps/[name], services/[name]. Each shard should be the smallest unit that has its own lockfile or dependency manifest.

Expected output: a list of workspace directories, each independently scannable.

### 3. Scan one shard at a time

```bash
for d in packages/*/; do
  grype dir:"$d" -o json > "scan-$(basename $d).json"
done
```

Each grype process only loads the packages of one workspace, so peak memory stays proportional to the largest shard, not the whole repo.

Expected output: one JSON result file per workspace, and no OOM kills.

### 4. Go SBOM-first per shard for the big ones

For the heaviest workspaces, split cataloging from matching:

```bash
syft packages/big-app -o spdx-json > sbom-big-app.json
grype sbom:sbom-big-app.json -o json > scan-big-app.json
```

Expected output: the matching step uses far less memory because package cataloging is already done.

### 5. Merge the shard results

Combine the per-shard JSON files into one report. Count unique CVEs by id across shards so a dependency shared by three workspaces is reported once, not three times.

Expected output: a single merged report with deduplicated CVE counts.

### 6. Cap memory as a backstop

Set GOMEMLIMIT or the container memory limit on the scan job so a pathological shard fails fast with a clear error instead of taking the whole runner down with it.

Expected output: worst case is a failed shard with a clear message, not a dead runner.

### 7. Pre-download the DB once

```bash
grype db check
```

Run this before the shard loop so every shard reuses the local DB instead of each process trying to fetch or verify it.

Expected output: the DB is current locally; shard scans start matching immediately.

## Variant phrasings

### grype out of memory on monorepo

Steps 1 through 3: confirm the kill, find the shard boundaries, scan per workspace.

### how to scan a large repo with grype without OOM

The whole playbook. Sharding is the answer; there is no flag that makes one giant scan cheap.

### grype dir scan uses too much RAM

Steps 3 and 4: narrow the target to one workspace, and go SBOM-first for the heavy ones.

## Why it happens

Grype's matching step compares every discovered package version against the vulnerability database, and the data structures for that grow with the number of packages times the number of candidate vulnerabilities. A monorepo multiplies this: 4000 transitive dependencies across dozens of workspaces, many of them duplicates of the same library at different versions, each getting its own match evaluation. Memory grows superlinearly with the package count, so the repo-root scan needs 9GB while no single workspace needs more than a fraction of that. Sharding turns one impossible scan into many trivial ones.

## Edge cases

- Shared vendored directories scanned once per shard: if every workspace vendors the same tree, exclude the vendor dir from all but one shard or the SBOM step catalogs it repeatedly.
- A single workspace that is itself huge: shard it further by subdirectory, or accept that one shard needs a bigger runner.
- grype db updates mid-loop: pin the DB version for the whole run so shards do not disagree with each other.
- Deduplicating across shards: the same CVE in three workspaces is one finding for the security team but three for the workspace owners. Produce both views.
- Dev dependencies and test fixtures inflate the package count: exclude them from the shard scans if your policy only tracks production dependencies, and say so in the report.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_i4oxpr235PnEdD3daRt-og
