## TL;DR
Point syft at the lockfile, not the directory tree. Cataloging 90,000 files one by one is what eats the memory; the package-lock already lists every package and version, so scanning it directly gives the same SBOM at a fraction of the cost.

```text
syft OOM'd generating an SBOM for a node_modules tree with 90,000 files - the agent's SBOM step kept dying
```

## Steps

1. Confirm the OOM is file-count driven. Check how many files sit under node_modules versus how many packages the lockfile declares.
   - Run: count files under the dependency directory and compare with the package count in package-lock.json.
   - Expected: tens of thousands of files for a few thousand packages - the files are the problem.

2. Scan the lockfile directly instead of the directory. Point syft at package-lock.json (or the equivalent for your ecosystem) so it catalogs packages from the manifest.
   - Run: `syft packages file:./package-lock.json -o cyclonedx-json`, writing the SBOM to a file.
   - Expected: the SBOM completes in seconds with the same package list.

3. If you must scan the tree (lockfile missing or untrusted), exclude the heaviest directories with syft's exclude option and raise the container's memory limit for the SBOM step.
   - Run: `syft dir:. --exclude './node_modules/.cache' -o cyclonedx-json`, writing the SBOM to a file.
   - Expected: the scan finishes; excluded paths are documented in the run log.

4. Verify parity: diff the package list from the lockfile scan against a successful tree scan of a smaller project, or against the package manager's own listing.
   - Expected: same packages and versions in both.

5. Make the lockfile scan the agent's default for this repo and keep the tree scan as a fallback only when no lockfile exists.
   - Expected: the SBOM step stops OOMing on every run.

## Use this when

- Syft gets OOM-killed on a large node_modules or vendor tree.
- The SBOM step is the flakiest part of the agent's pipeline.
- A lockfile exists but the agent scans the directory anyway.
- Memory limits cannot be raised (shared CI runners, small containers).

## Not for this skill when

- The OOM happens on a container image scan; that is layer-handling memory, and the fix is scanning the image manifest or raising memory.
- The lockfile is out of sync with what is installed; fix the lockfile first, or the SBOM will be wrong.
- Syft crashes with a parse error rather than OOM; that is a cataloger bug, not a memory problem.

## Variant phrasings

- syft out of memory scanning node_modules
- SBOM generation killed OOM on large repo
- syft scan too slow on huge dependency tree
- how to make syft use less memory

## Why it happens

Syft's directory catalogers walk every file to detect packages, and a big node_modules tree means tens of thousands of file reads, hashes, and license detections held in memory. The lockfile already contains the exact package inventory the package manager resolved, so the file walk is pure overhead - the agent was paying the cost of discovery for information that was already written down.

## Edge cases

- A lockfile scan misses packages installed outside the package manager (vendored tarballs, manual copies); spot-check for those.
- Workspaces and monorepos have multiple lockfiles; scan each one or the SBOM will be incomplete.
- If the lockfile is generated mid-run by the agent itself, validate it installed cleanly before trusting it as SBOM input.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_JQT7IbUSXFfZ99THrcc9Tw
