## TL;DR

Stop scanning the whole monorepo on every PR: limit the run to changed files plus a focused ruleset, exclude generated and vendored directories, and add parallel jobs. That combination usually takes a timing-out scan down to minutes. Raising the CI timeout is the last resort, not the fix.

## The error

```text
$ semgrep scan --config=auto --baseline-ref=main
...
Error: scan exceeded the job time limit (20m) - timed out
```

## Fix it

1. Prove scope is the problem. List the changed files with `git diff --name-only main...HEAD`, then run semgrep on just those files with your normal config.
   Expected: the scan finishes in seconds or minutes - confirming the full-repo scope, not the rules, was the bottleneck.
2. Exclude the noise permanently. Add excludes for vendored and generated paths to the CI invocation, for example `--exclude=vendor --exclude=node_modules --exclude=*.min.js --exclude=*.lock`.
   Expected: a sharp drop in scan time on the next run.
3. Trim the ruleset for the PR gate. Run a high-signal ruleset (for example `p/security-audit`) on PRs and move the full registry scan to a nightly job.
   Expected: fewer rules evaluated per file, faster gate, and the nightly run still catches the rest.
4. Parallelize. Add `-j 4` (or match your runner's core count) to the scan command.
   Expected: near-linear speedup up to the available cores.
5. Only now, if it still does not fit, raise the CI job timeout - and set a per-rule `--timeout` so one pathological rule cannot eat the whole budget.
   Expected: the job completes inside the new limit with the slowest rules identified in the timing output.

## Use this when

- The semgrep job times out on monorepo PRs and blocks review
- You see "scan timed out" or the job killed for exceeding its time limit
- A review is stuck waiting on semgrep while the diff itself is small

## Not for this skill when

- Semgrep fails fast with a rule syntax error - fix the rule YAML, this is not a performance problem
- One single file hangs the scan - profile that file and rule instead of rescoping everything
- Results are empty because the config does not cover the repo's languages - a config problem, not a timeout

## Variant phrasings

- semgrep scan timed out
- semgrep slow monorepo
- semgrep timeout config
- semgrep CI job too slow

## Why it happens

`--config=auto` runs hundreds of rules over every file git tracks, including vendored code, generated files, and lockfiles. On a monorepo that is hours of rule evaluation per run. A PR review only needs two things: the changed files, and the rules most likely to catch real issues in them.

## Edge cases

- `--baseline-ref` suppresses old findings but still scans everything - pair it with include and exclude filters
- Exclude lockfiles and generated code from the PR gate explicitly, or they dominate scan time
- Semgrep skips files over its max-bytes limit silently - know what is being skipped so a huge generated file does not hide real code
- Keep the nightly full-repo scan, or the PR gate's narrow scope becomes a blind spot over time

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_aqIJ5RHdZy8hwSC9uUgotw
