## TL;DR
A Semgrep rule is a small YAML file with an id, the languages it applies to, a severity, and a pattern written in the target language with metavariables like [ARG] for the parts that vary. Write the pattern to match the dangerous code shape, add a clear message explaining the safe alternative, test it with semgrep --test against good and bad examples, then drop it into your CI config. One focused rule per bad pattern beats one clever rule that tries to catch everything.

```text
how to write a custom Semgrep rule
```

## Use this when
- The same bug class keeps recurring and you want CI to catch it forever
- You need to ban a dangerous API (eval-like calls, raw SQL concatenation) in your codebase
- A security review produced a 'never do this again' pattern
- You are evaluating whether Semgrep fits your team's workflow

## Not for
- Installing or configuring Semgrep in CI (setup docs cover that)
- Writing rules for other SAST engines
- Browsing the public Semgrep rule registry

## Steps
1. Write the smallest YAML skeleton. Every rule needs rules, an id, languages, severity (INFO, WARNING, or ERROR), a message, and a pattern. Keep ids namespaced like team-name.no-eval so they sort and read well.
   Expected output: a YAML file that semgrep --validate accepts without errors.

2. Express the bad pattern in the target language. Use metavariables in brackets for the varying parts, e.g. a pattern like os.system([CMD]) matches any call to os.system regardless of argument. Semgrep understands the languages grammar, so it matches code shape, not text, and ignores formatting differences.
   Expected output: the pattern matches the vulnerable snippet from your last incident.

3. Add pattern-not for the safe exceptions. If there is a legitimate use (a wrapper that sanitizes first), exclude it with pattern-not so the rule doesnt cry wolf on reviewed code.
   Expected output: the sanitized wrapper call no longer triggers the rule.

4. Test with semgrep --test. Put vulnerable examples in a test file marked with ruleid comments for expected hits, and safe examples that must not match. Run the test command; it fails if any expectation is wrong.
   Expected output: semgrep --test reports all expectations met, with zero unexpected matches.

5. Write a message that teaches. The message shows up in CI output, so make it say what is wrong and what to do instead in one or two sentences, with a link to your internal guidance if you have it.
   Expected output: a developer who has never seen the rule understands the fix from the message alone.

6. Ship it as WARNING first, promote to ERROR later. Run the rule in CI for a week or two, fix or waive the backlog it finds, then flip it to blocking once the codebase is clean.
   Expected output: the first ERROR-mode run breaks no innocent builds because the backlog was already handled.

## Variant phrasings
- "Semgrep custom rule tutorial"
- "Semgrep pattern vs pattern-either"
- "how to test Semgrep rules locally"
- "ban function with Semgrep rule"

## Edge cases and pitfalls
- Overly broad patterns (matching every function call with one argument) drown the team in noise; anchor on the specific dangerous callee.
- Different language versions parse differently; pin the languages list to what you actually ship.
- Rules rot: review custom rules yearly and delete the ones for APIs you no longer use.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_8y7ibS39tqUeUDTNqo1Fgw
