# secrets detection in pre-commit hooks

## TL;DR
Secrets get committed by accident in seconds and live in git history forever, so catch them before the commit lands. Add a secrets scanner to the pre-commit config, tune out the false positives once, and make the hook fast enough that nobody disables it. Anything the hook misses is a rotation job, not a revert.

```text
secrets detection in pre-commit hooks
```

## Use this when
- A credential was committed recently and you want it to never happen again
- You are setting up a new repo or standardizing across repos
- Onboarding docs need the "install our hooks" section
- An agent is auditing repo hygiene or developer tooling
- CI catches secrets but you want the feedback earlier

## Not for this skill when
- You are scanning dependencies for CVEs (different tooling)
- You are choosing where secrets live at runtime (vault territory)
- The leak already happened and you need incident response (rotate first)
- You are detecting secrets in production logs (different pipeline)

## Steps

### 1. Install pre-commit and add the scanner hook
Add a secrets detection hook to the repo's pre-commit config alongside your formatters. Keep it in the shared config so every clone gets it, not just your machine.

```bash
pre-commit install
pre-commit run --all-files
```

Expected: the install succeeds and the first full run completes. Note how long it takes; if it is slow, developers will bypass it, so speed matters as much as accuracy.

### 2. Do one baseline run and triage the findings
The first run across the whole repo will flag historical items. Triage each: real secret (rotate it), test fixture (allowlist it explicitly), or false positive (tune the rule).

```bash
pre-commit run --all-files | head -40
```

Expected: a triaged list with an owner and action for every finding. "We will deal with it later" is how secrets stay in history for years.

### 3. Tune the rules instead of disabling the hook
Every scanner ships noisy default patterns. Add explicit allow entries for test fixtures and example keys, and tighten patterns that fire on your codebase's idioms. The config should show its work: each exception named and dated.

Expected: a second full run is clean, and the allowlist is small and reviewed. An allowlist longer than the findings list means the rules are wrong, not the code.

### 4. Handle a blocked commit gracefully
When the hook blocks a real secret, the developer removes it from the staged files, moves it to the secret manager or env file, and recommits. Document this flow in the repo readme so nobody is surprised at 6pm on a Friday.

Expected: the documented flow is three steps or fewer. If working around the hook is easier than following it, the flow is wrong.

### 5. Rotate anything that was committed before the hook existed
The hook only protects the future. For every real secret the baseline found in history, rotate it now: revoke the old value, issue a new one, and purge or accept the history exposure.

Expected: each historical finding has a rotation ticket closed, not just an allowlist entry. An allowlisted real secret is still a live credential in a public-ish place.

### 6. Mirror the check in CI
Pre-commit hooks can be skipped, so run the same scanner in CI on every PR as the backstop. The local hook gives fast feedback; CI gives the guarantee.

Expected: a PR that skips hooks still gets scanned in CI, and the CI job fails the build on new findings.

### Variant: monorepos with many teams
Put the hook config at the repo root and let teams add scoped exceptions in their directories. Centralize the scanner version so rule updates roll out everywhere at once.

### Variant: repos with lots of test fixtures
Fixtures full of fake keys are the top false-positive source. Keep them in clearly named fixture directories with a documented fake-key format, and allowlist the directory rather than sprinkling exceptions.

### Variant: scanning without pre-commit
If the team will not adopt pre-commit, a git pre-push hook or an IDE plugin running the same scanner covers most of the value. The worst option is CI-only, because the secret still lands in the branch history.

## Why this happens
Secrets end up in code because the distance between "works on my machine" and "committed" is one command, and the developer's attention is on the feature, not the config. Detection at commit time works because it is the last moment the secret is still private to one machine; after the push it is replicated to every clone, CI log, and backup.

## Edge cases and pitfalls
- Rebasing and cherry-picking can reintroduce cleaned secrets; the CI backstop catches what local hooks miss.
- Encrypted secret files are fine to commit, but the decryption key next to them is not; scan for both halves.
- Screenshots and demo videos in the repo can contain visible credentials; no text scanner catches those, so keep them out.
- Scanner updates add new patterns that flag old code; pin the version and review the changelog on upgrade.
- Developers working offline or in containers may not have hooks installed; the CI mirror is the guarantee, the hook is the convenience.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_qDFepz33cGgFKZNCbRXz6A
