TL;DR: The same code cannot give two answers, so something in the inputs moved - usually the tool version, a stale incremental call-graph index, or the code snapshot the check ran against. Pin the commit, pin the tool, cache the call graph keyed by code hash, and the flip-flop stops.

```text
the reachability check flip-flopped: vulnerable function reachable on Monday, "not called" on Tuesday, same code
```

## Steps

1. Record the exact inputs of both runs: commit hash, call-graph tool name and version, source tree or SBOM hash, and any cache or index paths used.
   Expected: at least one input differs between Monday and Tuesday - that is your variable.
2. Re-run the check on the pinned Monday commit with the pinned tool version, from a clean cache.
   Expected: the verdict matches Monday's; if it matches Tuesday's instead, the cache was the variable.
3. Store the call graph as a build artifact keyed by the code hash, and make the reachability check consume the stored artifact instead of rebuilding incrementally.
   Expected: an identical code hash always yields the identical graph.
4. Pin the tool version in the pipeline config (not "latest") and log it with every verdict.
   Expected: version bumps become visible events, not silent verdict changes.
5. Add a regression guard: if a verdict changes while the code hash and tool version are unchanged, flag the run as suspect instead of updating the ticket.
   Expected: flip-flops page the pipeline owner instead of rewriting triage history.

## Use this when

- Reachability verdicts alternate on the same commit
- "Not called" appears on days the full scan did not run (incremental or indexed runs)
- Verdicts changed right after a tool or base-image update
- Two environments (CI vs the agent's sandbox) disagree on the same code

## Not for this skill when

- The code did change (even a dependency bump counts) - re-baseline first
- The check is an LLM summarizing advisories rather than a real call-graph analysis - that needs the evidence-bound verdict fix instead
- The function sits behind a feature flag the check never evaluated - that is a coverage gap, not a flip-flop

## Variant phrasings

- reachability analysis gives different results on rerun, same commit
- call graph says not called one day, reachable the next
- how to make vulnerability reachability checks reproducible

## Why it happens

Reachability is a function of code, tool, and graph inputs - change any one and the answer can change. The usual movers: Monday's check ran on a fresh full-program graph while Tuesday's ran on a stale incremental index that missed the caller; the tool auto-updated overnight and its graph builder changed; or the two runs snapshotted different revisions because main moved. The verdict looked like it was about the code, but it was really about the pipeline.

## Edge cases

- In monorepos "the commit" is repo-wide but the check may only index one package - pin the package-level hash too.
- Generated code (protobuf, mocks) can shift call edges without a source change - include generated files in the hash, or exclude them explicitly and document it.
- A stored graph goes stale the moment code changes - key strictly by hash and rebuild on miss, never reuse a graph across hashes.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_OjNIFT2tdZys65wTTSlktA
