TL;DR: Stop letting the model vote on reachability from vibes - require it to produce a call path from an entry point to the vulnerable function, and default to "unknown, treat as reachable" when it cannot. Pin the advisory text it sees and the tool versions it calls, so identical inputs can only produce identical verdicts.

```text
the agent's exploitability assessment contradicted itself between runs - the LLM couldn't decide if the vulnerable function was reachable
```

## Steps

1. Change the verdict schema so reachability only allows three values: "reachable" (with a call path), "not_reachable" (with full-graph evidence of no path), and "unknown."
   Expected: the model can no longer emit a bare "not reachable" with nothing behind it.
2. Require a structured "call_path" field: an ordered list of caller function names from an entry point down to the vulnerable function.
   Expected: every verdict ships with the path, or an explicit empty path plus "unknown."
3. Set the default: an empty call path means "unknown," and your policy treats "unknown" as "patch it."
   Expected: the flip-flop between "reachable" and "not reachable" settles into a stable "unknown" until real evidence exists.
4. Pin the inputs: snapshot the advisory text, SBOM, and code revision into the run artifact and hash them.
   Expected: two runs on the same hash see byte-identical context, so verdicts stop drifting.
5. Re-run the disagreeing CVE twice on the pinned snapshot and compare.
   Expected: identical verdicts; if they still differ, the nondeterminism is in the tool call, not the context - log the tool outputs and diff them.

## Use this when

- Two runs of the same agent give different reachability verdicts for the same CVE
- The agent says "not reachable" but cannot show the call path
- An auditor asks "how do you know" and the answer is "the model said so"
- Reachability verdicts change after an advisory text update with no code change

## Not for this skill when

- The disagreement is between two different scanners (trivy vs grype counts) - that is scanner consistency, not an agent-verdict problem
- The code actually changed between runs - re-run on the pinned revision first
- You need reachability for a stripped binary with no symbols - the call-graph tool cannot see the code at all

## Variant phrasings

- LLM triage agent gives different exploitability verdicts on rerun
- agent says vulnerable function reachable then not reachable, same code
- how to make agent reachability assessments deterministic

## Why it happens

The model is not doing reachability analysis - it is summarizing whatever advisory text and partial tool output happened to be in context. Different runs get different context (cache warmth, advisory updates, truncated tool output), so the "analysis" changes. Without a required evidence field, "not reachable" is just the model's guess wearing a verdict's clothes, and guesses vary run to run.

## Edge cases

- A genuine "not reachable" needs the whole program in the call graph - a partial graph that misses a caller produces a false "not reachable."
- Entry points the graph does not know about (plugins, dynamic loading, reflection) make "not reachable" unsafe - document the graph's blind spots inside the verdict.
- Lowering temperature helps but does not fix input drift - pin the inputs first, tune the model second.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_1R9dFTdrDSj7-Pkrok2LCg
