two agent runs scored the same CVE's exploitability differently because the second run had a warmer cache of advisory...
Makes agent exploitability verdicts independent of cache state by snapshotting the exact advisory text into each run artifact and hashing the full model context. Use it when reruns of the same triage agent disagree and the only difference is cache warmth. Key trigger: identical CVE and code, different verdict, warmer cache on the second run.
TL;DR: The second run was not smarter - it just saw more advisory text, so its "analysis" differed. Snapshot the exact advisory set into the run artifact, hash the full model context, and require matching context hashes before you compare two verdicts. Cache becomes a recorded input instead of a hidden variable.
two agent runs scored the same CVE's exploitability differently because the second run had a warmer cache of advisory textSteps
- For both runs, dump what the agent actually saw: the advisory texts retrieved, the SBOM or package list, and the code snippets in context.
Expected: the advisory sets differ - the warmer run had texts the first run never fetched.
- Change the pipeline to snapshot every advisory text it retrieves into the run artifact before the verdict step, and pass the snapshot (not a live cache lookup) to the agent.
Expected: the verdict step always sees the same frozen advisory set for a given run.
- Hash the full context (advisories plus SBOM plus code revision plus prompt) and store the hash alongside the verdict.
Expected: two verdicts are only comparable when their context hashes match.
- Re-run the CVE twice from the same snapshot.
Expected: identical verdicts; if they still differ, the nondeterminism is in the model call itself (temperature, tool-call ordering), not the cache.
- Add a staleness rule: if the advisory snapshot is older than your refresh window, refresh it and re-run rather than reusing a warm cache silently.
Expected: cache hits are explicit and timestamped, never invisible.
Use this when
- Reruns of the same agent disagree on the same CVE with no code change
- Verdicts seem to get "better" the more the agent has run (a warming pattern)
- You cannot reproduce a verdict because you do not know what context it saw
- Advisory text updates change old verdicts without any re-analysis being logged
Not for this skill when
- The inputs genuinely changed (new advisory published, code bumped) - that is a legitimate re-score, just log what changed
- The disagreement is between two different agents or tools - compare their evidence, not their caches
- The verdict flip-flops on identical context - look at model temperature and tool-call nondeterminism instead
Variant phrasings
- agent exploitability score changes between runs, same CVE
- how to make LLM vulnerability triage reproducible
- cache warming changed the agent's CVE verdict
Why it happens
The agent's verdict is a function of its context, and a cache is context that changes silently. Run one sees three advisories; run two's cache has warmed and it sees seven, including the one with the proof-of-concept note. The model did not reason differently - it read different material. Because the cache state was never recorded, the two runs looked identical from the outside and the disagreement looked like flakiness.
Edge cases
- Snapshots bloat artifacts fast - store advisory texts once per pipeline run and reference them by hash from each verdict.
- A frozen snapshot can go stale mid-incident - timestamp it and define when "frozen" becomes "expired."
- Tool-call results (search hits, fetched pages) are context too - snapshot those the same way, or the cache sneaks back in through the tools.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstEwRQYtPuHn9bxi6TUc-Wg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.