the agent's triage state didn't persist between runs - it re-scored the same 300-CVE backlog from scratch every night
Fixes a CVE triage agent that loses its triage state between runs and re-scores the same 300-CVE backlog every night. Use it when every run starts from an empty or stale state even though the previous run completed. Key trigger: triage results exist in logs but never appear in the database or state store on the next run.
TL;DR: Persist every triage decision to a durable store keyed by CVE ID plus package version, and read it back at the start of each run. The agent re-scores because nothing durable survives between runs - the state lived in memory, in an ephemeral container filesystem, or in a write that never committed. A single read-back check at run start proves the fix.
run 2026-10-07: scoring 300 CVEs from scratch (0 previously triaged found in store)- At the start of a run, query the triage store for a known CVE and log the stored decision. Expected: you see the stored decision; an empty result proves the store is empty or unreachable.
- Find the write call in the triage code and confirm the connection points at the production database, not a container-local file or in-memory structure. Expected: writes target a durable host that survives restarts.
- Make the write synchronous and confirmed: after each batch of decisions, read one record back. Expected: the read-back returns the decision just written.
- Re-run the nightly job twice in a row. Expected: the second run reports something like "280 previously triaged, 20 new" instead of scoring all 300.
- Add a startup assertion: if the store is unreachable, fail loudly instead of silently starting from empty. Expected: a bad database connection pages someone instead of burning a night of re-scoring.
Use this when
- triage state resets between agent runs
- the backlog gets fully re-scored every night
- "previously triaged" counts read zero at run start
- completed runs leave no trace in the state store
Not for this skill when
- state persists but the decisions themselves are wrong
- the scanner re-reports findings the agent already closed (a dedup-key problem)
- runs crash mid-way and lose partial progress (a checkpointing problem)
- the database itself is down (an infra problem)
Variant phrasings
- agent forgets triage results between runs
- triage database empty after successful run
- re-scoring same CVE backlog every night
- triage state not durable across runs
Why it happens
The agent wrote state to process memory, to a temp file inside an ephemeral container, or to a database connection that silently failed or never committed. The next run boots clean with no history, so it treats every CVE as new and burns a full scoring pass.
Edge cases
- writing to a read replica that lags makes state look missing - read state from the primary
- partial writes on mid-run crashes need per-batch checkpoints, not only end-of-run saves
- a schema migration that drops the triage table looks exactly like this - check migration logs first
- two agents sharing one store need namespaced keys or they overwrite each other's state
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_GLTCErNw9fa-EloxVmd4cQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.