the EPSS CSV feed changed its header row format and the agent's percentile math silently used the wrong column
Fixes EPSS enrichment that computed exploitability percentiles from the wrong CSV column after a header format change. Use when an agent's EPSS percentile math silently reads the wrong column. Key trigger: EPSS percentile values look wrong or out of range after a feed update.
EPSS CSV header change broke the percentile math
TL;DR
Index the EPSS CSV by header name ("epss" and "percentile") instead of column position, and add a range guard asserting every percentile falls between 0 and 1. The agent read the wrong column after a header format change, so every exploitability percentile fed into triage was garbage. The range guard turns the next silent shift into a loud failure on the first bad row.
The failure
EPSS percentile math silently used the wrong column after a header row format change
(percentile values out of range or nonsensical, triage scores corrupted, no error raised)Steps
- Switch the EPSS CSV reader to header-name lookup for the epss and percentile columns, using the exact header names from the current feed. Expected: percentiles fall between 0 and 1 on every row.
- Add a range assertion on every row: abort the sync if any percentile is outside 0 to 1. Expected: the next column shift fails immediately on the first row instead of corrupting the whole backlog.
- Re-score everything triaged during the broken window with the corrected parser. Expected: backlog percentiles match the published EPSS values again.
- Log the EPSS header row on every sync and alert when it changes. Expected: the next format change is noticed before triage consumes it.
Use this when
- EPSS percentiles look wrong, constant, or out of range
- exploitability scores changed sharply with no change in the threat landscape
- any numeric feed column is read by position rather than by name
- triage priority order stops correlating with real-world exploitation
Not for this skill when
- EPSS values are correct but your scoring formula misuses them (fix the formula)
- the EPSS API (not the CSV) is the source (same principle, JSON keys instead of headers)
- percentiles are missing rather than wrong (that is a coverage gap, not a shift)
- you do not use EPSS in triage at all
Variant phrasings
- EPSS percentile column wrong after CSV format change
- exploit prediction scoring reading wrong EPSS column
- EPSS feed parser header change broke scoring
Why it happens
EPSS publishes a CSV whose header row is not a stable contract, and positional readers assume it is. When the header format changed, the percentile column moved and the parser kept reading the old position, which now held a different metric. Because the wrong column still contained numbers, no type check caught it. Percentiles have a natural range of 0 to 1, which makes them ideal for a range guard: any shift that lands on a non-percentile column almost always produces out-of-range values, so the guard catches what the parser cannot.
Edge cases
- A shift onto another 0-to-1 column (like the epss score itself) passes the range guard. Spot-check a few known CVEs against the published EPSS values after any header change.
- EPSS updates daily. A stale local copy combined with a header change can produce confusing diffs. Always sync fresh before debugging the parser.
- If you cache EPSS scores, invalidate the cache for the broken window or corrected values will never reach triage.
- The range guard should log the offending row, not just abort. You need the actual values to diagnose which column moved where.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_NZKWDXZyK7gjviJkqSgO3g
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.