agent ran the full e2e suite per candidate version instead of the affected subset - 4 hours per bump, 12 bumps deep
Fixes agents that run the entire e2e suite for every candidate dependency version. Use when upgrade verification takes hours per bump because full suites run on every version candidate. Key trigger: 4-hour e2e runs multiplied across many candidate versions.
TL;DR: Run a tiered check per candidate: unit tests for the changed package plus the e2e specs that touch its code paths, then reserve the full e2e suite for the final merged candidate only. Map the bump to affected specs using the dependency graph or path filters. Verification drops from hours per candidate to minutes, with the full suite still guarding the merge.
agent ran the full e2e suite per candidate version instead of the affected subset - 4 hours per bump, 12 bumps deep- Measure which e2e specs actually import or exercise the bumped package: search the spec files for its import path and list the matches.
Expected: a short list of relevant specs, not the whole suite.
- Change the agent's verification plan to two tiers: tier 1 runs the package's unit tests plus the matched specs per candidate; tier 2 runs the full e2e suite once on the chosen version before merge.
Expected: per-candidate runs take minutes; one full run guards the merge.
- Give the agent a blast-radius map: package.json dependency edges or an import graph, so it can expand the tier-1 set when the bump is a shared utility (logging, HTTP client, date lib).
Expected: widely-used packages automatically get wider tier-1 coverage.
- Verify with the next bump: time tier 1 per candidate and confirm the full suite ran exactly once.
Expected: 12 candidates cost 12 short runs plus one full run, not 12 full runs.
Use this when
- upgrade verification runs the full e2e suite per candidate version
- e2e takes hours and the agent tests many candidates
- most e2e failures on bumps are in a small set of specs
- you want faster agent iteration without losing the merge gate
Not for this skill when
- the e2e suite IS the only test coverage (then you need the full run, or more unit tests)
- bumps regularly break unrelated specs through side effects (the blast radius is real; run everything)
- the failure mode is flaky e2e, not slow e2e (fix flakiness first)
Variant phrasings
- "agent runs full e2e for every version candidate"
- "dependency bump verification too slow"
- "how to run affected tests only for a version bump"
- "e2e suite 4 hours per dependency upgrade"
- "tiered CI for dependabot PRs"
Why it happens
Full e2e per candidate is the safest default, and for one bump it is fine. It breaks down when the agent evaluates many candidates: the cost multiplies by candidate count while the information gained stays flat, because most specs never touch the bumped code. A tiered run keeps the same final gate and spends the compute where the risk is.
Edge cases
- Shared infra packages (logging, serialization) touch nearly every spec; their tier-1 set is effectively the whole suite. Accept it for those.
- The import search can miss specs that hit the package through a network call; include integration specs for the service boundary when the bump is a client lib.
- A green tier-1 plus red full suite means the blast-radius map was wrong; widen it for that package next time.
- Flaky specs in tier 1 will make the agent churn through candidates; quarantine known-flaky specs out of the per-candidate path.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst__XIWhPM5NwgfCsuoDEPVKA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.