TL;DR: Retry the install with a longer timeout and backoff before declaring anything broken. A 5-minute timeout during verification is a network event, not a verdict on the PR: raise the fetch timeout, retry a few times, and only mark the PR broken after repeated failures with logs attached. Record timeouts as their own status, separate from real failures.

```text
npm install timed out after 5 minutes while the agent verified an upgrade
```

1. Re-run the install with a longer fetch timeout and a few retries configured on the package manager.
   Expected: transient slowness no longer kills the verification run.
2. Check the registry is reachable with a quick metadata request for the same package.
   Expected: confirms whether the problem is the registry, the local network, or the PR itself.
3. Retry the verification step with exponential backoff, up to three attempts.
   Expected: one slow attempt does not doom the PR; persistent slowness gets proven, not assumed.
4. Only after repeated failures, mark the PR broken, and attach the install log showing the timeouts.
   Expected: humans see "timed out 3 times" with evidence instead of a bare broken label.

## Use this when
- npm install timed out during upgrade verification
- The agent marked a PR broken after a single timeout with no retry
- A 5-minute install timeout ended the verification run
- "Marked the PR broken and moved on" after one slow install

## Not for this skill when
- The install fails fast with a real error (ERESOLVE, 404, checksum mismatch)
- The registry returns 429 (that needs backoff on rate limiting specifically)
- The timeout is on the whole CI job rather than the install step
- The failure reproduces instantly and deterministically

## Variant phrasings
- npm install timeout with no retry
- agent marked PR broken after timeout
- install timed out during verification
- npm fetch timeout on upgrade check

## Why it happens
The agent had no retry policy on its verification step, so it collapsed "the network was slow once" into "the upgrade is broken". Timeouts are the most transient failure class in the whole pipeline and they got the harshest verdict, which is exactly backwards.

## Edge cases
- Do not raise the timeout without bound; three attempts with backoff, then a human-visible timeout status.
- Cache registry tarballs between verification runs so retries are cheap and fast.
- If timeouts cluster across many PRs at once, that is a registry incident: pause verification instead of failing each PR individually.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_KirXjxzlPUtIQLUu0XYDtQ
