# TL;DR

Treat simultaneous 5xx failures across unrelated packages as a registry incident, not a resolution problem: stop, back off, and reschedule the run. A 404 means the version does not exist; a 500 means the registry is sick, and no amount of retrying in the same minute will fix it. The agent burned its retries against a broken registry and then wrote "unresolvable" next to twelve perfectly fine packages.

## Error

```text
agent's upgrade scheduled for off-peak ran during a registry incident; every retry hit the same 500 and it marked 12 packages unresolvable
```

## Steps

1. Recognize the incident pattern. Check the failures: if unrelated packages all return 500/502/503 at the same time, that is a registry problem, not a dependency problem. Expected: you confirm the failures are correlated in time and spread across unrelated packages.
2. Stop retrying immediately. Continuing to hammer a struggling registry makes it worse and wastes the run. Expected: the agent halts the upgrade loop instead of marking packages unresolvable.
3. Check the registry status. Look at the registry's status page or try one manual install. Expected: you confirm the outage (or confirm it is over).
4. Clear the bad labels. Remove the "unresolvable" marks the agent recorded for those packages. Expected: the packages are back in the candidate pool with no stigma attached.
5. Reschedule with a circuit breaker. Re-run the upgrades after the incident clears, and add a rule: N consecutive 5xx responses across unrelated packages pauses the run and reports "registry incident" instead of continuing. Expected: the next incident pauses the run cleanly instead of corrupting the data.

## Use this when

- Many unrelated packages fail at once with 500-level errors
- An agent labeled packages "unresolvable" during a registry outage
- You are adding retry/backoff logic to an upgrade agent
- Upgrade runs are scheduled during off-peak hours when nobody notices an outage

## Not for this skill when

- Failures are 404s or version-not-found errors - those really are unresolvable versions
- Only one package fails while others succeed - that is a package problem, not an incident
- The registry is rate-limiting you (429) - back off, but do not treat it as an outage

## Variant phrasings

- npm registry 500 errors during automated upgrade
- agent marked packages unresolvable after registry outage
- how to distinguish registry incident from bad version
- retry storm during npm outage

## Why it happens

Agents are built to keep going: a failed install triggers a retry, and a failed retry triggers a label. Nothing in that loop asks "are ALL packages failing right now," so a registry-wide 500 looks like twelve individual package problems. Off-peak scheduling makes it worse because no human is around to notice the status page is red.

## Edge cases

- Partial outages return 500 for some packages and 200 for others - the circuit breaker should trigger on the error rate, not require total failure.
- A 500 on one specific tarball URL while metadata requests succeed can mean a corrupted publish, not an outage - investigate single-package 500s individually.
- Do not let the circuit breaker stay tripped forever: it should reset after a cooldown so the next scheduled run proceeds normally.
- Log incident pauses distinctly from real unresolvable labels so future audits do not confuse the two.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_8VB3lK6a835V8PAV416K5g
