auth token expired mid-deploy: agent left worker in half-published failed state
Recovers a Worker left half-published when the API credential expired mid-deploy. Covers confirming the dead credential, re-authenticating, inspecting live state, and running one clean full redeploy. Use when deploy output shows auth errors partway through after earlier steps succeeded. Key trigger: 401 or expired-credential error in the middle of wrangler deploy output.
TL;DR: Stop retrying and re-authenticate first. A deploy is a sequence of API calls sharing one credential, so a token expiring mid-deploy leaves the worker half-published: new script with old bindings, or vice versa. Verify the credential is dead with wrangler whoami, get a fresh one, inspect what's live with wrangler deployments list, then run one clean full deploy. Retrying with the dead credential just burns rate limit.
auth token expired mid-deploy: agent left worker in half-published failed stateSteps
- Stop all retries immediately. Run
wrangler whoami. Expected: an auth error, confirming the credential is dead. No deploy attempts until this passes. - Get a fresh credential: create a new API credential with Workers write scope, or re-run the login flow. Run
wrangler whoamiagain. Expected: it shows the correct account with no errors. - Inspect the live state:
wrangler deployments list. Expected: you can see which deployment is live and its timestamp, so you know whether it predates the failed run. - Run one clean full deploy:
wrangler deploy. Expected: success. A full deploy replaces the script, bindings, and routes together, which clears the partial state. Don't try to resume by replaying only the failed sub-step. - Verify: check the deployments list again and hit the worker's route. Expected: the live version matches the deploy you just pushed, and the worker behaves as the new code intends.
- Add the guard: the agent checks credential validity with
wrangler whoamibefore every deploy, and treats any 401 mid-run as "stop and re-authenticate", never "retry the same call". Expected: this failure mode stops recurring.
Use this when
- deploy output shows an auth error partway through, after earlier steps succeeded
- the live worker's behavior doesn't match any single known deploy
- bindings or routes look stale while the script looks new (or vice versa)
- "half-published" or "partially deployed" state after a failed run
Not for this skill when
- the deploy failed on validation before anything published - that's a config error
- the credential was already dead before the deploy started - just re-authenticate, no recovery needed
- the failure is a 429 rate limit with a valid credential - back off instead
- the worker is fine and only the local state is confused - verify before "fixing"
Variant phrasings
- wrangler deploy died with an authentication error halfway through
- worker left in inconsistent state after expired token during deploy
- deploy partially applied then failed on auth
- half-published worker after credential expiry
Why it happens
Wrangler's deploy is not one atomic API call - it uploads the script, sets bindings, uploads assets, and updates routes as separate calls. A credential expiring between calls commits whatever already landed and fails the rest. Blind retries with the same dead credential can't converge and each attempt spends rate-limit budget.
Edge cases
- Secrets rotated between the failed deploy and recovery may belong to the half-state. Verify secrets are the intended values too.
- If another agent deployed concurrently, the deployments list shows which one won. Coordinate before redeploying over it.
- Service credentials with long lifetimes avoid expiry mid-deploy for scheduled agent runs.
- A deploy that failed at the route-update step can leave the new script live on old routes. Step 3 catches this.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst2ycAeC8IVF_aSdfqiCRLw