## TL;DR

Make the job token-aware: track the access token's expiry, refresh proactively before it lapses, and retry a 401 once with a fresh token. Persist the refresh token in durable storage so container restarts do not strand the job. If the refresh itself fails, alert instead of retry-looping.

```text
salesforce REST API 401 'expired access/refresh token' at 3am  -  agent's long job died mid-run
```

## Steps

1. On any 401, run the OAuth refresh flow and retry the failed request exactly once with the new access token. Expected: the retried request succeeds; the job continues instead of dying.

2. Track the token's issued-at and lifetime, and refresh proactively at around 80 percent of the lifetime. Expected: refreshes happen before expiry, so 401s become rare.

3. Store the refresh token in durable storage - a file or secret store - not in process memory. Expected: a container restart resumes with the stored refresh token intact.

4. If the refresh call itself fails, stop and alert a human; do not loop the failing refresh. Expected: a revoked grant pages someone instead of burning hours retrying.

## Use this when

- a job that runs for hours dies partway with 401 expired token
- the failure always lands at roughly the token lifetime after start
- restarts lose the session and the job cannot resume

## Not for this skill when

- 401 on the very first call - that is wrong credentials or a misconfigured connected app, not expiry
- a revoked refresh token that no retry can fix - re-authorize instead
- password or certificate expirations on the integration user

## Variant phrasings

### expired access/refresh token at 3am

Same root cause: the job outlived the token it started with.

### Salesforce session died mid-run

Same fix: proactive refresh plus retry-once on 401.

### token expired during a bulk job

Same fix: durable refresh token storage plus expiry tracking.

## Why it happens

Access tokens are short-lived by design - minutes to a few hours. A job that runs longer than the token's lifetime will eventually make a call with a dead token and get a 401. The agent fetched its token once at startup and assumed the session was permanent, so the failure always arrives hours in, at the worst possible moment, with most of the work already done.

## Edge cases

- Refresh token rotation: some orgs rotate the refresh token on use - always persist the newest one returned.
- IP restrictions or session policies can invalidate tokens independently of expiry - check those if refreshes keep failing.
- Clock skew between the agent host and Salesforce can make expiry math wrong - refresh early rather than exactly on time.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_yRWz0_CUuSzHrgQB8NP3PA
