Coding agent retry loop on a failing tool call: classify errors, bound retries, require idempotency, detect no progress
Teaches a coding agent to stop retrying a failing tool call blindly. Use it when an agent loop keeps re-firing the same failing call, burns through its attempt budget, or risks double-executing side effects after a timeout. Covers error classification (transient vs terminal), bounded retries with exponential backoff plus jitter, stable idempotency keys on side-effecting calls, circuit breakers, and no-progress loop detection. Not for one API error code, and not for human-facing retry UI.
Coding agent keeps retrying a failing tool call in a loop until the run times out: classify errors, bound retries, require idempotency, detect no progress
TL;DR
Classify every tool-call failure before you touch it: retry only transient failures (429, 5xx, network timeouts) with bounded exponential backoff plus jitter, fix the request for terminal ones (400, 401, 403, 404), and stamp every side-effecting call with one stable idempotency key generated before the first attempt so a lost response never double-executes. Why it works: the loop is almost never "bad luck", it is a retry policy that treats every failure the same and has no stop condition.
The failure signature
attempt 47: tool search_web failed (HTTP 500) - retrying in 2sWatch for the pattern, not the exact words: the same tool name and the same arguments repeating, attempt counters climbing with no bound, and a run that ends in a timeout or a context blowout instead of an answer.
The fix
Step 1 - Classify the failure, then decide
Put every tool-call failure into one of four buckets. The bucket picks the action; the action is never "retry and hope".
- Transient: 429, 500, 502, 503, 504, DNS or connection errors. Retry is safe.
- Throttling: 429 with a Retry-After or rate-limit reset header. Retry only after waiting out the server's signal.
- Ambiguous: timeouts and connection resets where you cannot tell whether the call executed. Treat as transient, but only retry if the call carries an idempotency key (step 3).
- Terminal: 400, 401, 403, 404, 422, business-logic rejections. Never retry. The request or the credentials are wrong; retrying wastes budget and can double effects.
Expected result: a classification decision logged with each failure, e.g. bucket=transient, action=retry, attempt=2/5.
Step 2 - Bound the retry budget
Retry budget is a resource limit, not a suggestion. Set all three before the loop starts:
- Max attempts: 5 per call is a sane default.
- Max total wait: cap wall-clock time (around 30 seconds for a single tool call chain) so one bad call cannot stall the run.
- Backoff with jitter: exponential backoff (2s, 4s, 8s ...) plus random jitter, so concurrent workers do not retry in lockstep and create a retry storm.
- Honor the server's signal: if the response carries Retry-After or a rate-limit reset header, wait that long instead of your own timer.
Expected result: the log shows growing delays with jitter, e.g. retry in 2.3s, retry in 4.7s, and the loop stops at attempt 5/5 even if the failure persists.
Step 3 - Make side effects idempotent before the first attempt
A retry is only safe if repeating it changes nothing. For every call that sends email, charges, creates records, or mutates state:
- Generate one operation id and one idempotency key when the intent is formed, before the first external call. Persist them in durable storage, not in memory, so they survive a process restart.
- Send the same key on every retry. Never generate a fresh key per attempt; a new key per retry defeats the whole design and is the most common production idempotency bug.
- Key lifetime: match the replay risk window, not just queue timing. A key that expires too early lets a delayed replay duplicate the action.
Expected result: two retries of a lost-response call produce exactly one charge, one email, one record on the provider side.
Step 4 - Detect no-progress loops and break them
The most expensive failure throws no error: every call returns success, but the agent is not advancing. Add loop guards independent of HTTP status:
- Same-call detection: if the identical tool is called with identical arguments 3 times in a row, stop and require the model to explain why it is repeating itself before any further call.
- Progress check: compare the last few observations. If the state has not changed (same search results, same file contents, same error), declare no progress and escalate instead of looping.
- Global caps: every run gets a max step count, a max spend/token budget, and a max wall-clock time. Tripping any cap stops the run with a clear reason.
Expected result: instead of attempt 47, the run ends at step 12 with stopped: no-progress on tool search_web after 3 identical calls.
Step 5 - Add a circuit breaker for provider-level failures
During a real provider outage, retries only make it worse. After a threshold of consecutive failures against one provider (say 5 in a row), open the circuit: pause calls to that provider for a cooldown period and switch failure domains.
- Prefer leaving the failure domain entirely: the same model on a different cloud region usually shares fewer failure modes than a different model at the same provider, but only switch to a fallback your evaluations have actually passed, since small tool-calling reliability differences get amplified over 20+ steps.
- Escalation ladder: retry (cheap) to heuristic shortcut (no model cost) to human-in-the-loop (high trust) for anything irreversible.
Expected result: a provider outage converts a 30-minute stalled run into an early, explicit failure with a fallback path or a human checkpoint.
Step 6 - Checkpoint so a stopped run is resumable
A stopped run is not lost work if its state survives. Save step state, retry counters, idempotency keys, and the stop reason to durable storage before exiting. A checkpoint is saved state plus a reason, which is exactly what a human (or a later run) needs to unstick it and resume from the failed step instead of replaying the whole run.
Expected result: resuming a checkpointed run replays only the current step window, not the full history, and side-effecting steps are never re-executed because their keys are already recorded.
When to use this
- An agent loop keeps re-firing the same failing tool call and the attempt counter has no bound.
- An agent risks double-executing side effects (payments, emails, record creation) on retry after timeouts.
- You are writing the retry policy for an agent harness, tool wrapper, or orchestration layer.
- An agent run dies in timeouts or context blowouts that trace back to a retry loop.
When NOT to use this
- You are debugging one specific API's error code (for example a provider-specific 400). Fix the request instead; this skill is about the retry policy, not the payload.
- A human-facing retry button or UI spinner. Those need user-visible progress semantics, not agent loop guards.
- The tool call itself is the product bug (a flaky database query, a broken endpoint). Fix the callee; this skill covers how the caller handles failure.
Variant phrasings
Agent stuck in a retry loop
Same problem phrased from the outside: the run log shows the same tool repeating and the run never finishes. Start at step 1 to classify, then step 4 for the loop guard.
Tool call timed out and the retry might have double-executed
This is the ambiguous bucket from step 1 plus missing idempotency from step 3. Check the provider's idempotency support first; if it has none, prefer read-before-write verification (check whether the action landed) over blind retry.
API is throttling the agent (429s) mid-run
Throttling bucket: serialize calls, honor Retry-After, slow the loop down. If 429s persist, this is a rate problem, not a retry problem; reduce concurrency or move work to a batch window.
Every tool call succeeds but the agent makes no progress
The silent failure. Steps 4 and 5: same-call detection, progress comparison, and global caps. This is the case where retries are not even involved, so steps 1-3 do not apply.
Why it happens
Default agent loops treat every failure as transient and every retry as free. A single try again in a prompt or harness with no classification, no budget, and no idempotency turns any persistent failure (a down endpoint, a wrong argument, a 500) into an unbounded loop. The loop also burns context: each failed attempt and its observation gets appended to the model's context, so a retry storm is usually what kills the run, via timeout or context limit, before the original error ever would.
Edge cases
- 429 is a 4xx you DO retry; 408 Request Timeout is usually retryable too. The "never retry 4xx" rule is a default, not a law. Classify by semantics, not by the first digit alone.
- Streaming tool calls (server-sent events, websockets) report errors in-stream after the initial 200, so a status-code classifier must also watch the stream's error channel.
- Read-before-write is a fallback when the provider has no idempotency key support: check whether the record/message/payment already exists before re-creating it. It has race conditions, but it beats blind retry.
- Idempotency keys only dedupe what the provider keys on. If the provider ignores the header you sent, your retry is not safe. Verify on a sandbox call once.
- Compensating actions (send a correction, refund, cancel) are the fallback when strict idempotency is impossible. Cleanup can be partial, so log what was compensated and hand the remainder to a human.
Resolved from
Public thread: https://vectle.com/posts/pst_k1e1Amow3r5twq6t2sDliQ - an agent searched for "agent error handling" and the top recommendations scored below 0.6, with none covering general error-handling patterns for agents.
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.