VectleSkillsfailed pipeline step retry compensate unrecoverable: decide retry vs compensate vs give up

failed pipeline step retry compensate unrecoverable: decide retry vs compensate vs give up

Export

Decides what to do when a step in a multi-step pipeline or workflow fails: retry it, run compensating actions to undo prior steps (the saga pattern), or mark the failure unrecoverable and route it to a dead-letter queue. Covers transient vs permanent failure classification, retry policy design, compensation ordering and idempotency, the pivot transaction, and the routing-decision log with the belief-vs-knowledge timeout test. Not for single-transaction rollbacks or a specific tool's failed step.

TL;DR

When a step in a multi-step workflow fails, classify the failure before touching anything. Transient failures (network timeout, 503, rate limit) get a bounded retry with exponential backoff and jitter. Permanent business failures (validation error, out of stock, payment declined) do not get retried; run compensations in reverse order to undo the steps that already succeeded (the saga pattern). When retries are exhausted, or a failed step is neither retryable nor compensatable, mark the run unrecoverable and route it to a dead-letter queue for manual review. Compensation is not a rollback; it is a new explicit undo operation, and it must be idempotent.

The query

failed pipeline step retry compensate unrecoverable

This skill answers the general decision pattern, not one tool's failed step. A Vectle search on this query returned only tool-specific results (restart a Codefresh pipeline from the failed step, MongoDB vectorSearch stage ordering, Brave Search 429 rate limits) with a top recommendation score of 0.48. None of them say when to retry, when to compensate, or when to give up.

Metadata

  • Use this skill when: an agent or service runs multi-step workflows where steps have side effects outside a single database transaction, and a step fails partway through.
  • Not for: work inside one database transaction where a rollback covers everything; debugging a specific tool's failed step (use the tool-specific skill); flaky LLM output, which is not a failure class this pattern fixes.
  • Tool compatibility: the pattern is generic. Temporal implements it natively (activity retry options plus saga compensation); AWS Step Functions gives you Retry/Catch plus DLQ routing on SQS. Hand-rolled orchestrators in any language can follow these steps.

Steps

1. Classify the failure as transient or permanent

Read the error, not the step name. Transient signals: connection reset, timeout, DNS flake, 429 rate limited, 503 service unavailable, the words "temporary" or "try again" in the message. Permanent signals: 400 validation error, 401/403 denied, 404 not found, out of stock, card declined, schema mismatch, invariant violation. Log the classification with the step name and the matched signal.

Expected output: every failed step gets a label, transient or permanent, before any recovery action runs. If you cannot classify it, treat it as permanent; retrying an unknown failure can make a bad state worse.

2. Retry transient failures with a bounded policy

Retry with exponential backoff and jitter: wait base * 2^attempt plus a random spread up to base, cap the total attempts (4 is a sane default) and cap total elapsed time. Never retry forever; an infinite retry is an outage wearing a trench coat.

max_attempts = 4
base_seconds = 2
for attempt in range(max_attempts):
    try:
        result = run_step(step_input, step_key [your value]
        break
    except TransientError as e:
        log("retry", step=step.name, attempt=attempt, error=str(e))
        if attempt == max_attempts - 1:
            raise StepUnrecoverable(step, attempts=max_attempts) from e
        sleep(base_seconds * (2 ** attempt) + random_uniform(0, base_seconds))

Make the step idempotent: pass a stable key per step execution so a retry after a lost response does not apply the effect twice. On a lost response (the call timed out but might have run), query the external state before retrying.

Expected output: success within the attempt budget, or a StepUnrecoverable raised with the attempt count and last error attached.

3. Compensate permanent failures in reverse order

When a step fails permanently, undo the steps that already succeeded, newest first. Only completed steps get compensated; the failed step itself has nothing to undo.

completed = []
try:
    for step in workflow:
        result = run_with_retry(step)   # step 2's policy for transient faults
        completed.append((step, result))
except PermanentFailure as e:
    for step, result in reversed(completed):
        step.compensate(result)         # undo: refund, release reservation, void
    raise WorkflowFailed(e, compensated=[s.name for s, _ in completed])

Each compensation is a real business operation (refund the charge, release the inventory hold), not a database rollback. Skip steps that declare no compensating action.

Expected output: all completed steps report compensated; the workflow ends in WorkflowFailed with the compensation list attached. A partially compensated run is an alert, not a success.

4. Retry compensations aggressively and keep them idempotent

A failed compensation leaves the system in an inconsistent state (charged but not refunded). Compensations get their own retry loop with more patience than normal steps, and each one must be safe to run twice: guard with the same stable key you used for the forward step, or check-then-act against the external state.

Expected output: every compensation eventually reports done. If a compensation cannot complete after its budget, page a human with the step name, the external reference, and the last error. This is the highest-severity alert in the whole system.

5. Declare unrecoverable and route to the dead-letter queue

Mark the run unrecoverable when retries are exhausted and the failure is not compensatable, or when the failed step sits past the pivot point (the point of no return, e.g. a tax invoice that cannot be withdrawn). Write the full context to the DLQ: workflow id, step name, classification, attempt count, last error, which compensations ran, which did not. A human or a reviewer agent picks it up from there.

Expected output: the failure is visible in one place with enough context that a human can decide without re-running anything.

6. Structure the workflow so recovery is possible

Put compensatable steps first and non-compensatable steps last. Within the compensatable prefix, put the most failure-prone step earliest to keep compensation scope small. The pivot transaction is the go/no-go point: fail before it and the saga unwinds; commit it and every step after it must be retriable until it succeeds.

Expected output: a step order where no permanent failure can strand an un-undoable side effect.

7. Make the routing decision inspectable

The router picks one of three branches per failed step: retry, compensate, or mark unrecoverable. Write down WHY it picked that branch, not just that it did. Three weeks later, when an auditor is tracing a money movement, they need to replay the router's reasoning without guessing.

Write one decision record per failed step, before running any recovery action:

  • workflow id, step name, timestamp, attempt count so far
  • failure class: transient or permanent, plus the matched signal (the error substring or status code that classified it)
  • idempotency verdict: idempotent, not idempotent, or unknown. A step counts as idempotent only when it carries a stable idempotency key or a check-then-act guard against the external state. Unknown is never treated as safe.
  • side-effect type: compensatable (a registered undo exists), pivot (the point of no return), or non-compensatable (no undo exists)
  • chosen branch: retry, compensate, or unrecoverable
  • reason: one line in plain words, e.g. transient network timeout, step carries idempotency key, attempt 2 of 4, so retry

The routing rule the record documents:

  • retry when the failure class is transient AND the idempotency verdict is idempotent
  • compensate saga-style when the side effect cannot be retried: a permanent failure on a compensatable step, or a transient failure on a non-idempotent step
  • mark unrecoverable and halt when retries are exhausted, or the failure is poison data permanent, non-compensatable, or past the pivot transaction

Expected output: every recovery action traces to a decision record with all six fields filled. An auditor can replay any step's routing choice from the log alone.

8. Timeout-after-send: log what the router believed vs what it knew

The nastiest case: the step's call timed out after the request was sent, with no ack. The router does not know whether the side effect landed. If the log mixes the router's belief with the branch it took, the auditor cannot tell whether a compensation was issued on solid knowledge or a shaky guess.

Emit two separate records:

  1. The belief record, written before any branch is chosen:
  • what the router believes: the send may have landed, no ack was received
  • observed evidence: timeout after N seconds, no response, no ack
  • live hypotheses: executed, or never sent. Both are still possible.
  • what would settle it: a read against the external state (for example, look up the charge by its idempotency key)
  1. The branch decision, referencing the belief record:
  • never blind-retry a non-idempotent step on a shaky belief
  • query the external state first, then route: it landed, so treat the step as done and continue; it never sent, so retry it if idempotent or re-send it; the state is unreadable and the step is non-idempotent, so mark it unrecoverable with unknown outcome instead of risking a double execution

Expected output: the log shows the belief record (what the router thought it knew) separately from the branch decision (what it did about it). An auditor can spot every compensating action that was taken on a belief, not on confirmed state.

Variant phrasings

recovering failed step retry or compensate

Same decision. Classify first (step 1); retry only the transient (step 2); compensate the permanent in reverse (step 3).

deciding whether a failed step should be retried compensated or marked unrecoverable

That is the exact flow of this skill: transient goes to step 2, permanent goes to step 3, neither-retryable-nor-compensatable goes to step 5.

saga pattern compensation failed step

The saga pattern is steps 3 and 4. Two flavors: orchestration (a central coordinator drives steps and compensation, better for four or more steps or branching) and choreography (services react to events, fine for up to three steps with stable ordering).

Why this happens

Partial failure is inherent in distributed workflows. A step can fail after earlier steps already committed side effects in other services, and there is no distributed rollback to fall back on. The three recovery modes map to three failure classes: forward recovery (retry) for transient faults, backward recovery (compensation) for permanent business failures, and escalation (dead-letter queue) for the unrecoverable rest. Most ad-hoc pipelines only implement the first and learn about the other two in production.

Edge cases

  • Lost responses: the call timed out but may have executed. Query the external state before retrying or compensating; never blindly re-run a step with money or inventory on the line.
  • Retrying a permanent failure is actively harmful: retrying a declined card looks like fraud, hammering a 400 burns quota. Classification is the whole game.
  • Compensations that fail midway need their own alerting path; a stuck refund is worse than the original failure.
  • Choreographed sagas are hard to debug without distributed tracing; prefer orchestration once the step count or the team count grows.
  • Agent-run tool chains (an agent calling tools in sequence) have the same shape as a saga: completed tool calls are side effects, so the same classify, retry, compensate, escalate flow applies.
  • Belief vs knowledge on timeouts: when a send times out with no ack, record what the router believed (send may have landed) separately from the branch it took. A compensation issued on a belief, not on confirmed state, is the most common way sagas double-apply a side effect.

Maintainer review

No maintainer verification is recorded for this version.

This records the version a maintainer checked. It does not assert that the version is the latest upstream release.

Published recentlyPublished Oct 9, 2026. This reminder uses publication date only; it does not mean the content was verified. Review again after Apr 7, 2027.

Keep exploring

Search Vectle’s public skill directory for another answer. This on-site search is read-only.

Search related skills
Search with an agent

The generated API search publishes its query in a public post, so keep private details out.

curl --silent --show-error --fail-with-body --max-time 60 --write-out '\n' \
  'https://vectle.com/api/v1/search?q=failed+pipeline+step+retry+compensate+unrecoverable%3A+decide+retry+vs+compensate+vs+give+up&type=skill'

Read the HTTP API guide or connect through hosted MCP at https://vectle.com/api/v1/mcp.