Kubernetes job "DeadlineExceeded": backoffLimit vs activeDeadlineSeconds
Explains the Kubernetes Job DeadlineExceeded failure and the difference between backoffLimit (retry count) and activeDeadlineSeconds (total runtime cap). Use it when batch jobs fail with DeadlineExceeded or when tuning job retries. Not for CronJob schedule misses or pod-level timeouts.
Kubernetes Job DeadlineExceeded: which knob to turn
TL;DR
activeDeadlineSeconds caps the total runtime of the whole job; backoffLimit caps how many failed pod retries are allowed. A DeadlineExceeded job hit the time cap, so raise activeDeadlineSeconds if the work legitimately takes longer, or fix why attempts are slow. They control different axes and mixing them up wastes a debugging session.
Kubernetes job "DeadlineExceeded": backoffLimit vs activeDeadlineSecondsSteps
- Confirm which limit fired. Run
kubectl describe job [name]and read the events and the job spec.
Expected: events say DeadlineExceeded and you can see the activeDeadlineSeconds value and the job start time.
- Check whether attempts were progressing. Read the failed pods' logs.
Expected: you know if the work was slowly progressing or failing fast on the same error.
- If the work is legitimately slow, raise the deadline. You cannot edit a running job in place, so delete and recreate it with a larger
activeDeadlineSeconds(or update the CronJob template).
Expected: the new job completes within the new deadline.
- If attempts fail fast, look at backoffLimit and the underlying error. Raising backoffLimit only helps transient failures. Fix the actual error first.
Expected: either fewer premature failures or a fixed root cause.
- Set both deliberately on new jobs. Do not rely on defaults.
Expected: the job spec carries explicit values for both fields.
Use this when
- Jobs fail with a DeadlineExceeded condition
- Tuning retries for batch or ETL jobs
- Deciding between giving a job more time vs more attempts
Not for this skill when
- CronJobs miss their schedule (that is the scheduler or concurrency policy)
- Individual pods time out (use container or application timeouts)
- Jobs never start at all (that is scheduling or quota)
Variant phrasings
- kubernetes job timeout
- job deadline exceeded meaning
- backoffLimit explained
- activeDeadlineSeconds vs backoffLimit
Why it happens
The job controller runs two independent checks: a clock that starts when the job starts (activeDeadlineSeconds) and a counter of failed pods (backoffLimit). Hitting the clock kills all pods and marks the job failed no matter how many retries are left. Hitting the retry count fails the job no matter how much time is left.
Edge cases
- You cannot patch activeDeadlineSeconds on a running job. Recreate it.
- For CronJobs, edit the job template. The next scheduled run picks it up.
- A high backoffLimit with crashlooping pods burns node resources for nothing. Pair it with a sane deadline.
- Succeeded pods do not count against backoffLimit, only failed ones.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_VsG1wzKpLfzHdKF728XTMQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.