executor reports task instance finished" zombie tasks
Diagnoses Airflow zombie tasks where the executor says the task finished but the task thinks it is still running. Use when you see the 'was the task killed externally' message, when tasks show as running forever with no logs, or after worker evictions and OOM kills. Not for tasks stuck in queued or scheduled states, for normal task failures with logs, or for scheduler heartbeat issues.
TL;DR
Something outside Airflow killed the worker process, usually the kernel OOM killer or a node eviction, while the task instance still believed it was running. The fix is to find the kill reason in the worker or infrastructure logs, address it (usually more memory or eviction-proof scheduling), then clear the zombie task instance and rerun. The message is Airflow noticing the mismatch, not the cause.
Executor reports task instance finished (failed) although the task says its running. Was the task killed externally?Use this when
- Task logs contain "was the task killed externally"
- A task shows as running forever but produces no new log lines
- This started after moving to KubernetesExecutor or after a cluster event
Not for this skill when
- Tasks are stuck in queued, scheduled, or upforretry (scheduler-side issue)
- The task failed normally with a traceback in its logs
- The scheduler itself is unhealthy or not heartbeating
Steps
- Read the tail of the task log around the time it died:
airflow tasks logs my_dag my_task 2026-10-01 --log-level INFO | tail -50Expected output: the last lines before death. An abrupt cutoff with no traceback points to an external kill; a Python traceback points to an in-task bug instead.
- Check the infrastructure for the kill reason. On Kubernetes:
kubectl describe pod [worker pod] | grep -A5 "Last State"Expected output: the last container state, often Reason: OOMKilled or Reason: Evicted. That line names your culprit.
- If it was OOMKilled, either give the task more memory or make it use less:
MyOperator(
task_id="my_task",
executor_config={"KubernetesExecutor": {"request_memory": "2Gi", "limit_memory": "4Gi"}},
)Expected output: the pod spec requests enough memory that the kernel stops killing it. Alternatively chunk the work so peak memory drops.
- Clear the zombie task instance so the DAG can move on:
airflow tasks clear my_dag -t my_task -s 2026-10-01 -e 2026-10-01 --yesExpected output: the stuck "running" state is cleared. The task becomes eligible to run again on the next scheduler loop.
- Harden against recurrence: set realistic memory requests on memory-hungry tasks, add retries for preemption-prone node pools, and alert on the zombie message pattern:
default_args = {"retries": 2, "retry_delay": timedelta(minutes=10)}Expected output: the occasional eviction becomes a quiet retry instead of a stuck DAG and a confused on-call.
Variant phrasings
was the task killed externally airflow
Yes, almost always. That message is the executor reporting that the process vanished without the task instance state changing. Start at the infrastructure layer, not the DAG code.
airflow task says running but finished
The DB row says running because nothing updated it when the process died. Clearing the task instance (step 4) resets the state; the scheduler will not do it on its own.
zombie task airflow kubernetes executor
The KubernetesExecutor is the most common source because pods get OOMKilled and evicted routinely. Memory requests and limits per task are the durable fix; the default pod spec fits nobody's workload.
Why it happens
Airflow tracks task state in two places: the metadata DB row ("running") and the actual worker process. The executor reconciles them by checking whether the process is alive. When the kernel or orchestrator kills the process from outside, the DB row is never updated, so the next reconciliation finds a "running" task with no living process and reports the mismatch. It is a symptom of infrastructure pressure, not an Airflow bug.
Edge cases
- With CeleryExecutor, a worker restart or SIGKILL produces the same symptom; check worker logs and the broker for disconnects.
airflow jobs checkcan find scheduler and worker jobs that died without cleanup; run it if zombies recur.- A task that forks subprocesses can leave orphans that hold resources after the kill; make sure cleanup handles child processes.
- Database locks or a full metadata disk can mimic zombie behavior by preventing state updates; check DB health if no kill reason appears anywhere.
- Marking the zombie as failed manually without clearing can confuse downstream trigger rules; prefer
tasks clearso the task reruns cleanly.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_lEB9ikL2Wg3My8IAt9PHEw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.