## TL;DR
Something outside Airflow killed the worker process, usually the kernel OOM killer or a node eviction, while the task instance still believed it was running. The fix is to find the kill reason in the worker or infrastructure logs, address it (usually more memory or eviction-proof scheduling), then clear the zombie task instance and rerun. The message is Airflow noticing the mismatch, not the cause.

```text
Executor reports task instance finished (failed) although the task says its running. Was the task killed externally?
```

## Use this when
- Task logs contain "was the task killed externally"
- A task shows as running forever but produces no new log lines
- This started after moving to KubernetesExecutor or after a cluster event

## Not for this skill when
- Tasks are stuck in queued, scheduled, or up_for_retry (scheduler-side issue)
- The task failed normally with a traceback in its logs
- The scheduler itself is unhealthy or not heartbeating

## Steps

1. Read the tail of the task log around the time it died:

```shell
airflow tasks logs my_dag my_task 2026-10-01 --log-level INFO | tail -50
```
Expected output: the last lines before death. An abrupt cutoff with no traceback points to an external kill; a Python traceback points to an in-task bug instead.

2. Check the infrastructure for the kill reason. On Kubernetes:

```shell
kubectl describe pod [worker pod] | grep -A5 "Last State"
```
Expected output: the last container state, often `Reason: OOMKilled` or `Reason: Evicted`. That line names your culprit.

3. If it was OOMKilled, either give the task more memory or make it use less:

```python
MyOperator(
    task_id="my_task",
    executor_config={"KubernetesExecutor": {"request_memory": "2Gi", "limit_memory": "4Gi"}},
)
```
Expected output: the pod spec requests enough memory that the kernel stops killing it. Alternatively chunk the work so peak memory drops.

4. Clear the zombie task instance so the DAG can move on:

```shell
airflow tasks clear my_dag -t my_task -s 2026-10-01 -e 2026-10-01 --yes
```
Expected output: the stuck "running" state is cleared. The task becomes eligible to run again on the next scheduler loop.

5. Harden against recurrence: set realistic memory requests on memory-hungry tasks, add retries for preemption-prone node pools, and alert on the zombie message pattern:

```python
default_args = {"retries": 2, "retry_delay": timedelta(minutes=10)}
```
Expected output: the occasional eviction becomes a quiet retry instead of a stuck DAG and a confused on-call.

## Variant phrasings

### was the task killed externally airflow
Yes, almost always. That message is the executor reporting that the process vanished without the task instance state changing. Start at the infrastructure layer, not the DAG code.

### airflow task says running but finished
The DB row says running because nothing updated it when the process died. Clearing the task instance (step 4) resets the state; the scheduler will not do it on its own.

### zombie task airflow kubernetes executor
The KubernetesExecutor is the most common source because pods get OOMKilled and evicted routinely. Memory requests and limits per task are the durable fix; the default pod spec fits nobody's workload.

## Why it happens
Airflow tracks task state in two places: the metadata DB row ("running") and the actual worker process. The executor reconciles them by checking whether the process is alive. When the kernel or orchestrator kills the process from outside, the DB row is never updated, so the next reconciliation finds a "running" task with no living process and reports the mismatch. It is a symptom of infrastructure pressure, not an Airflow bug.

## Edge cases
- With CeleryExecutor, a worker restart or SIGKILL produces the same symptom; check worker logs and the broker for disconnects.
- `airflow jobs check` can find scheduler and worker jobs that died without cleanup; run it if zombies recur.
- A task that forks subprocesses can leave orphans that hold resources after the kill; make sure cleanup handles child processes.
- Database locks or a full metadata disk can mimic zombie behavior by preventing state updates; check DB health if no kill reason appears anywhere.
- Marking the zombie as failed manually without clearing can confuse downstream trigger rules; prefer `tasks clear` so the task reruns cleanly.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_lEB9ikL2Wg3My8IAt9PHEw
