Airflow sensor "timeout" tuning guide
Guide to tuning Airflow sensor timeout, poke interval, and execution mode. Use when sensors time out before data arrives, when sensors hog worker slots for hours, or when deciding between poke and reschedule mode. Not for general task retries, for debugging a sensor that never succeeds, or for choosing whether to use a sensor at all.
TL;DR
Set timeout from measured upstream arrival times (p95 plus margin), pick poke_interval to match how fresh the data needs to be, and use mode="reschedule" for any wait longer than about 15 minutes so the sensor frees its worker slot between pokes. Most sensor pain comes from defaults: poke mode holding a slot for hours, or a timeout shorter than the upstream's bad days.
Airflow sensor "timeout" tuning guideUse this when
- A sensor times out even though the data eventually arrives
- Sensors sit in poke mode for hours, starving the worker pool
- You are adding a new sensor and guessing at the parameters
Not for this skill when
- The sensor never succeeds at all (check the poke logic and credentials first)
- You need general task retry policy (different knobs)
- You are deciding between a sensor and a triggered DAG (architecture question)
Steps
- Measure how long the upstream actually takes, using recent history:
SELECT percentile_cont(0.95) WITHIN GROUP (ORDER BY arrival_delay_minutes)
FROM upstream_arrivals WHERE arrival_date > CURRENT_DATE - INTERVAL '30 days';Expected output: the p95 arrival delay in minutes. This number, not a guess, is what your timeout should be based on.
- Set the timeout to p95 plus a comfortable margin:
S3KeySensor(
task_id="wait_for_file",
timeout=6 * 3600, # p95 was ~4h, margin to 6h
...
)Expected output: the sensor survives the upstream's bad days without timing out, while still failing loudly if data is truly missing.
- Choose the poke interval from freshness requirements, not from impatience:
poke_interval=15 * 60, # check every 15 minutesExpected output: downstream starts within 15 minutes of data arrival. Shorter intervals find data faster but hammer the source API; 5 to 15 minutes is the sane range for most sensors.
- Switch to reschedule mode for waits longer than about 15 minutes:
mode="reschedule",Expected output: between pokes the sensor releases its worker slot, so a 6-hour wait costs almost nothing in pool capacity. Poke mode would have held a slot the entire time.
- Decide what a timeout means: hard failure or graceful skip:
soft_fail=True,Expected output: with soft_fail, a timeout marks the task skipped instead of failed, so downstream tasks with appropriate trigger rules can proceed. Use this when missing data is a normal, handleable condition rather than an incident.
Variant phrasings
airflow sensor timed out but file arrived later
Your timeout was shorter than the upstream's worst case. Recompute from p95 arrival times (step 1) and add margin; also check whether the file arrived late because of an upstream incident you should have been alerted about.
sensor poke vs reschedule mode
Poke mode occupies a worker slot for the whole wait; reschedule mode frees it between pokes. Reschedule wins for long waits, poke is fine for short ones under ~15 minutes. Note reschedule creates a new task instance try per poke, which shows up in the UI.
S3KeySensor timeout tuning
Same framework: measure the file arrival p95, set timeout above it, poke every 5-15 minutes, reschedule mode. For partitioned keys, sensor on the specific partition path for the logical date so a late unrelated file doesnt satisfy the check early.
Why it happens
Sensors poll: they wake up, check a condition, and sleep. Two resources are at stake: worker slots (held in poke mode) and wall-clock patience (the timeout). Defaults are tuned for nobody in particular, so every sensor inherits a timeout and interval that fit some other team's upstream. Tuning is just aligning those two knobs with your actual upstream behavior and freshness needs.
Edge cases
- In reschedule mode each poke is a new try, so
retriesinteract oddly; keep retries low and let the timeout do the work. - A poke_interval of seconds against an API will get you rate-limited; the source's limits are the real floor.
- Timeout counts from the first poke, and in older versions rescheduling didnt reset it; verify behavior on your Airflow version if a rescheduled sensor times out unexpectedly.
- Sensors with
timeoutlonger than the DAG's schedule interval can overlap with the next run; considermax_active_runs=1or shorter timeouts. soft_fail=Trueplusall_successtrigger rules downstream means a skipped sensor blocks everything; pair softfail with `nonefailed` or similar where skipping is acceptable.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_0ouqZfbCuJAo29r6Tgec5Q
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.