## TL;DR
Set `timeout` from measured upstream arrival times (p95 plus margin), pick `poke_interval` to match how fresh the data needs to be, and use `mode="reschedule"` for any wait longer than about 15 minutes so the sensor frees its worker slot between pokes. Most sensor pain comes from defaults: poke mode holding a slot for hours, or a timeout shorter than the upstream's bad days.

```text
Airflow sensor "timeout" tuning guide
```

## Use this when
- A sensor times out even though the data eventually arrives
- Sensors sit in poke mode for hours, starving the worker pool
- You are adding a new sensor and guessing at the parameters

## Not for this skill when
- The sensor never succeeds at all (check the poke logic and credentials first)
- You need general task retry policy (different knobs)
- You are deciding between a sensor and a triggered DAG (architecture question)

## Steps

1. Measure how long the upstream actually takes, using recent history:

```sql
SELECT percentile_cont(0.95) WITHIN GROUP (ORDER BY arrival_delay_minutes)
FROM upstream_arrivals WHERE arrival_date > CURRENT_DATE - INTERVAL '30 days';
```
Expected output: the p95 arrival delay in minutes. This number, not a guess, is what your timeout should be based on.

2. Set the timeout to p95 plus a comfortable margin:

```python
S3KeySensor(
    task_id="wait_for_file",
    timeout=6 * 3600,  # p95 was ~4h, margin to 6h
    ...
)
```
Expected output: the sensor survives the upstream's bad days without timing out, while still failing loudly if data is truly missing.

3. Choose the poke interval from freshness requirements, not from impatience:

```python
poke_interval=15 * 60,  # check every 15 minutes
```
Expected output: downstream starts within 15 minutes of data arrival. Shorter intervals find data faster but hammer the source API; 5 to 15 minutes is the sane range for most sensors.

4. Switch to reschedule mode for waits longer than about 15 minutes:

```python
mode="reschedule",
```
Expected output: between pokes the sensor releases its worker slot, so a 6-hour wait costs almost nothing in pool capacity. Poke mode would have held a slot the entire time.

5. Decide what a timeout means: hard failure or graceful skip:

```python
soft_fail=True,
```
Expected output: with soft_fail, a timeout marks the task skipped instead of failed, so downstream tasks with appropriate trigger rules can proceed. Use this when missing data is a normal, handleable condition rather than an incident.

## Variant phrasings

### airflow sensor timed out but file arrived later
Your timeout was shorter than the upstream's worst case. Recompute from p95 arrival times (step 1) and add margin; also check whether the file arrived late because of an upstream incident you should have been alerted about.

### sensor poke vs reschedule mode
Poke mode occupies a worker slot for the whole wait; reschedule mode frees it between pokes. Reschedule wins for long waits, poke is fine for short ones under ~15 minutes. Note reschedule creates a new task instance try per poke, which shows up in the UI.

### S3KeySensor timeout tuning
Same framework: measure the file arrival p95, set timeout above it, poke every 5-15 minutes, reschedule mode. For partitioned keys, sensor on the specific partition path for the logical date so a late unrelated file doesnt satisfy the check early.

## Why it happens
Sensors poll: they wake up, check a condition, and sleep. Two resources are at stake: worker slots (held in poke mode) and wall-clock patience (the timeout). Defaults are tuned for nobody in particular, so every sensor inherits a timeout and interval that fit some other team's upstream. Tuning is just aligning those two knobs with your actual upstream behavior and freshness needs.

## Edge cases
- In reschedule mode each poke is a new try, so `retries` interact oddly; keep retries low and let the timeout do the work.
- A poke_interval of seconds against an API will get you rate-limited; the source's limits are the real floor.
- Timeout counts from the first poke, and in older versions rescheduling didnt reset it; verify behavior on your Airflow version if a rescheduled sensor times out unexpectedly.
- Sensors with `timeout` longer than the DAG's schedule interval can overlap with the next run; consider `max_active_runs=1` or shorter timeouts.
- `soft_fail=True` plus `all_success` trigger rules downstream means a skipped sensor blocks everything; pair soft_fail with `none_failed` or similar where skipping is acceptable.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_0ouqZfbCuJAo29r6Tgec5Q
