## TL;DR
Give tasks a few retries with exponential backoff so transient failures heal themselves, and only alert when retries are exhausted or an SLA is missed. The fix is `retries` plus `retry_exponential_backoff` in default_args, with alert callbacks that check whether this was the final attempt. Alerting on every attempt trains the team to ignore alerts, which defeats the point.

```text
how to set retries and alerts that aren't noisy
```

## Use this when
- Every retried task sends an email or Slack message and the team mutes the channel
- You are setting the failure policy for a new DAG
- You want failures to page but flakes to stay quiet

## Not for this skill when
- You are debugging why a specific task fails (fix the task first)
- You need SLA-miss alerting specifically (related but separate callbacks)
- Sensor timeouts need tuning (different knobs)

## Steps

1. Set sane retry defaults on the DAG so transient flakes self-heal:

```python
default_args = {
    "retries": 3,
    "retry_delay": timedelta(minutes=5),
    "retry_exponential_backoff": True,
    "max_retry_delay": timedelta(minutes=60),
}
```
Expected output: a task that fails on a transient error waits 5, then 10, then 20 minutes before giving up, instead of failing loudly on the first hiccup.

2. Cap the backoff so a retry storm cant push a task into the next day:

```python
"max_retry_delay": timedelta(minutes=60),
```
Expected output: even with exponential growth, no single wait exceeds an hour. Tune this against your DAG's schedule interval.

3. Alert only on the final failure, not on every attempt, by checking the try number in the callback:

```python
def alert_on_final_failure(context):
    ti = context["ti"]
    if ti.try_number > ti.max_tries:
        send_alert(f"{ti.dag_id}.{ti.task_id} failed after all retries")
```
Expected output: one alert per genuinely failed task instead of one per attempt. The team starts reading alerts again.

4. Route by severity: page for final failures of critical tasks, send a daily digest for the rest:

```python
critical = MyOperator(task_id="load_fact", on_failure_callback=page_oncall, ...)
routine = MyOperator(task_id="refresh_staging", on_failure_callback=log_only, ...)
```
Expected output: the on-call phone stays quiet for low-stakes tasks while truly important failures still page.

5. Track the retry rate per task and fix chronic retry-ers instead of raising their retries:

```sql
SELECT task_id, COUNT(*) FILTER (WHERE try_number > 1) * 1.0 / COUNT(*) AS retry_rate
FROM task_instance
WHERE dag_id = '[dag id]' AND execution_date > CURRENT_DATE - INTERVAL '30 days'
GROUP BY task_id ORDER BY retry_rate DESC;
```
Expected output: a ranked list. Any task retrying more than about 10 percent of runs has a bug worth fixing; retries are a bandage, not a strategy.

## Variant phrasings

### airflow too many retry alert emails
You are alerting on attempts instead of outcomes. Move the alert into an on_failure callback gated on the final try, and the noise drops to one message per real failure.

### alert only after retries exhausted airflow
Check `ti.try_number` against `ti.max_tries` in `on_failure_callback`, or use SLA callbacks which fire independently of retries. Both patterns give you outcome-level alerting.

### how many retries should an airflow task have
Three retries with exponential backoff covers the vast majority of transient failures (network blips, brief upstream outages). More than five usually means the task or its dependency is broken and retries are just delaying the inevitable alert.

## Why it happens
Airflow's default alerting fires per failure event, and a retry is preceded by a failure event, so naive setups page on attempt one of a task that would have succeeded on attempt two. Meanwhile teams set retries high to paper over flaky tasks, multiplying the noise. The fix is conceptual: retries handle transient causes, alerts handle persistent ones, and the callback is where you separate the two.

## Edge cases
- `on_retry_callback` exists for logging retries quietly; use it for observability without paging.
- SLA miss callbacks fire even when the task eventually succeeds, so keep SLA alerting separate from failure alerting.
- `retries=0` on sensors with long timeouts can be correct; a sensor timing out is often the signal you want, not something to retry.
- Email alerting requires SMTP config on the workers; if alerts silently never arrive, check that before tuning anything else.
- A task with `depends_on_past` and retries can wedge the whole DAG; prefer idempotent tasks over long retry chains for ordering-sensitive pipelines.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_u7ZqFCV7xbtzctZBrnGhcA
