## TL;DR
Cron schedules guess when data is ready; Datasets react to it. A producer task declares an `outlet` Dataset (a URI like a table or file path), and the consumer DAG lists that Dataset in its `schedule=[...]` so it runs only when the dataset actually updates. This kills the "wait 2 hours just in case" buffer most teams bake into cron chains.

## The query
```
airflow dataset triggered dag scheduling
```

## Use this when
- a downstream DAG should run when an upstream task finishes writing data, not at a fixed time
- you are chaining DAGs and tired of cron offsets drifting
- multiple producers feed one consumer (schedule takes a list, or combine with & conditions)

## Not for
- simple time-based schedules (a cron expression is less machinery)
- triggering across separate Airflow instances (Datasets are local to one metastore)
- Airflow versions before 2.4 (Datasets did not exist there)

## Steps
1. Define the Dataset with a URI that describes the data, not the task. URIs are just strings, so pick a convention like `s3://bucket/table` or `snowflake://db/schema/table` and stick to it.
```python
from airflow.datasets import Dataset
raw_orders = Dataset("s3://data-lake/raw/orders")
```
Expected output: the Dataset shows up in the Airflow UI under Datasets.

2. Declare it as an `outlet` on the producer task so Airflow records an update event when the task succeeds.
```python
write_task = PythonOperator(
    task_id="write_orders",
    python_callable=write_orders,
    outlets=[raw_orders],
)
```
Expected output: each successful run creates a dataset event in the UI.

3. Set the consumer DAG's `schedule` to the Dataset (or a list of them) instead of a cron string.
```python
with DAG(dag_id="transform_orders", schedule=[raw_orders], ...):
    ...
```
Expected output: the consumer DAG creates a run only after the producer's dataset event lands.

4. For multiple inputs, pass a list for "any of them" semantics, or combine with `&` for "all of them" (Airflow 2.9+).
Expected output: runs trigger only when the combined condition is satisfied.

5. Check the Datasets view in the UI to confirm events are flowing, and confirm the consumer's run history lines up with producer completions.
Expected output: dataset events and consumer DAG runs match one to one, with no cron-driven empty runs.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_DVXM7Wximyw27nBIpPFH6Q
