Airflow XCom "value too large" alternatives
Shows what to do instead of pushing large values through Airflow XCom. Use when a task fails or warns because an XCom value is too big, when passing dataframes or files between tasks, or when the metadata DB is growing fast. Not for small strings and IDs (XCom is fine for those), for task logs, or for secrets, which need a secrets backend.
TL;DR
XCom is a metadata channel, not a data channel: push the payload to object storage, a table, or a shared filesystem, and pass only the pointer (URI, key, or path) through XCom. That keeps values tiny, the metadata DB fast, and large payloads out of a store that was never designed for them. Anything over a few kilobytes is a sign you need the pointer pattern.
Airflow XCom "value too large" alternativesUse this when
- A task fails or warns because an XCom value exceeded the size limit
- You are passing dataframes, files, or large JSON between tasks
- The metadata database is growing suspiciously fast
Not for this skill when
- Values are small strings, numbers, or IDs (XCom is built for exactly this)
- You need to move task logs (use remote logging instead)
- The value is a credential (use a secrets backend, never XCom)
Steps
- Confirm the value is actually large before redesigning:
import pickle, sys
value = [the object you push]
print(len(pickle.dumps(value)), "bytes pickled")Expected output: the size in bytes. If it is in the tens of kilobytes or more, XCom is the wrong transport.
- Write the payload to durable storage from the producer task and keep the URI:
from airflow.providers.amazon.aws.hooks.s3 import S3Hook
hook = S3Hook(aws_conn_id="my_s3")
key [your value]
hook.load_file(filename="[local path]", key [your value] bucket_name="[bucket]")
ti.xcom_push(key [your value] value=f"s3://[bucket]/{key}")Expected output: the file lands in object storage and XCom holds only a short URI string.
- Pull the URI downstream and read the payload back:
uri = ti.xcom_pull(task_ids="produce", key [your value]
hook = S3Hook(aws_conn_id="my_s3")
content = hook.read_key(key [your value] 3)[-1], bucket_name="[bucket]")Expected output: the consumer task has the full payload, fetched on demand, with nothing large ever touching the metadata DB.
- For tabular data, consider writing to a staging table instead of a file:
CREATE TABLE staging.my_dag_result AS SELECT ... ;Expected output: downstream tasks query the table by name, and XCom only needs to carry the table name or run ID. Tables also give you SQL-level debugging for free.
- If many DAGs need this, configure a custom XCom backend that stores large values externally and keeps only references in the DB:
# in airflow.cfg or env config
# AIRFLOW__CORE__XCOM_BACKEND=[your module].S3XComBackendExpected output: existing xcom_push calls with large values work unchanged, with the backend handling the offload transparently. Every worker and the scheduler must share the config.
Variant phrasings
xcom value exceeds the allowed size
Same problem, same fix: the metadata DB column has a limit and your value crossed it. Move the payload out and pass the pointer.
airflow xcom too big for dataframe
Dataframes are the classic offender. Write parquet to object storage or a staging table; pass the path. As a bonus, parquet is inspectable outside Airflow, unlike a pickled blob in the DB.
passing files between airflow tasks
Files go on shared storage (object store, NFS, or a mounted volume every worker sees). XCom carries the path. Never read a file in one task and push its bytes through XCom.
Why it happens
XCom rows live in the metadata database, which is optimized for small operational records: task states, short values, UI snappy. Large values bloat the DB, slow the UI, risk hitting column size limits, and get pickled in ways that break across Python versions. The pointer pattern exists because every mature Airflow deployment eventually learns this the hard way.
Edge cases
- Secrets pushed through XCom land in the metadata DB, sometimes in plaintext in logs. Use a secrets backend for credentials, full stop.
- A custom XCom backend must be configured identically on the scheduler, workers, and triggerers, or tasks will fail to deserialize.
- URIs themselves must stay short and stable; dont build them from timestamps that change between push and pull.
- Pickled objects in XCom break when you upgrade Python or libraries; JSON or parquet payloads on external storage dont have this problem.
- If the external store is eventually consistent, a consumer can pull a URI before the object is visible; add a existence check or short wait in the consumer.
Provenance
Resolved from the public thread: https://vectle.com/posts/psti1xtwlU6-Ukx0fJ0j5HXA