pandas explode list column into rows
Explodes list-valued pandas columns into one row per element. Use when a column holds lists or arrays and you need a long-format frame, when explode produces NaN rows you did not expect, or when you need to keep the other columns aligned. Not for splitting delimited strings (split first), for pivoting, or for asof timestamp matching.
TL;DR
Call df.explode("col") to turn each list element into its own row, after making sure empty lists and non-list values are handled the way you want. Explode keeps every other column aligned automatically and resets nothing, so fix the index after.
pandas explode list column into rowsUse this when
- A column contains lists and you need one row per element
- You want long-format data for grouping or joining
- Explode output has surprise NaN rows or a duplicated index
Not for this skill when
- The column holds delimited strings like "a,b,c", split to lists first
- You need to pivot rows into columns
- You need nearest-timestamp matching between series
Steps
- Inspect what is actually in the column before exploding:
print(df["tags"].apply(type).value_counts())
print(df["tags"].apply(lambda x: len(x) if isinstance(x, list) else None).describe())Expected output: the mix of types (lists, None, maybe stray strings) and the distribution of list lengths. Non-list values and empty lists are what cause surprises.
- Normalize non-list values so explode behaves predictably:
df["tags"] = df["tags"].apply(lambda x: x if isinstance(x, list) else [])Expected output: every cell is a list. Without this, scalar values explode into one row each (fine) but None becomes NaN (often not what you want).
- Explode and reset the index, since explode duplicates index labels:
long = df.explode("tags").reset_index(drop=True)
print(long.head())
print(len(long))Expected output: one row per list element, other columns repeated, and a clean 0..n index. The row count equals the total number of elements across all lists.
- Decide what empty lists should become. By default they produce one NaN row:
# keep the NaN rows (marks "had no tags")
long = df.explode("tags")
# or drop them
long = df.explode("tags").dropna(subset=["tags"])Expected output: either NaN placeholder rows or a frame with only real elements. Pick deliberately, silent NaN rows corrupt counts.
- Explode multiple list columns together only when the lists align element-wise:
long = df.explode(["tags", "scores"])Expected output: paired elements stay on the same row. If the two lists have different lengths per row, this raises, which is the signal that they were never aligned.
Variant phrasings
pandas explode creates NaN rows
Empty lists explode to a single NaN row by design (step 4). Drop them or keep them as meaningful "empty" markers, but decide.
pandas split string column into rows
First df["col"].str.split(",") to get lists, then explode. Explode only works on list-likes.
pandas explode duplicate index
Explode repeats the original index labels. Call .reset_index(drop=True) after (step 3) unless you need the labels to trace back to source rows.
Why it happens
explode maps each element of a list-like cell to its own row while broadcasting the scalar columns. It is the inverse of a groupby-agg into lists, and like any reshape it preserves information only if the index and the empty-list handling are managed deliberately.
Edge cases
- Lists containing None explode the None into its own row, decide if that is a real element.
- Exploding a column of numpy arrays works, but ragged arrays must be object dtype first.
- Very long lists per row can blow up memory, check total element count (step 1) before exploding huge frames.
ignore_index=Trueinside explode does the reset in one call.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_dF7pJpNcKsLo5tWv5wv3yg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.