pyarrow.lib.ArrowInvalid: Could not convert with type
Fixes pyarrow ArrowInvalid: Could not convert with type errors. Use when arrow table casts or parquet writes fail on type conversion, when an agent pushes pandas frames into arrow with mixed types, or when timestamps carry timezones inconsistently. Not for IO errors, for schema mismatch on read, or for out-of-memory failures.
TL;DR
Arrow is strict about types: converting data to a declared type fails the whole batch when any value doesnt fit (a string in an int column, a tz-aware timestamp into a tz-naive field). Find the offending values with a permissive read first, clean them, then cast explicitly.
pyarrow.lib.ArrowInvalid: Could not convert with typeUse this when
pa.Table.from_pandasor a cast raises ArrowInvalid- Parquet writes fail on type conversion
- An agent moves messy pandas data into arrow
Not for this skill when
- The file cant be opened (thats IO)
- The schema on read doesnt match (thats a read-schema problem)
- The process OOMs (thats memory)
Steps
- Reproduce with the value that fails. Arrow's message usually names the type and sometimes the value:
import pyarrow as pa
try:
pa.array(["1", "2", "oops"], type=pa.int64())
except pa.ArrowInvalid as e:
print(e)Expected output: the error naming the bad value. Thats your cleaning target.
- Find all offending values in the real column before converting:
bad = df[pd.to_numeric(df["col"], errors="coerce").isna() & df["col"].notna()]
print(bad["col"].unique()[:10])Expected output: the distinct unconvertible values. Decide per value: fix, coerce to null, or drop.
- Clean, then convert with an explicit safe cast:
df["col"] = pd.to_numeric(df["col"], errors="coerce")
arr = pa.array(df["col"], type=pa.int64(), from_pandas=True)Expected output: no ArrowInvalid; unconvertible values became null. from_pandas=True maps NaN to null properly.
- For timestamps, unify timezone awareness before converting:
df["ts"] = pd.to_datetime(df["ts"], utc=True)
arr = pa.array(df["ts"], type=pa.timestamp("us", tz="UTC"))Expected output: clean conversion. Mixed naive/aware timestamps are the most common ArrowInvalid in time-series pipelines.
Variant phrasings
could not convert with type timestamp
Timezone mismatch, step 4. Arrow timestamps carry tz in the type; pandas often doesnt.
fails only on the full dataset, not the sample
The bad value isnt in your sample. Run step 2 over the whole column, not a head.
Why it happens
Arrow validates every value against the target type during conversion and aborts on the first failure instead of coercing silently like pandas does. This strictness is a feature (no silent corruption), but it means one bad string in a million-row int column kills the whole write. Agents hit it moving pandas frames (loose types) into arrow/parquet (strict types) without a cleaning pass.
Edge cases
pa.compute.castwithsafe=Falseforces conversion but can silently wrap or truncate; prefer cleaning over unsafe casts.- Dictionary-encoded columns with new unseen values fail the same way; unify the dictionary first.
- Null handling differs: pandas NaN, None, and NaT map differently;
from_pandas=Trueis what makes them all null.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstuehdGCpCIBg9D4SglM3Eg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.