## TL;DR
Arrow is strict about types: converting data to a declared type fails the whole batch when any value doesnt fit (a string in an int column, a tz-aware timestamp into a tz-naive field). Find the offending values with a permissive read first, clean them, then cast explicitly.

```text
pyarrow.lib.ArrowInvalid: Could not convert with type
```

## Use this when
- `pa.Table.from_pandas` or a cast raises ArrowInvalid
- Parquet writes fail on type conversion
- An agent moves messy pandas data into arrow

## Not for this skill when
- The file cant be opened (thats IO)
- The schema on read doesnt match (thats a read-schema problem)
- The process OOMs (thats memory)

## Steps

1. Reproduce with the value that fails. Arrow's message usually names the type and sometimes the value:

```python
import pyarrow as pa
try:
    pa.array(["1", "2", "oops"], type=pa.int64())
except pa.ArrowInvalid as e:
    print(e)
```
Expected output: the error naming the bad value. Thats your cleaning target.

2. Find all offending values in the real column before converting:

```python
bad = df[pd.to_numeric(df["col"], errors="coerce").isna() & df["col"].notna()]
print(bad["col"].unique()[:10])
```
Expected output: the distinct unconvertible values. Decide per value: fix, coerce to null, or drop.

3. Clean, then convert with an explicit safe cast:

```python
df["col"] = pd.to_numeric(df["col"], errors="coerce")
arr = pa.array(df["col"], type=pa.int64(), from_pandas=True)
```
Expected output: no ArrowInvalid; unconvertible values became null. `from_pandas=True` maps NaN to null properly.

4. For timestamps, unify timezone awareness before converting:

```python
df["ts"] = pd.to_datetime(df["ts"], utc=True)
arr = pa.array(df["ts"], type=pa.timestamp("us", tz="UTC"))
```
Expected output: clean conversion. Mixed naive/aware timestamps are the most common ArrowInvalid in time-series pipelines.

## Variant phrasings

### could not convert with type timestamp
Timezone mismatch, step 4. Arrow timestamps carry tz in the type; pandas often doesnt.

### fails only on the full dataset, not the sample
The bad value isnt in your sample. Run step 2 over the whole column, not a head.

## Why it happens
Arrow validates every value against the target type during conversion and aborts on the first failure instead of coercing silently like pandas does. This strictness is a feature (no silent corruption), but it means one bad string in a million-row int column kills the whole write. Agents hit it moving pandas frames (loose types) into arrow/parquet (strict types) without a cleaning pass.

## Edge cases
- `pa.compute.cast` with `safe=False` forces conversion but can silently wrap or truncate; prefer cleaning over unsafe casts.
- Dictionary-encoded columns with new unseen values fail the same way; unify the dictionary first.
- Null handling differs: pandas NaN, None, and NaT map differently; `from_pandas=True` is what makes them all null.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_ueh_dGCpCIBg9D4SglM3Eg
