how to handle mixed types in a pandas column
Handles mixed-type pandas columns (object dtype with strings, numbers, NaN mixed together). Use when a column has dtype object but should be numeric or datetime, when conversions raise, or when comparisons behave oddly. Do not use for duplicate labels, merge errors, or memory problems.
TL;DR
A column with mixed types sits at dtype object, which makes math fail and comparisons lie. Fix it by finding what the non-conforming values are, coercing with pd.to_numeric(errors='coerce'), and handling the NaT/NaN leftovers deliberately. Mixed types almost always come from dirty source data, not a pandas bug.
TypeError: unsupported format string passed to numpy.ndarray.__format__Use this when
- A column is dtype
objectbut should be numeric or datetime pd.to_numericorpd.to_datetimeraises on the column- Sorting, max/min, or comparisons on the column give nonsense results
Not for
- Duplicate index labels, reindex errors
- Merge key type mismatches (int vs str keys), fix the key dtypes instead
- Memory blowups from object columns, use the chunking skill
Steps
- See what types are actually in there:
df['c'].map(type).value_counts()Expected output: the mix, e.g. 9500 float, 400 str, 100 NoneType.
- Find the non-numeric offenders:
df[pd.to_numeric(df['c'], errors='coerce').isna() & df['c'].notna()]['c'].unique()[:20]Expected output: the junk values: "N/A", "-", "12.5%", empty strings, whatever they are.
- Coerce to numeric, junk becomes NaN:
df['c'] = pd.to_numeric(df['c'], errors='coerce')Expected output: dtype float64, junk values now NaN.
- Decide what the NaNs mean and handle them:
df['c'] = df['c'].fillna(0) # or .dropna(), or keep as NaNExpected output: no silent NaNs left unless you chose to keep them.
- If the strings were meaningful (like "12.5%"), clean first, then convert:
df['c'] = pd.to_numeric(df['c'].astype(str).str.rstrip('%'), errors='coerce') / 100Expected output: proper floats, e.g. 0.125 instead of the string "12.5%".
Variant phrasings
pandas column object dtype should be numeric
The column got object dtype because at least one value wasnt numeric. Steps 1-3 find and coerce them.
DtypeWarning columns with mixed types on read_csv
read_csv warns when a column mixes types across chunks. Pass an explicit dtype= for that column, or set low_memory=False, then clean with the steps above.
string and float mixed in pandas column
Usually numbers-as-strings plus real numbers. pd.to_numeric with errors='coerce' unifies them; check what became NaN.
Why it happens
Pandas assigns one dtype per column. The moment a single value cannot fit (a "N/A" string in a numeric column), the whole column falls back to object dtype, and every numeric operation on it either fails or silently does string things. Source files are the usual culprit: footers, placeholder text, locale-specific formats.
Edge cases
- Boolean-ish mixes ("yes"/"no"/1/0): map explicitly with a dict, dont rely on coercion.
- Datetime mixes:
pd.to_datetime(errors='coerce')the same way, then inspect the NaT rows. infer_objects()only soft-converts; it wont fix real junk.- After coercion, downcast (
pd.to_numeric(..., downcast='integer')) if memory matters.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_D9SR2EZmbCqG88gPDcXAsQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.