pandas categorical vs object dtype: when it matters
Explains when pandas categorical vs object dtype matters. Use when deciding whether to convert string columns to category, when groupby or merge on strings is slow, or when memory from string columns is high. Do not use for numeric downcasting, for enum types in SQL, or for general dtype questions.
TL;DR
Use category dtype for string columns with low cardinality (few repeated values like status, region, type); keep object for free text, IDs, and near-unique strings. Category stores one small integer per row plus a lookup table, so it saves memory and speeds up groupby and sorting. It matters most when the same strings repeat thousands of times.
df['status'] = df['status'].astype('category')Use this when
- A string column repeats a small set of values
- Groupby/sort on strings is slow
- String columns dominate memory usage
Not for
- Numeric downcasting (int64 to int32), separate concern
- SQL enum types
- Free-text columns (names, descriptions, IDs)
Steps
- Check cardinality before converting:
df['col'].nunique(), len(df)Expected output: unique count vs row count. Category wins when unique is much smaller than rows (rule of thumb: under 50 percent, ideally under 5 percent).
- Convert the low-cardinality columns:
df['status'] = df['status'].astype('category')Expected output: dtype category; memory for the column drops sharply.
- Verify the memory win:
df.memory_usage(deep=True)Expected output: the converted column now costs roughly 1-2 bytes per row plus the category table.
- For ordered data (sizes, ratings), declare the order:
df['size'] = pd.Categorical(df['size'], categories=['S','M','L','XL'], ordered=True)Expected output: comparisons like df['size'] > 'M' now work logically.
- Benchmark the groupby speedup:
%timeit df.groupby('status')['v'].sum()Expected output: noticeably faster than on object dtype, since grouping hashes small ints instead of strings.
Variant phrasings
pandas category dtype when to use
Low-cardinality repeated strings. That is the whole rule.
convert object to category pandas memory
Same conversion; the memory win is the usual motivation. Watch out for near-unique columns where it backfires.
pandas categorical vs string dtype
The newer string dtype is for text operations; category is for repeated labels. They solve different problems and can combine.
Why it happens
Object dtype stores a full Python string object per cell. Category stores each distinct value once and a tiny integer code per row. Less memory, and operations like groupby, sort, and merge compare ints instead of hashing strings.
Edge cases
- Near-unique strings (UUIDs, emails): category ADDS overhead (the lookup table) with no benefit. Keep object.
- New unseen values: assigning a value not in the categories raises; use
cat.add_categoriesfirst or leave as object. - Merging on categoricals requires matching categories on both sides or pandas upcasts; align with
.astype('category')on both. pd.get_dummieson category includes all categories even unobserved ones unlessobserved=True; groupby has the sameobservedparameter.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_L5Rwnau5nBYFWXjQbIoxig
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.