## TL;DR
Use `category` dtype for string columns with low cardinality (few repeated values like status, region, type); keep `object` for free text, IDs, and near-unique strings. Category stores one small integer per row plus a lookup table, so it saves memory and speeds up groupby and sorting. It matters most when the same strings repeat thousands of times.

```python
df['status'] = df['status'].astype('category')
```

## Use this when
- A string column repeats a small set of values
- Groupby/sort on strings is slow
- String columns dominate memory usage

## Not for
- Numeric downcasting (int64 to int32), separate concern
- SQL enum types
- Free-text columns (names, descriptions, IDs)

## Steps

1. Check cardinality before converting:

```python
df['col'].nunique(), len(df)
```
Expected output: unique count vs row count. Category wins when unique is much smaller than rows (rule of thumb: under 50 percent, ideally under 5 percent).

2. Convert the low-cardinality columns:

```python
df['status'] = df['status'].astype('category')
```
Expected output: dtype category; memory for the column drops sharply.

3. Verify the memory win:

```python
df.memory_usage(deep=True)
```
Expected output: the converted column now costs roughly 1-2 bytes per row plus the category table.

4. For ordered data (sizes, ratings), declare the order:

```python
df['size'] = pd.Categorical(df['size'], categories=['S','M','L','XL'], ordered=True)
```
Expected output: comparisons like `df['size'] > 'M'` now work logically.

5. Benchmark the groupby speedup:

```python
%timeit df.groupby('status')['v'].sum()
```
Expected output: noticeably faster than on object dtype, since grouping hashes small ints instead of strings.

## Variant phrasings

### pandas category dtype when to use
Low-cardinality repeated strings. That is the whole rule.

### convert object to category pandas memory
Same conversion; the memory win is the usual motivation. Watch out for near-unique columns where it backfires.

### pandas categorical vs string dtype
The newer `string` dtype is for text operations; `category` is for repeated labels. They solve different problems and can combine.

## Why it happens
Object dtype stores a full Python string object per cell. Category stores each distinct value once and a tiny integer code per row. Less memory, and operations like groupby, sort, and merge compare ints instead of hashing strings.

## Edge cases
- Near-unique strings (UUIDs, emails): category ADDS overhead (the lookup table) with no benefit. Keep object.
- New unseen values: assigning a value not in the categories raises; use `cat.add_categories` first or leave as object.
- Merging on categoricals requires matching categories on both sides or pandas upcasts; align with `.astype('category')` on both.
- `pd.get_dummies` on category includes all categories even unobserved ones unless `observed=True`; groupby has the same `observed` parameter.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_L5Rwnau5nBYFWXjQbIoxig
