## TL;DR
Run `df.memory_usage(deep=True)` to see bytes per column, sort it, and attack the top offenders: downcast numerics, convert repeated strings to category, and drop what you dont need. `deep=True` is the key detail; without it, object columns report only pointer sizes and lie about their real cost.

```python
df.memory_usage(deep=True).sort_values(ascending=False)
```

## Use this when
- A DataFrame uses more RAM than expected
- You need to find which columns to downcast or drop
- Before/after comparison for a dtype optimization pass

## Not for
- MemoryError on files that dont fit in RAM, use the chunking skill
- Profiling Python code outside pandas
- CPU/runtime profiling, use cProfile or line_profiler

## Steps

1. Get per-column usage with deep introspection:

```python
usage = df.memory_usage(deep=True, index=True).sort_values(ascending=False)
print(usage)
```
Expected output: bytes per column, index included, biggest first.

2. Compare against the shallow number to see hidden string costs:

```python
df.memory_usage(deep=False).sum(), df.memory_usage(deep=True).sum()
```
Expected output: two totals. A big gap means object columns hold most of the real memory.

3. Downcast the numeric offenders:

```python
df['n'] = pd.to_numeric(df['n'], downcast='integer')
df['f'] = pd.to_numeric(df['f'], downcast='float')
```
Expected output: int64 to int32/int16/int8 where values fit; rerun step 1 to confirm the drop.

4. Convert repeated-string columns to category:

```python
df['region'] = df['region'].astype('category')
```
Expected output: the column shrinks to roughly one small int per row plus the category table.

5. Report the total win:

```python
print(f"{(before - after) / before:.0%} smaller")
```
Expected output: a percentage. 50-80 percent is typical on messy real-world frames.

## Variant phrasings

### pandas dataframe memory usage by column
`df.memory_usage(deep=True)` is the one-liner. Sort it and work top-down.

### why is my dataframe so large pandas
Usually object-dtype strings and int64/float64 defaults. Steps 3-4 fix both.

### reduce pandas dataframe memory footprint
Profile first (step 1), then downcast, categorize, and drop unused columns, in that order.

## Why it happens
Pandas defaults are generous: int64 for any integer, float64 for any float, and full Python objects for strings. Each is 2-8x bigger than the data needs. `deep=False` hides the string cost because it only counts the 8-byte pointers, which is why people underestimate object columns.

## Edge cases
- `deep=True` is slower on huge frames; sample 100k rows for a quick estimate.
- Category dtype has overhead per category; dont use it for near-unique strings like IDs or free text.
- Downcasting can overflow: verify `df['n'].max()` fits the target dtype before committing.
- The index itself can be big (a large string index); `set_index` on an int column or `reset_index` to a RangeIndex helps.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_daeY-YYVANuMjuRBZU-CgA
