## TL;DR
You ran out of RAM because pandas loads everything into memory at once. Fix it by processing the file in chunks: pass `chunksize=` to `read_csv` and aggregate per chunk, and downcast dtypes before anything else. Most MemoryErrors disappear with those two moves alone.

```text
MemoryError
```

## Use this when
- `pd.read_csv` raises `MemoryError` or the process is OOM-killed
- A merge, groupby, or concat blows past available RAM
- You need row counts, sums, or per-group stats on a file bigger than memory

## Not for
- Slow but working pandas code, vectorize instead
- Rewriting the job in Dask or Spark (do that when chunking gets too painful)
- MemoryErrors from numpy or model training, different territory

## Steps

1. Check how big the frame actually is once loaded (or estimate):

```python
df.memory_usage(deep=True).sum() / 1e9
```
Expected output: gigabytes used. Object columns usually dominate.

2. Downcast dtypes on load. This is the cheapest win:

```python
df = pd.read_csv('big.csv', dtype={'id': 'int32', 'flag': 'bool'})
```
Expected output: same data, often 50-70 percent less memory than default int64/object inference.

3. Read only the columns you need:

```python
df = pd.read_csv('big.csv', usecols=['a', 'b', 'c'])
```
Expected output: a frame with just those columns, proportionally smaller.

4. Process in chunks and combine the partial results:

```python
chunks = pd.read_csv('big.csv', chunksize=100_000)
result = pd.concat([c.groupby('k')['v'].sum() for c in chunks]).groupby(level=0).sum()
```
Expected output: the correct total, computed with peak memory of one chunk plus the result.

5. Convert repeated strings to category:

```python
df['region'] = df['region'].astype('category')
```
Expected output: memory for that column drops to a fraction, groupby on it gets faster too.

## Variant phrasings

### pandas out of memory large csv
Chunksize plus usecols plus category dtypes handles nearly all of these.

### python killed reading csv pandas
"Killed" usually means the OS OOM killer, not a Python exception. Same fix. Check `dmesg` for oom-killer lines to confirm.

### MemoryError on pandas merge
Merge multiplies memory: both frames plus the join result. Downcast both frames first, merge on the smallest key set, and consider chunking the larger side.

## Why it happens
Pandas keeps the whole frame in RAM, and its default dtypes are generous: int64 for every integer, full Python objects for strings. A CSV that is 2 GB on disk can easily become 8-10 GB in memory. Chunking keeps the working set to one piece at a time, and downcasting shrinks every piece.

## Edge cases
- `chunksize` returns an iterator: you must loop it, you cannot call `.head()` on the whole thing.
- Mixed-type columns force object dtype and kill the category trick; clean or specify dtypes.
- If one group dominates (skewed keys), the per-chunk aggregation in step 4 can still blow up; pre-filter or shard by key.
- When chunking is too awkward (window functions, cross-chunk joins), that is the signal to move to Dask or DuckDB, which stream for you.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_3i_dTwKT6m8vnL8xAsvxnw
