## TL;DR
`apply` with `axis=1` runs a Python function once per row, which is slow by design. Replace it with vectorized column operations (arithmetic, `.str` methods, `np.where`), which run in compiled code and are typically 10-100x faster. It works because pandas/numpy execute the same math on whole arrays at once instead of looping in Python.

```text
# the slow pattern
df['c'] = df.apply(lambda row: row['a'] * 2 if row['b'] > 0 else 0, axis=1)
```

## Use this when
- `df.apply(..., axis=1)` is slow on a large frame
- A per-row Python function dominates runtime
- `.apply` on a Series could be a `.str` or arithmetic operation instead

## Not for
- MemoryError on big files, use the chunking skill
- Groupby `.agg`/`.apply` shape errors, separate skill
- Code that is already vectorized but slow, the algorithm needs rethinking

## Steps

1. Measure the baseline so you know the win:

```python
%timeit df.apply(lambda row: row['a'] * 2 if row['b'] > 0 else 0, axis=1)
```
Expected output: a per-loop time, usually seconds on 100k+ rows.

2. Replace if/else row logic with `np.where`:

```python
import numpy as np
df['c'] = np.where(df['b'] > 0, df['a'] * 2, 0)
```
Expected output: identical values, runs in milliseconds.

3. Replace string munging with `.str` accessors:

```python
df['domain'] = df['email'].str.split('@').str[1]
```
Expected output: vectorized string ops, no per-row lambda.

4. Replace row math with plain column arithmetic:

```python
df['total'] = df['price'] * df['qty'] * (1 - df['discount'])
```
Expected output: the whole column computed in one compiled pass.

5. Verify equality on a sample before committing:

```python
(df['c_new'] == df['c_old']).all()
```
Expected output: True. Vectorized rewrites can differ on NaN handling, so check.

## Variant phrasings

### pandas apply too slow alternative
Vectorize first. If the logic truly cannot vectorize (complex branching, external calls), `numba` or `np.vectorize` are the next rungs, still faster than apply.

### speed up row-wise operations pandas
Same answer. Also consider: do you need all rows? Filtering before the expensive op beats speeding up the op.

### pandas apply axis=1 slow
`axis=1` is the worst case: it builds a Series per row. Column-wise `apply` (default axis=0) is less bad but still Python-speed.

## Why it happens
`apply` calls your Python function once per row (or per column), with full Python call overhead each time plus Series construction. Vectorized ops push the loop into C: one call, the whole array, no per-element overhead. The speedup is structural, not a tweak.

## Edge cases
- Truly unvectorizable logic (regex with backrefs, API calls per row): use `numba.jit` on numpy arrays, or accept apply and parallelize with swifter/Dask.
- `np.where` with NaN conditions: NaN comparisons are False, verify that matches your old lambda's behavior.
- Datetime row logic: use `.dt` accessors (`df['d'].dt.year`) instead of per-row parsing.
- Watch dtype changes: vectorized string ops can upcast; check `dtypes` after.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_KcVlsCMk8JJTi74Bec6jrw
