## TL;DR
`pd.concat` aligns on column names, so if two frames spell a column differently (`"user_id"` vs `"user_id "` with a trailing space), you get one column of real values plus one column of NaNs, silently. No error, no warning. The fix is boring but reliable: diff the column sets before concatenating, normalize names (strip whitespace, unify case), then concat and assert the NaN pattern looks sane.

## The query
```
pandas concat mismatched columns silent NaN fix
```

## Use this when
- a concat result has mysterious NaN columns that shouldn't be there
- you are stacking frames from different sources (CSVs, APIs, exports)
- a data agent builds frames programmatically and needs a concat sanity check

## Not for
- merge/join key mismatches (different function, different fix)
- NaNs that come from the source data itself (check the inputs first)
- intentionally ragged frames where NaN fill is the desired behavior

## Steps
1. Diff the column sets before you concat. This one-liner finds the problem in seconds:
```python
set(df1.columns) ^ set(df2.columns)
```
Expected output: the symmetric difference shows exactly which names don't line up.

2. Inspect the offenders with `repr()`. Trailing spaces, non-breaking spaces, and case differences are invisible in normal printing but obvious in repr.
```python
[repr(c) for c in df1.columns if c not in set(df2.columns)]
```
Expected output: you can see the hidden characters causing the mismatch.

3. Normalize both frames' columns the same way: `df.columns = df.columns.str.strip().str.lower()` is the usual 90% fix. Apply it right after loading, before any concat.
Expected output: the column-set diff from step 1 is now empty.

4. Concat, then verify with intent. Check `result.isna().sum()` on the previously mismatched columns and confirm the NaN counts match what the source data actually contains, not alignment artifacts.
Expected output: no NaN columns that can't be explained by the inputs.

5. Make it a habit: a tiny helper that asserts equal column sets (or logs the diff) before every concat in a pipeline. Silent NaNs are a data-quality bug that compounds downstream.
Expected output: future mismatches raise loudly at concat time instead of poisoning results quietly.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_TT4MIKJy_IqwLuzJQkj6ag
