## TL;DR
Pandas could not guess your date format, so it gave up. Fix it by telling pandas the format explicitly with `format='%Y-%m-%d'` (match your data), and add `errors='coerce'` while debugging. It happens because mixed or unusual formats defeat the auto-parser, and in pandas 2.x the strict parser raises instead of guessing.

```text
pandas._libs.tslibs.parsing.DateParseError: Unknown string format
```

## Use this when
- `pd.to_datetime(col)` raises "unknown string format" or ParserError
- Dates parse as NaT even though the strings look fine
- A column mixes formats like "2024-01-05" and "01/05/2024"

## Not for
- Timezone-aware vs naive comparison errors, separate skill
- Parsing dates in SQL or Spark
- Datetime math after parsing works fine, different topic

## Steps

1. Look at the actual unique formats in the column:

```python
df['d'].dropna().astype(str).str[:10].unique()[:20]
```
Expected output: the raw string shapes, so you can see the mix.

2. Parse with an explicit format:

```python
pd.to_datetime(df['d'], format='%Y-%m-%d')
```
Expected output: a datetime64 column, no error. Adjust the format string to match what step 1 showed.

3. If formats are mixed, try the flexible parser with coercion:

```python
pd.to_datetime(df['d'], format='mixed', errors='coerce')
```
Expected output: parses what it can, NaT where it cannot. Then inspect the NaT rows.

4. Find the rows that failed:

```python
df[pd.to_datetime(df['d'], format='mixed', errors='coerce').isna()]
```
Expected output: the offending rows. Usually junk strings, blanks, or a second format.

5. For day-first ambiguity ("05/01/2024"), say so explicitly:

```python
pd.to_datetime(df['d'], format='%d/%m/%Y')
```
Expected output: correct dates instead of month/day mixups. Never rely on the default guess for these.

## Variant phrasings

### pandas ParserError unknown string format present at position
Same fix. The "position" tells you which row broke the parser; look at that row.

### to_datetime returns NaT for valid dates
Usually `errors='coerce'` silently ate unparsable values, or the format arg did not match. Run step 4 to see what failed.

### pandas 2.0 to_datetime now raises
In pandas 2.x the default parser got stricter. Code that worked in 1.x needs an explicit `format=` now. Use `format='mixed'` as the drop-in flexible option.

## Why it happens
Auto-detecting date formats is guesswork, and pandas 2.x made the parser strict: if the first chunk of values does not match one clean format, it raises instead of muddling through. Explicit `format=` removes the guesswork entirely and is also faster.

## Edge cases
- Unix timestamps as integers need `unit='s'` (or 'ms'): `pd.to_datetime(df['ts'], unit='s')`.
- Excel serial dates need an origin: `pd.to_datetime(df['x'], unit='D', origin='1899-12-30')`.
- Two-digit years ("24-01-05"): pandas assumes a pivot year; spell it out with `%y` and verify a sample.
- `errors='coerce'` in production hides data problems; log the count of NaT it produces.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_bianusWbGNHQHHdUqtmopQ
