## TL;DR
Pass explicit dtypes for the misinferred columns via the `schema` or `dtypes` argument, and set `infer_schema_length` high enough to see the real data. polars infers types from the first rows it samples, so a column that starts with nulls or short strings gets the wrong type.

```text
polars read_csv misinferring column types fix
```

## Use this when
- A numeric column reads in as String (or the reverse)
- Columns you know have values come back full of nulls
- `read_csv` throws a parse error on one column

## Not for this skill when
- The file itself is corrupt or has genuinely mixed garbage rows
- The CSV is too large to fit in memory at all
- You are reading parquet with a schema conflict, not CSV

## Steps

1. Confirm what polars inferred and where it went wrong:

```python
df = pl.read_csv("orders.csv")
print(df.schema)
print(df.null_count())
```
Expected output: the inferred schema plus per-column null counts. Columns with suspiciously high null counts are the misinferred ones.

2. Look at the raw values in the bad column to see what confused inference:

```python
df_bad = pl.read_csv("orders.csv", dtypes={"price": pl.String})
print(df_bad["price"].unique().sort().head(20))
```
Expected output: the raw strings, often revealing currency symbols, commas, `"N/A"` markers, or empty strings that broke numeric inference.

3. Read the column as String first, then clean and cast in a controlled way:

```python
df = pl.read_csv("orders.csv", dtypes={"price": pl.String})
df = df.with_columns(
    pl.col("price").str.replace_all(",", "").cast(pl.Float64, strict=False)
)
print(df["price"].null_count())
```
Expected output: a clean Float64 column. `strict=False` turns unparseable values into null instead of throwing, and the null count tells you how many rows had junk.

4. If inference is fine but the sample was too small, raise the inference window:

```python
df = pl.read_csv("orders.csv", infer_schema_length=10000)
```
Expected output: correct dtypes when the misleading rows were past the default 100-row sample. Full-file inference costs a slower read, so prefer explicit dtypes for files you read repeatedly.

5. For files you read on a schedule, pin the schema explicitly so it never drifts:

```python
schema = {"order_id": pl.Int64, "price": pl.Float64, "placed_at": pl.Datetime}
df = pl.read_csv("orders.csv", schema=schema)
```
Expected output: the same schema every run, regardless of file content. A column whose values cannot parse now throws loudly instead of silently misinferring.

## Variant phrasings

### polars read_csv column all null
Read the column as String (step 2). The values are probably formatted in a way the inferred type rejected, like `"12,500"` for an Int64.

### polars parse error on read_csv
The inferred dtype rejected some value further down the file. Either widen with `strict=False` casting (step 3) or pin the schema (step 5) to find the exact bad rows.

### polars date column read as string
Pass `try_parse_dates=True` to read_csv, or cast explicitly with `pl.col("d").str.to_date("%Y-%m-%d")` matching your actual format.

## Why it happens
polars infers each column's dtype by parsing a sample of rows (100 by default). If the sample contains only nulls, short strings, or oddly formatted numbers, inference picks the wrong type and every later value either becomes null or throws. The behavior is deterministic given the sample, which is why the same file misinfers the same way every time.

## Edge cases
- Mixed-type columns where some rows are numbers and some are text: there is no numeric dtype for those, keep String and clean.
- Large ints above 2^63 overflow Int64 silently in some paths, use String for IDs you never do math on.
- `null_values=["N/A", "-", ""]` handles custom null markers at read time so they do not poison inference.
- BOM at the start of UTF-8 files can corrupt the first column name, strip it when writing the file.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_rh3aPYIERU6-GtbIegRCDw
