## TL;DR
Set per-column types in `dbt_project.yml` under `seeds:` with `+column_types`, because dbt infers seed types from the CSV and guesses wrong on ambiguous columns. Explicit types make seed schemas deterministic.

```text
dbt seeds csv column type overrides
```

## Use this when
- Seeds infer wrong column types
- You need deterministic seed schemas
- CSV columns are ambiguous (ids with leading zeros, mixed formats)

## Not for this skill when
- You are configuring source freshness
- You are writing unit tests
- You are choosing snapshot strategies

## Steps

1. See what dbt inferred. After `dbt seed`, check the created table:

```bash
dbt seed --select my_seed
# then inspect the table schema in the warehouse
```
Expected output: the inferred types. The classic surprises: zip codes inferred as integers (losing leading zeros), and date columns inferred as text.

2. Pin the types explicitly in `dbt_project.yml`:

```yaml
seeds:
  my_project:
    my_seed:
      +column_types:
        zip_code: varchar(10)
        signup_date: date
        user_id: integer
```
Expected output: the seed table builds with exactly these types on the next `dbt seed --full-refresh`. Types are database-specific, so use your warehouse's type names.

3. Handle the leading-zero case deliberately:

```yaml
seeds:
  my_project:
    regions:
      +column_types:
        region_code: varchar(10)  # '007' stays '007'
```
Expected output: codes keep their formatting. This is the single most common seed type bug: numeric inference destroying string identifiers.

4. For CSVs with genuinely messy columns, clean at the source instead of fighting types:

```text
If a column mixes dates and free text, no type override fixes it.
Fix the CSV generation; seeds are for small, clean reference data,
not for raw extracts.
```
Expected output: a policy. Seeds work well under a few thousand rows of curated data; they are the wrong tool for messy bulk loads.

5. Remember seeds rebuild semantics:

```bash
dbt seed                  # incremental-ish: only new/changed seeds
  dbt seed --full-refresh  # drops and recreates all seeds
```
Expected output: type changes require `--full-refresh` to take effect, since the table must be recreated. A plain `dbt seed` after editing `column_types` silently keeps the old schema.

## Variant phrasings

### dbt seed column types wrong
Override with `+column_types` (step 2) and full-refresh (step 5).

### dbt seed leading zeros stripped
The column inferred as numeric. Pin it to varchar (step 3).

### dbt_project.yml seeds config
The `seeds:` block supports `+column_types`, `+schema`, `+delimiter`, and more per seed. See the dbt docs for the full list for your adapter version.

## Why it happens
dbt seeds are CSVs, and CSVs carry no type information. dbt samples the values and guesses, which works for clean data and fails for identifiers that look numeric, dates in odd formats, and mixed columns. Explicit types remove the guessing.

## Edge cases
- `column_types` syntax varies slightly by adapter; check your adapter's docs for the exact key.
- Seeds live in version control, so a 10MB CSV in git is a smell; keep seeds small.
- Delimiter issues (commas inside quoted fields) break parsing before types even matter; set `+delimiter` explicitly for non-comma files.
- Empty strings vs NULL: dbt treats empty strings per adapter defaults; verify nullable columns behave as expected.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_noIlgm5MrTY2dNRI9lQJdg
