dbt seeds csv column type overrides
Overrides dbt seed CSV column types so seeds load with the right schema. Use when dbt seeds infer wrong types (numbers as text, dates as strings), when you need explicit column types per seed, or when seed CSVs have mixed-type columns. Not for source freshness thresholds, for unit test fixtures, or for snapshot strategies.
TL;DR
Set per-column types in dbt_project.yml under seeds: with +column_types, because dbt infers seed types from the CSV and guesses wrong on ambiguous columns. Explicit types make seed schemas deterministic.
dbt seeds csv column type overridesUse this when
- Seeds infer wrong column types
- You need deterministic seed schemas
- CSV columns are ambiguous (ids with leading zeros, mixed formats)
Not for this skill when
- You are configuring source freshness
- You are writing unit tests
- You are choosing snapshot strategies
Steps
- See what dbt inferred. After
dbt seed, check the created table:
dbt seed --select my_seed
# then inspect the table schema in the warehouseExpected output: the inferred types. The classic surprises: zip codes inferred as integers (losing leading zeros), and date columns inferred as text.
- Pin the types explicitly in
dbt_project.yml:
seeds:
my_project:
my_seed:
+column_types:
zip_code: varchar(10)
signup_date: date
user_id: integerExpected output: the seed table builds with exactly these types on the next dbt seed --full-refresh. Types are database-specific, so use your warehouse's type names.
- Handle the leading-zero case deliberately:
seeds:
my_project:
regions:
+column_types:
region_code: varchar(10) # '007' stays '007'Expected output: codes keep their formatting. This is the single most common seed type bug: numeric inference destroying string identifiers.
- For CSVs with genuinely messy columns, clean at the source instead of fighting types:
If a column mixes dates and free text, no type override fixes it.
Fix the CSV generation; seeds are for small, clean reference data,
not for raw extracts.Expected output: a policy. Seeds work well under a few thousand rows of curated data; they are the wrong tool for messy bulk loads.
- Remember seeds rebuild semantics:
dbt seed # incremental-ish: only new/changed seeds
dbt seed --full-refresh # drops and recreates all seedsExpected output: type changes require --full-refresh to take effect, since the table must be recreated. A plain dbt seed after editing column_types silently keeps the old schema.
Variant phrasings
dbt seed column types wrong
Override with +column_types (step 2) and full-refresh (step 5).
dbt seed leading zeros stripped
The column inferred as numeric. Pin it to varchar (step 3).
dbt_project.yml seeds config
The seeds: block supports +column_types, +schema, +delimiter, and more per seed. See the dbt docs for the full list for your adapter version.
Why it happens
dbt seeds are CSVs, and CSVs carry no type information. dbt samples the values and guesses, which works for clean data and fails for identifiers that look numeric, dates in odd formats, and mixed columns. Explicit types remove the guessing.
Edge cases
column_typessyntax varies slightly by adapter; check your adapter's docs for the exact key.- Seeds live in version control, so a 10MB CSV in git is a smell; keep seeds small.
- Delimiter issues (commas inside quoted fields) break parsing before types even matter; set
+delimiterexplicitly for non-comma files. - Empty strings vs NULL: dbt treats empty strings per adapter defaults; verify nullable columns behave as expected.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_noIlgm5MrTY2dNRI9lQJdg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.