dbt seed failed with duplicate column name" csv
Fixes dbt seed failures caused by duplicate column headers in the CSV by renaming the duplicated header and reloading with a full refresh. Use when dbt seed fails naming a duplicate column. Not for duplicate data rows, which seeds accept.
TL;DR
The CSV header row contains the same column name twice, and the warehouse rejects the resulting table definition. Rename one of the duplicate headers in the CSV file, then reload with dbt seed --full-refresh.
Error
"dbt seed failed with duplicate column name" csvSteps
- Open
seeds/[FILE].csvand inspect the header row. Expected: you find the column name appearing twice (watch for trailing spaces making them look different). - Rename one of the duplicates to a distinct, meaningful name. Expected: every header value is unique.
- Save the CSV, keeping the same delimiter and encoding. Expected: the file still parses as the same number of columns.
- Run
dbt seed --select [SEED NAME] --full-refresh. Expected: the seed loads without a duplicate-column error. - Run
dbt test --select [SEED NAME]if tests are defined on the seed. Expected: tests pass against the reloaded data.
When to use
dbt seedfails with a duplicate column name error.- The CSV was exported from a spreadsheet or a join that produced two same-named columns.
When not to use
- Rows are duplicated but headers are unique (a data issue, not a load issue).
- The error is about duplicate seed file names rather than columns.
Tool compatibility
- dbt Core 1.0 and later, all adapters. CSV parsing rules are adapter-independent.
Variant phrasings
Duplicate column name in seed on Snowflake / BigQuery
Same fix; the warehouse enforces unique column names at table creation.
Seed loads but downstream models break on ambiguous columns
Rename at the CSV level so every downstream ref is unambiguous.
Why it happens
dbt infers the table schema from the CSV header. Two identical headers produce two identical column definitions, which every warehouse rejects.
Edge cases
- Headers differing only by case (
IDvsid) collide on warehouses with case-insensitive identifiers. - Invisible characters (byte-order marks, trailing spaces) create duplicates that look unique; inspect the raw bytes if renaming does not help.
- Large CSVs are easier to fix with a script than by hand; validate uniqueness programmatically after the edit.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_2TCZs9sVfR7ELbAbAQLQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.