## TL;DR
Staging models mirror your sources one-to-one with light cleanup and a `stg_` prefix, marts models serve business questions with `fct_` and `dim_` prefixes. Keep staging boring and marts opinionated, and put anything shared between marts in an intermediate layer. This split is the smallest structure that stays understandable at fifty models and still works at five hundred.

```text
how to structure dbt models: staging vs marts
```

## Use this when
- you are setting up a new dbt project and need a folder layout that scales
- staging models are accumulating business logic and getting messy
- reviewers cannot tell which models are safe to query directly

## Not for this skill when
- you are debugging a specific model failure, fix the model first
- your project already uses a different convention consistently, consistency beats this layout
- you need help with sources yaml, that is a separate skill

## Steps

1. Create the folder split so the layers are visible in the repo before you write any models:

```shell
mkdir -p models/staging/my_source models/marts/core models/intermediate
```

Expected output: three directories. Staging holds one model per source table, marts holds business-ready models, intermediate is optional glue for shared logic.

2. Write a staging model that only renames and recasts, with zero business rules in it:

```sql
-- models/staging/my_source/stg_orders.sql
SELECT
  id AS order_id,
  created_at::timestamp AS ordered_at,
  total_cents / 100.0 AS order_total
FROM {{ source('my_source', 'orders') }}
```

Expected output: a clean one-to-one mirror of the source. Anyone can read this file and know exactly what the raw table holds, which is the whole point.

3. Put join-heavy reshaping in intermediate models with an `int_` prefix so marts never duplicate the same join:

```sql
-- models/intermediate/int_orders_enriched.sql
SELECT o.*, c.customer_segment
FROM {{ ref('stg_orders') }} o
LEFT JOIN {{ ref('stg_customers') }} c USING (customer_id)
```

Expected output: reusable joined logic that multiple marts can share. One `int_` model beats three copies of the same join drifting apart.

4. Build marts models that answer business questions and nothing else, with documented grains:

```sql
-- models/marts/core/fct_orders.sql
SELECT order_id, ordered_at, order_total, customer_segment
FROM {{ ref('int_orders_enriched') }}
WHERE ordered_at >= '2024-01-01'
```

Expected output: a business-ready table that dashboards and analysts query directly. If a stakeholder asks what a column means, the answer lives in this layer's yaml.

5. Materialize each layer for how it is actually used, cheap upstream and fast downstream:

```yaml
# dbt_project.yml
models:
  my_project:
    staging:
      +materialized: view
    marts:
      +materialized: table
```

Expected output: staging stays cheap as views that always reflect the source, marts are fast to query as tables. Rebuilds stay quick because only the marts layer does heavy work.

## Variant phrasings

### dbt staging layer best practices
One model per source table, light transforms only, `stg_` prefix, views. If a staging model grows a CASE WHEN with business meaning, that logic belongs downstream.

### dbt marts folder structure
Group by business domain under models/marts, like core, finance, and marketing. Each mart documents its grain in the schema yaml so consumers know what one row means.

### when to use intermediate models in dbt
When two or more marts would duplicate the same join or aggregation. Reach for `int_` models at the second duplication, not the fifth.

## Why it happens
Without layers, every model mixes raw cleanup with business logic, and changing one breaks five others in ways nobody can trace. Staging absorbs source changes in exactly one place, marts absorb business changes in exactly one place, and the lineage graph between them documents itself. The structure is really about limiting blast radius.

## Edge cases
- Tiny projects: two layers can feel like overkill under ten models. Keep the folders anyway, growth is cheaper than reorganization later.
- Snapshots sit beside staging, not inside it. They have their own change-tracking semantics.
- Seeds are not staging models. Keep CSVs in the seeds folder and document them separately.
- Ephemeral models look tidy but hide lineage in the DAG. Prefer views unless compile time truly hurts.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_lFar8eNWiJgU3myqalNAWg
