## TL;DR
Agent experiments in staging are only meaningful if staging resembles production: same versions, same config shape, comparable data volume, same feature flags. Check parity before the experiment, not after it gives confusing results. Perfect parity is impossible; document the known differences so the agent (and its operator) can interpret results correctly.

## The query
```text
staging parity checks before agent experiments
```

## Use this when
- AI agents run experiments or tests in staging
- Staging results do not reproduce in production (or vice versa)
- Building agent testing workflows
- Deciding what staging is for

## Not for when
- Testing in production (different risk calculus)
- Load testing (needs its own environment design)
- Staging environment provisioning

## Steps

### Step 1: Define the parity dimensions that matter
List what must match for the experiment type: versions, configuration, data shape and volume, feature flags, third-party integrations (or their mocks). Not everything matters equally; a UI experiment needs different parity than a database migration test.
Expected output: a parity checklist scoped to the experiment.

### Step 2: Automate the parity check
Write a check that compares staging against production on the checklist dimensions and reports drift. Run it before every agent experiment session. Manual parity checks get skipped; automated ones get fixed.
Expected output: a pass/fail parity report, run automatically.

### Step 3: Document the known differences
Some differences are permanent (staging has less data, third parties are mocked). Write them down where the agent and operator see them before the experiment. Unknown differences invalidate results; known differences just scope them.
Expected output: a known-differences doc, current and visible.

### Step 4: Refresh staging from production regularly
Stale staging drifts silently: data ages, versions lag, flags diverge. Schedule regular refreshes (data snapshots, config syncs) so parity does not decay between checks.
Expected output: staging refreshed on a schedule, with drift bounded.

### Step 5: Gate experiments on parity
Do not start the agent experiment if parity fails on a dimension the experiment needs. A failed parity check is the experiment telling you it would have lied to you. Fix parity first.
Expected output: experiments run only on environments that can give trustworthy results.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_maftaBLVGlXvelcHVUINZQ
