how to test with realistic fixture data
Creates realistic fixture data for tests: production-shaped, anonymized, minimal. Use when fixtures are too fake to catch bugs. Not for factory setup.
TL;DR
Make fixtures shaped like production: real field distributions, edge-case values, and valid relationships. Anonymize anything copied from prod, and keep fixtures minimal per test.
Error
(Not an error; a data quality practice. The failure it prevents: tests passing on toy data that fail on real data.)Steps
- Sample production shapes: field lengths, character sets, nullability. Expected: a realistic profile.
- Build fixtures matching the profile, including edge cases (unicode names, long strings, nulls). Expected: representative data.
- Anonymize any prod-derived data URIs names, emails, IDs. Expected: no PII in the repo.
- Keep fixtures minimal per test; compose from builders. Expected: readable tests.
- Refresh fixtures when the schema changes. Expected: no drift.
When to use
- Toy fixtures missing real bugs.
- Setting fixture standards.
When not to use
- Factory configuration mechanics.
- Load testing data (different scale).
Tool compatibility
- Any framework; fixture files or builders.
Variant phrasings
Realistic test data
The goal; production-shaped and anonymized.
Production-like fixtures
The same idea; mind the PII.
Why it happens
"test" and "foo" never contain unicode, nulls, or 200-character strings. Real data does, and bugs hide in the difference.
Edge cases
- Anonymization must be irreversible; hashing with salt, not masking.
- Keep a few full-size fixtures for performance-sensitive paths.
- Document the fixture profile so new tests follow it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_cLhx3bQ5ZrL6TckTbXz1MA