how to test feature flags without flakiness
Tests feature flags deterministically: flag control in tests. Use when flags cause test variance. Not for flag rollout strategy.
TL;DR
Control flags explicitly in tests: set them via API, query params, or config before the test, and test both states. Never let tests depend on the ambient flag state.
Error
(Not an error; a determinism practice. Symptom: tests pass or fail depending on flag state.)Steps
- Inventory which flags each test depends on. Expected: the flag surface mapped.
- Set flags explicitly in test setup: API call, URL param, or env var. Expected: deterministic state.
- Test both on and off for flag-gated behavior. Expected: both paths covered.
- Isolate flag state per test; reset after each. Expected: no leakage.
- In CI, run the matrix for critical flags. Expected: combinations covered.
When to use
- Flag-dependent test variance.
- Testing flag-gated features.
When not to use
- Flag rollout or targeting strategy.
- Tests with no flag dependence.
Tool compatibility
- Any framework; flag providers (LaunchDarkly, Unleash, homegrown).
Variant phrasings
Feature flag test isolation
The core practice; explicit control.
Testing with flags on and off
The coverage goal; matrix the critical ones.
Why it happens
Flags are global mutable state. Tests that read ambient flags inherit whatever state the environment has.
Edge cases
- Flag evaluation caching needs cache-busting between tests.
- Percentage rollouts are non-deterministic by design; force 0% or 100% in tests.
- Stale flags (always on) should be removed, not tested forever.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst9ijl63ySH0JwBvYIAi4uA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.