TL;DR
A flake diagnosed with 8 workers has a totally different timing profile than CI's 2-worker shards, so the agent tuned the wrong thing. Reproduce at CI's worker count and the real flake shows up.

```text
agent diagnosed a flake locally with --workers=8 but CI shards with --workers=2, so the timing profile was completely different
```

## Steps
1. Find CI's real worker count. Check the CI config for the shard or `--workers` flag (here it is 2). Do not trust your local default.
   Expected: you can quote the exact worker count CI uses for that shard.
2. Rerun locally with CI's count. Run the suite the same way CI does, with `--workers=2` and the same sharding flags.
   Expected: the flake reproduces, or at least the timing profile (per-test durations) looks different from your 8-worker run.
3. Re-diagnose from the slow profile, not the fast one. Look for shared-resource contention, serialized DB access, or startup races that only appear when tests queue behind each other.
   Expected: you find a contention point (e.g. tests waiting on a single DB connection) that 8 workers hid by finishing fast enough.
4. Fix the contention, not the worker count. Options: isolate the shared resource per worker, add proper readiness waits, or mark the truly serial tests so they do not share a worker.
   Expected: the suite passes at both 2 and 8 workers without retries.
5. Lock the worker count in CI config and document it for the agent's future runs.
   Expected: the agent's runbook records the exact CI command, so it never diagnoses at the wrong count again.

## Use this when
- the agent diagnosed locally at a different parallelism than CI shards at
- the flake vanishes when you add workers and returns when you remove them
- timing-sensitive tests behave differently between local and CI runs
- you suspect a race that only shows under serialized or contended execution

## Not for this skill when
- the failure happens at every worker count including 1
- the problem is a hard resource limit (memory, disk), not timing
- local and CI already use the same worker count
- the flake is order-dependent rather than timing-dependent

## Variant phrasings
- "flake only reproduces with fewer workers than my machine uses"
- "CI shards with different parallelism and the timing profile changed"
- "test passes with many workers, fails with few workers"

## Why it happens
Worker count changes how tests interleave in time. With 8 workers, CPU-bound tests finish before contended resources get stressed; with 2, tests queue, share connections longer, and hit timeouts and races that the fast profile never showed. Diagnosing at the wrong count tunes timeouts and retries for a world that does not exist in CI.

## Edge cases
- Some runners also throttle CPU per worker, so matching the count still is not a perfect replica; match the runner size too if you can.
- Playwright and jest shard differently (test-level vs file-level); copy CI's exact flags, not just the count.
- A test that only passes with MORE workers usually has a different bug (a hidden dependency on fast execution); do not just bump workers in CI and call it fixed.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_JsXr3TzgbtsqEmcM4c2VOA
