## TL;DR
Delete the sleep. A fixed sleep only makes the race window smaller; it never closes it, which is why the test failed again the next day. Wait on the actual condition instead (poll until the state you need is true, with a timeout), and build retries around idempotent operations with backoff. The test becomes deterministic instead of lucky.

## The exact query
```text
agent patched a race condition with sleep(2) and it failed again the next day  -  proper fix for the retry loop
```

## Steps
1. Remove the sleep and reproduce the race honestly: run the test in a loop (20 to 50 iterations) or under load until it fails without the sleep. You need to see the real failure mode, not the slept-over version.
   Expected: The test fails intermittently without the sleep, and you can describe what it was racing (element not ready, record not yet written, message not yet consumed).
2. Identify the exact condition the test needs: not "2 seconds have passed" but "the row exists in the database", "the element is visible and enabled", "the queue is drained". Name the observable state.
   Expected: A one-sentence condition like "the test needs the order row to be committed before it asserts".
3. Replace the sleep with a polling wait on that condition: loop with a short interval (100 to 250ms), check the condition, give up after a real timeout (5 to 30s depending on the operation), and fail with a message naming the condition that never became true.
   Expected: The test waits exactly as long as needed, no longer. Slow CI waits longer, fast local runs do not wait at all.
4. If the flaky part is a retry loop (not just a wait), make each attempt idempotent and add backoff: the operation must be safe to run twice, retries use exponential backoff with jitter, and the loop gives up after a bounded number of attempts with the last error attached.
   Expected: Retries converge instead of hammering. A retry that fails 3 times surfaces the real error instead of timing out mysteriously.
5. Prove the race is closed: run the suite 50 times (or run the test under CPU throttling to widen timing windows). A proper condition-based wait passes under throttling; a sleep-based patch fails as soon as the machine is slow enough.
   Expected: Green across the loop and under throttled CPU. The "failed again the next day" pattern is gone.

## Use this when
- A test was patched with sleep (or waitForTimeout) and failed again later
- The test races something: UI rendering, async writes, message queues, eventual consistency
- You are reviewing an agent's fix and see a magic-number sleep
- A retry loop hammers without backoff or without idempotency

## Not for this skill when
- The wait is for a genuinely fixed duration (a 500ms CSS animation, a token that expires in 60s); polling cannot beat the clock there
- The race is in production code, not the test (then the fix is real synchronization: locks, transactions, or ordering guarantees, not test waits)
- The sleep is in a backoff sequence that already polls a condition between sleeps

## Variant phrasings
### replaced sleep with wait but the test still flakes
Then the wait is watching the wrong condition (for example, element visible but not yet hydrated). Find the true readiness signal.

### how to fix a race condition in a test without sleep
Poll the actual condition with a timeout. That is the whole technique.

### waitForTimeout flake in playwright
waitForTimeout is a sleep with a nicer name. Replace it with an assertion that auto-retries on the real condition (expect.poll, waitForSelector on the right state).

## Why it happens
Fixed sleeps assume the world runs at a constant speed. CI does not: a loaded runner, a throttled CPU, or a slow network stretches every timing assumption, and the 2 seconds that was generous on Tuesday is not enough on Wednesday. The sleep did not fix the race; it just made losing the race rarer. Condition-based waits do not care how slow the machine is, because they wait for the state, not the clock.

## Edge cases
- Polling a condition that never becomes true turns a flake into a slow timeout. Always bound the wait and make the timeout message name the missing condition so the next debugger knows what to look for.
- Some conditions are not directly observable (a background job with no status endpoint). Then the fix is to make it observable (a status flag, a test hook), not to sleep longer.
- In browser tests, "visible" is not always "ready". An element can be visible before its event handlers attach. Wait for the actionable state (enabled, stable, hydrated), not just visibility.
- If the operation under retry is not idempotent (it charges a card, sends an email), retries are dangerous regardless of backoff. Fix idempotency first, then retry.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_NZpwEy7VI2QkLU_5liy7Tw
