my agent labeled intermittent DB deadlocks as flaky instead of filing the transaction ordering bug
Fixes agents that dismiss intermittent database deadlocks as flakes instead of filing the transaction-ordering bug. Use when parallel tests deadlock sporadically and the deadlock graph shows the same tables locked in different order by two transactions. Capture the deadlock report, enforce one consistent lock order in code, and keep retries as a backstop only. Not for deadlocks from an overloaded database, and not for single-threaded runs with no concurrency.
TL;DR
A deadlock is never a flake; it is two transactions grabbing locks in opposite order. Pull the deadlock graph from the database logs, find the two code paths that lock the same tables in different order, and make the order consistent everywhere. Retries are a bandage, not the fix.
The exact query
my agent labeled intermittent DB deadlocks as flaky instead of filing the transaction ordering bugSteps
- Stop calling it a flake and pull the evidence: grab the deadlock report from the database (postgres logs it with deadlock_timeout, mysql prints it in SHOW ENGINE INNODB STATUS). Identify the two transactions and the exact tables and rows each held and wanted.
Expected: You can name the two transactions, the tables, and the lock order each one used. The deadlock is now a diagram, not a mystery.
- Map both transactions back to code: find the two code paths (often two different tests, or a test and a fixture) that touch the same tables. Write down the order each one acquires its locks.
Expected: You see the inversion plainly, for example path A locks users then orders, path B locks orders then users.
- Enforce one consistent lock order across the whole codebase: pick a canonical table order (alphabetical is fine) and rewrite the offending path to acquire locks in that order. Add a code comment naming the deadlock so the next person does not reintroduce it.
Expected: Both paths now lock tables in the same order. The deadlock is structurally impossible, not just unlikely.
- Reduce the lock footprint while you are there: keep transactions short, avoid locking rows you only read (use the right isolation level), and never hold a DB transaction open across a network call or a sleep in tests.
Expected: Transactions hold fewer locks for less time, so even untested paths are less likely to collide.
- Only after the ordering fix, add a bounded retry with backoff as a backstop for deadlocks you have not found yet. Log every retry so a new deadlock pattern shows up in monitoring instead of hiding behind retries.
Expected: Retries exist but almost never fire. If they start firing, the log tells you where the next ordering bug is.
Use this when
- Tests deadlock intermittently, especially under parallel runners
- The deadlock graph shows the same tables locked in different order by two transactions
- An agent keeps retrying or labeling the deadlock flaky instead of investigating
- Two tests or a test plus fixture touch the same tables concurrently
Not for this skill when
- Deadlocks come from a genuinely overloaded or undersized database (fix capacity first)
- Tests run single-threaded with no concurrency (then it is not a real deadlock, look elsewhere)
- The deadlock involves an external system you cannot reorder (then retries plus idempotency are the actual fix)
Variant phrasings
intermittent deadlock in parallel pytest runs
Same fix. Parallelism just makes the ordering inversion show up more often.
postgres deadlock detected in CI but never locally
Local runs are usually serial, CI runs parallel. The ordering bug was always there; CI just rolls the dice more often.
mysql deadlock on the same tables every night
If it is the same tables on a schedule, it is two jobs colliding. Same lock-order fix, applied to the jobs.
Why it happens
Databases detect the circular wait and kill one transaction, which surfaces as an intermittent error that looks exactly like a flake to an agent reading test output. The root cause is always in application code: two paths acquiring locks on the same resources in opposite order. Parallel test runners raise the collision odds, so the bug appears "in CI only" and the agent reaches for the flake label instead of reading the deadlock graph.
Edge cases
- Some deadlocks come from gap locks or next-key locks on ranges, not just row order. If the graph shows range locks, the fix may be a different index or isolation level rather than lock order.
- Retries without idempotency can double-apply writes. Make the retried transaction idempotent or the retry becomes a data-corruption bug.
- If the two colliding paths live in different services, you cannot enforce one lock order in code. Then the fix is a saga or a single-writer design, and the retry needs a proper backoff with jitter.
- Deadlock graphs get truncated in logs. Raise the log verbosity for the deadlock detector temporarily rather than guessing from a partial graph.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_IvbVDOjYjxgruBLCH55feQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.