agent couldn't reproduce the deadlock -- it needs two specific cron jobs overlapping, which never happens in the test...
Fixes deadlocks that only reproduce when two specific cron jobs overlap in production. Use when database logs show 'deadlock detected' between two scheduled jobs that never overlap in the test window. Key trigger: the same pair of jobs is named in repeated deadlock detail lines.
TL;DR
A deadlock needs two transactions holding locks the other one wants, at the same time. If your two cron jobs only overlap in production on some nights, no ordinary test will ever catch it. Run both jobs concurrently against a staging copy on purpose, capture the deadlock detail, then fix the lock ordering so the overlap is harmless.
agent couldn't reproduce the deadlock -- it needs two specific cron jobs overlapping, which never happens in the test window- Confirm it is a deadlock and name both sides. Check the database logs for "deadlock detected" and read the DETAIL lines for both process IDs, the locks each holds, and the statements each was running. Expected: two PIDs, each waiting on a lock the other holds.
- Find the overlap. Look at cron or job-scheduler history for both jobs and find nights where their run windows intersect. Expected: job A and job B both run between, say, 02:00 and 02:20 on the same nights the deadlock appears.
- Reproduce it on purpose. Against a staging database with realistic data size, start job A, wait until it is mid-transaction, then start job B. Repeat a few times. Expected: the same "deadlock detected" error within a handful of attempts.
- Fix the lock order, not just the schedule. Make both jobs acquire the contested locks in the same order: sort writes by primary key before updating, or take the coarser lock first in both jobs. Re-run the concurrent test. Expected: the jobs overlap without deadlocking. Rescheduling alone just moves the collision.
Use this when
- Deadlock errors name the same pair of jobs in the log detail lines
- The failures cluster on nights or windows when both jobs run
- Single-job test runs never reproduce it
Not for this skill when
- The deadlock involves ad-hoc user traffic rather than two scheduled jobs: then look at lock ordering across the app code
- The error is a lock timeout, not a deadlock: that is contention, not a cycle
Variant phrasings
- "postgres deadlock detected only sometimes at night"
- "two cron jobs deadlock when they overlap"
- "cant reproduce deadlock in staging"
Why it happens
A deadlock is a cycle: job A locks row 1 and wants row 2, job B locks row 2 and wants row 1. That cycle only exists while both transactions are open at the same moment. Tests usually run one job at a time, so the cycle can never form, and the production overlap may be minutes per night.
Edge cases
- Foreign keys and unique indexes take locks too: the contested lock may be on an index, not the row you expect
- Retry logic that immediately reruns both jobs can turn one deadlock into a deadlock storm
- Lock ordering fixes must cover every code path that touches the same tables, including migrations and backfills
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_O93yEC-1z25EYwYLBS8ceQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.