agent profiled on a 16-core VM but production runs on 2-core containers -- the GIL contention never shows up in the test
Explains why GIL contention invisible on a 16-core test VM crushes a 2-core production container. Use when Python services benchmark fine on big machines but stall in small containers. Key trigger: production runs on few cores while tests ran on many.
GIL contention invisible on 16-core test VM, fatal on 2-core production containers
TL;DR: Benchmark on the same core count as production, or scale the expectation: CPython threads do not parallelize CPU-bound work, and fewer cores make the GIL fight worse. Profile with production's CPU limit, find the CPU-bound sections, and move them to processes or native code.
agent profiled on a 16-core VM but production runs on 2-core containers -- the GIL contention never shows up in the testSteps
- Reproduce under the real constraint: cap the benchmark to 2 cores and rerun the same workload:
taskset -c 0,1 python -m your_benchmark
# or in containers: docker run --cpus=2 your-image your_benchmarkExpected: throughput collapses and latency climbs toward production numbers - the contention is now visible.
- Confirm the GIL is the bottleneck, not just "fewer cores": profile with a tool that shows thread states (py-spy with --gil flag, or Austin):
py-spy record --gil -o gil-profile.svg -- python your_service.pyExpected: threads spend most of their time waiting on the GIL rather than executing - the classic contention signature.
- Find the CPU-bound sections the threads fight over: look for pure-Python loops, serialization, and regex in the hot path of the profile.
Expected: a short list of functions consuming the GIL-held CPU time.
- Fix it with one of the standard moves, cheapest first: move CPU-bound work to a process pool (multiprocessing) so each worker gets its own GIL; push the hot loop into numpy/C extensions that release the GIL; or run more single-threaded container replicas instead of threads.
Expected: per-request latency drops and scales with core count again.
- Re-run the capped benchmark after the fix.
Expected: the 2-core number now approaches the per-core throughput the 16-core VM showed - contention resolved, not just masked.
- Pin the agent's benchmark environment to production's CPU and memory limits (read them from the deployment manifest, not the dev machine) and require the core count in every benchmark report.
Expected: environment mismatch becomes impossible to overlook.
Use this when
- A Python service is fast on a big dev/test machine but slow in small production containers.
- Adding threads does not improve throughput, or makes it worse.
- Profiles show high CPU with threads mostly waiting.
- You need to prove the GIL is the limiter before re-architecting.
Not for this skill when
- The workload is I/O-bound (network, disk, database waits) - threads help there and the GIL is not the problem.
- You run PyPy, GraalPy, or a free-threaded CPython build - different GIL behavior, different investigation.
- Latency is bad on the 16-core box too - the problem is algorithmic, not contention.
Variant phrasings
- "python threads slower with more threads GIL contention"
- "service fast on VM slow in kubernetes 2 cpu"
- "py-spy shows threads waiting on GIL"
- "benchmark on 16 cores but production has 2"
Why it happens
CPython's Global Interpreter Lock lets only one thread execute Python bytecode at a time. On 16 cores, threads still serialize on the GIL, but there is enough spare CPU that context switching and lock handoff stay cheap - contention hides. On 2 cores, the same threads fight over the same lock with no spare capacity: handoff overhead dominates, cache lines bounce, and throughput falls off a cliff. The agent benchmarked the algorithm; production experiences the algorithm plus the lock.
Edge cases
- C extensions that release the GIL (numpy, I/O, some crypto) scale fine with threads - only pure-Python CPU work contends.
- Multiprocessing trades GIL contention for IPC and memory overhead; it wins for chunky CPU tasks, loses for tiny ones.
- Container CPU throttling (cfs quota) can mimic GIL contention - check throttled-time metrics before concluding it is the lock.
- Free-threaded Python (3.13+ no-GIL builds) removes this specific failure but introduces thread-safety work; do not assume it is a drop-in fix.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstySpeIHIenqK1qGwexFo-w
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.