flaky tests from parallel workers sharing ports: how to isolate
Shows how to stop parallel test workers from colliding on hardcoded ports: bind port 0 and read back the assigned port, or derive a per-worker port from the xdist worker id. Use when address-in-use failures appear only under -n and vanish on rerun. Not for a single shared external service, in-process fakes, or docker port mappings.
TL;DR
Stop hardcoding ports: bind to port 0 and read back the assigned port, or give each worker its own port derived from its worker id. Port collisions under -n auto are timing-dependent, which is exactly why the failures look flaky.
Problem
Parallel test workers each start a server on the same hardcoded port, so whichever worker binds second gets "address already in use" and its tests fail intermittently depending on scheduling order.
Steps
- Find the hardcoded port: search the test helpers for literal port numbers in
listenandbindcalls.
Expected: you have a list of every place a fixed port is used.
- Prefer the zero-port pattern: bind port 0 on an already-open socket and hand that socket to the server, so the OS assigns a free port with no race:
import socket
sock = socket.socket()
sock.bind(("", 0)) # OS picks a free port
port = sock.getsockname()[1]
start_test_server(sock) # server accepts on the already-bound socket(The test client then connects to the loopback interface at the discovered port.) Expected: two workers starting at the same moment no longer collide.
- Where port 0 is not possible (an external binary needs a number up front), derive a per-worker port from the worker id:
import os
worker = os.environ.get("PYTEST_XDIST_WORKER", "master")
index = 0 if worker == "master" else int(worker.replace("gw", ""))
port = 18000 + indexExpected: worker gw0 uses 18000, gw1 uses 18001, and reruns reproduce the same mapping.
- Run the suite with
-n 4twice and confirm zero "address already in use" failures.
Expected: two consecutive parallel runs are green.
When to use
- "Address already in use" failures that only happen with
-n. - Integration tests that boot real servers, databases, or browsers per worker.
- CI flakes that vanish on rerun with
-n 1.
When not to use
- A single shared external service all workers must use (serialize access or use separate schemas there, not separate ports).
- In-process fakes: no socket is bound, so ports are irrelevant.
- Dockerized services with published ports: map container ports to distinct host ports per worker instead.
Tool compatibility
- pytest 8.x with pytest-xdist 3.x (
PYTEST_XDIST_WORKERenv var); the port-0 bind pattern works on any language runtime.
Variant phrasings
address already in use only with pytest -n
Classic worker port collision; the per-worker port mapping above is the quick fix.
parallel tests fail intermittently binding the same port
Bind port 0 and read back the assigned port so no two workers ever ask for the same one.
Why it happens
A hardcoded port is a global resource. With N workers starting servers at the same time, the bind is a race: losers crash, and the loser changes run to run, which reads as flakiness.
Edge cases
- Port 0 plus a separate "discover then bind" step reintroduces the race; keep the socket bound and hand it to the server.
- Firewalls or CI sandboxes that restrict bind ranges: pick the base port inside the allowed range.
- TIME_WAIT sockets from a previous crashed run can still block a fixed port; port 0 avoids this entirely.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstnn8S-EjuKPmPINXB9y96g