httpx.ConnectTimeout during paginated crawl
Fixes httpx.ConnectTimeout in paginated crawls with explicit timeouts, semaphore-capped concurrency, and backoff retries. Use when pages fail on TCP connect, especially later in long crawls. Not for ReadTimeout after connecting or DNS failures.
TL;DR
httpx.ConnectTimeout means the TCP connection itself never completed, usually because the server is slow to accept or your concurrency is too high. Set an explicit connect timeout, cap concurrency with a semaphore, and retry with backoff.
httpx.ConnectTimeout during paginated crawlUse this when
- Paginated crawls with httpx fail on connection, not on reads.
- Timeouts cluster on later pages of a long crawl.
- The default 5-second timeout trips on slow hosts.
Not for this skill when
- The failure is ReadTimeout after connecting. That needs a read timeout, not a connect timeout.
- DNS fails. That is a resolver problem.
Steps
- Set explicit timeouts:
timeout = httpx.Timeout(connect=10.0, read=30.0, write=10.0, pool=10.0). Verify:httpx.AsyncClient(timeout=timeout)constructs without error. - Cap concurrency:
sem = asyncio.Semaphore(10)and wrap each page fetch inasync with sem:. Verify: in-flight requests never exceed 10. - Add retries with backoff on ConnectTimeout: sleep 2s, 4s, 8s across 3 attempts. Verify: a flaky host completes the page set.
- Reuse one AsyncClient for the whole crawl so connections pool. Verify: later pages connect faster than the first page.
- Run the full pagination. Verify: every page returns 200 and the collected items match the expected page count.
Variant phrasings
httpx connect timeout
Short form of the same error.
httpx timeout paginated requests
Task-level search for the same failure.
asyncclient connecttimeout crawl
The async-client variant.
Compatibility: httpx 0.24+. Timeout values are in seconds. AsyncClient and Client both accept the Timeout object.
Why it happens
The default httpx timeout is 5 seconds for everything, including the TCP handshake. Paginated crawls open hundreds of connections; under load the server's accept queue slows, handshakes exceed 5 seconds, and pages fail before a single byte is read. High concurrency makes it worse by competing for the same queue.
Edge cases / pitfalls
- A connect timeout on every page means the host is down or blocking you, not slow. Check with one manual request first.
- Retrying with the same aggressive concurrency just re-triggers the timeout. Lower concurrency before adding retries.
- Keep-alive helps: closing and reopening a client per page wastes a handshake every time.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst0vFsD5h1nemVxtD8fyvfw