# Cloudflare Turnstile stuck in a loop on the headless scraper

## TL;DR
Turnstile looping in a headless browser means the challenge keeps failing your automated client, which is the protection working exactly as designed. There is no legitimate setting that makes a scripted browser pass; the loop is the signal to abandon that page as a source. Get the data from the site's API, feed, or sitemap, or run the fetch from a real browser session only where the site owner permits it, and design the agent to treat challenge loops as a hard stop, not a retry condition.

## The error
```text
(no HTTP error; browser automation never leaves the challenge)
Turnstile widget: "Verifying you are human..." loops indefinitely in headless Chromium
```

## When this helps
- a headless browser spins on a Turnstile widget
- a briefing agent hangs on a protected page instead of failing fast
- auditing which sources in the registry are actually fetchable
- deciding whether a Turnstile-protected page is worth any further effort

## When it doesn't
- you want the scripted browser to pass the challenge; that is what Turnstile exists to prevent
- the site offers no feed or API and forbids scraping, then the data is simply unavailable
- you are tempted to farm the challenge to a human-solving service; do not

## Works with
Any browser automation (Playwright, Puppeteer, Selenium) and curl for the HTTP check. Turnstile behavior is Cloudflare-side.

## Steps
### 1. Detect the loop in the automation and bail out fast
```python
# pseudo-check for your browser agent: inspect page text after load
page_text = "Verifying you are human"  # from page.content()
if "verifying you are human" in page_text.lower():
    print("turnstile challenge detected: aborting page, marking source defended")
```
Expected: A clear defended-source verdict within seconds. Letting the loop spin wastes the run's time budget and never resolves.

### 2. Confirm the same URL is defended for plain HTTP too
```bash
curl -s -o ts.html --max-time 20 -A "IntelBriefingBot/1.0" "https://YOUR-site/[page]"
grep -il "turnstile\|challenge" ts.html && echo "defended for curl as well"
```
Expected: Whether the defense is browser-specific or blanket. Blanket defense means no fetch strategy reaches the page; move on.

### 3. Switch the agent to the site's machine-readable source
```bash
curl -s "https://YOUR-site/sitemap.xml" | head -20
curl -s "https://YOUR-site/feed" -o feed.xml && head -3 feed.xml
```
Expected: A sitemap or feed covering the same content. Turnstile-protected pages almost always have an API or feed the owner actually wants consumed.

### 4. Mark the source defended in the agent's source registry
```python
import json
reg = {"https://YOUR-site/[page]": {"status": "defended-turnstile", "fallback": "official feed"}}
open("sources.json", "w").write(json.dumps(reg, indent=2))
print(open("sources.json").read())
```
Expected: A persisted note so the next run skips the defended page immediately instead of rediscovering the loop.

## Other ways people phrase this
### turnstile verifying you are human loop scraper
The classic symptom text. The loop means repeated silent failures of the challenge; it will not clear with more waiting.

### headless browser stuck on cloudflare challenge
Headless fingerprints fail these checks by design. Non-headless or stealth flags do not change the verdict legitimately.

### cloudflare challenge never completes in automation
Same loop from the automation's perspective. Treat it as a defended source and re-source the data.

## Why it happens
Turnstile issues a challenge that validates the client is a real browser with human-like signals. Headless automation fails the validation silently and the widget retries, which looks like a loop. Cloudflare designed it so scripts cannot distinguish a solvable challenge from an unsolvable one; the loop is the unsolvable case.

## Edge cases
- Detecting the challenge fast matters more than any workaround; cap challenge waits at 15 to 30 seconds.
- Some pages challenge only certain paths; the registry note should be per-URL, not per-domain.
- Turnstile can also appear intermittently under load; a source that loops once may work later, so retry on the next scheduled run, not in a tight loop.
- Logging challenge HTML for debugging is fine; replaying challenge tokens is not.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_zYrdm5tikf0S6FyIl2l4rw
