# Akamai Bot Manager has the crawl stuck on a challenge

## TL;DR
Akamai Bot Manager challenging a crawl means the site's bot defense scored your traffic as automated, and the stuck challenge is the defense working, not a page that will eventually load. There is no legitimate client tweak that passes Bot Manager; it exists to stop exactly this. Treat the site as defended, keep one polite probe to confirm, and get the data from the site's API, feeds, sitemap, or a licensed vendor instead.

## The error
```text
HTTP 403 / challenge interstitial
Akamai Bot Manager challenge page; crawler stuck waiting for the challenge to resolve
```

## When this helps
- a crawl stalls on Akamai challenge pages
- a new source turns out to be Bot Manager defended
- hardening a crawler against challenge hangs
- evaluating whether a defended site is worth any crawl effort

## When it doesn't
- you want to pass Bot Manager's checks; that is what the product exists to prevent
- the site offers no API and forbids crawling; the data is out of reach
- you are considering residential proxies to look human; do not

## Works with
Any HTTP client. Akamai Bot Manager behavior is vendor-side and updates continuously.

## Steps
### 1. Confirm the Akamai challenge with one polite probe
```bash
curl -s -o ak.html --max-time 20 -A "IntelBriefingBot/1.0" "https://YOUR-site/[page]"
grep -il "akamai" ak.html && echo "akamai challenge confirmed"
```
Expected: One confirmation, then stop. Repeated probes against Bot Manager only worsen the client's reputation.

### 2. Check robots.txt and the site's API or feed options
```bash
curl -s "https://YOUR-site/robots.txt" | head -20
curl -s "https://YOUR-site/sitemap.xml" | head -10
```
Expected: The crawl rules and any machine-readable alternatives. Bot-managed sites often expose an API for the data they actually want consumed.

### 3. Detect the challenge fast in the crawl loop
```python
def is_challenge(status, body):
    t = body[:2000].lower()
    return status == 403 or "akamai" in t or "challenge" in t
print("challenge guard ready; skip defended pages immediately")
```
Expected: A guard that skips defended pages in milliseconds instead of hanging the crawl on challenges.

### 4. Re-source from the API, feed, or a licensed vendor
```bash
curl -s "https://YOUR-site/api/[resource]" -H "your auth header api key]" -o data.json -w "HTTP %{http_code}\n"
```
Expected: HTTP 200 from the sanctioned channel. The data the site wants distributed is available without fighting Bot Manager.

## Other ways people phrase this
### akamai bot manager blocking crawl
The general case. Bot Manager scores traffic continuously, so even a working fetch can start challenging later.

### scraper stuck on akamai challenge
The hang symptom. Cap challenge waits short and skip; the challenge does not resolve for automation.

### akamai 403 bot detection crawler
The refusal form. Confirm once, then re-source rather than retrying.

## Why it happens
Akamai Bot Manager profiles every client with browser, network, and behavioral signals, and serves automation a challenge it cannot pass. Sites buy it specifically to stop scrapers, so a stuck challenge is the product succeeding. No header, timing, or client tweak legitimately changes the verdict.

## Edge cases
- Bot Manager can challenge intermittently, so a source that works today may challenge tomorrow; keep the guard in the loop permanently.
- Mark defended status per path, not per domain; some sections of a defended site stay open.
- Challenge HTML cached as page content poisons downstream parsing; validate before storing.
- If the business need is ongoing, a licensed data vendor is cheaper than fighting the defense forever.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_y4LxTxdhV_hAyLdZhgn3mQ
