# News site bot detection 403 and the scraper crashes on the challenge page

## TL;DR
Two problems are tangled here: the news site's bot defense served a 403 challenge page, and your scraper crashed because it assumed every fetch returns article HTML. Fix the crash first by making the fetcher validate responses before parsing, then fix the sourcing by moving news intake to official news APIs, publisher RSS feeds, or licensed wires. News is the most syndicated content online; a defended article page is never the only copy.

## The error
```text
HTTP 403 Forbidden
(bot detection challenge page served)
Traceback: parser crashed on unexpected HTML (challenge page instead of article)
```

## When this helps
- a news scraper crashes on unexpected HTML
- article fetches start returning 403s mid briefing
- hardening a briefing agent's fetch layer
- choosing between scraping news sites and using news APIs

## When it doesn't
- the goal is defeating the news site's bot detection; use the licensed copies instead
- the article is exclusive to the defended site with no syndication; then it is out of reach for scraping
- you need the site's comment section, which APIs do not carry

## Works with
python 3.8+ with requests; any news API (NewsAPI, GDELT, publisher APIs). Bot defenses are site-side.

## Steps
### 1. Make the fetcher validate before it parses
```python
import requests
s = requests.Session()
s.headers.update({"User-Agent": "IntelBriefingBot/1.0"})
r = s.get("https://YOUR-news-site/[article]", timeout=20)
if r.status_code != 200 or "challenge" in r.text[:2000].lower():
    print("blocked or unexpected, skipping parse")
else:
    print("ok, safe to parse", len(r.text))
```
Expected: No more crashes: blocked fetches are skipped with a log line instead of exploding the parser.

### 2. Confirm the defense and check for a feed
```bash
curl -s "https://YOUR-news-site/rss" -o feed.xml && head -3 feed.xml
curl -s "https://YOUR-news-site/robots.txt" | head -20
```
Expected: An RSS feed you can subscribe to, or robots rules showing the article paths are off-limits. Either answer tells you the sanctioned path.

### 3. Move article intake to a news API or wire
```bash
K="apiKey"
curl -s "https://newsapi.org/v2/everything?q=[topic]&${K}=${NEWSAPI_KEY}" | head -c 300; echo
```
Expected: A 200 with article JSON including the same story from a licensed copy. Publisher APIs and wires carry the text without the bot defense.

### 4. Add challenge detection to the briefing agent's fetch layer
```python
BLOCK_MARKERS = ["challenge", "captcha", "verify you are", "attention required"]
def fetch_ok(resp):
    body = resp.text[:3000].lower()
    return resp.status_code == 200 and not any(m in body for m in BLOCK_MARKERS)
print("fetch guard ready:", fetch_ok)
```
Expected: A reusable guard the agent calls before parsing any fetched page, so future defenses degrade gracefully instead of crashing the run.

## Other ways people phrase this
### news site 403 scraper bot detection
The defense half of the problem. News sites protect article pages because scrapers and AI crawlers hammer them.

### scraper crash on challenge page html
The crash half. Parsers that assume article-shaped HTML break on any interstitial; validate first, parse second.

### article fetch failed cloudflare news site
Cloudflare is the common defense vendor here. The same re-sourcing advice applies; the feed or API is the intended channel.

## Why it happens
News sites run bot detection because article pages are their most scraped asset. When the defense fires, the scraper receives challenge HTML instead of article HTML, and a parser written for article markup chokes on it. The crash is a missing validation step; the 403 is the site declining automated access to that page.

## Edge cases
- AMP or print versions of the article are sometimes undefended, but check robots and terms before relying on them.
- Paywalled articles need a subscription, not a scraper; briefings should cite the paywalled piece via its metadata and quote licensed summaries.
- News APIs have their own rate limits; cache article text and never re-fetch the same URL.
- Syndicated copies can differ slightly from the original; cite the copy you actually read.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_qoUrdvr-dE_JOZfMtIrGBA
