# Press release site 403 after the JavaScript check

## TL;DR
A 403 right after a JavaScript check means the site's bot defense decided your fetcher is not a real browser, and press release archives are protected this way because they get scraped hard. Do not try to pass the check; press releases are the most widely syndicated content on the web, so the same release exists on the wire services, the company's IR site, and news APIs. Point the agent at those sanctioned copies and the 403 stops mattering.

## The error
```text
HTTP 403 Forbidden
(request blocked after JavaScript challenge check; press release archive page not served)
```

## When this helps
- a press archive fetch returns 403 after a JS check
- briefing coverage of company announcements has gaps
- the same release appears on multiple sites and needs deduping
- choosing between scraping an archive and using wire copies

## When it doesn't
- the goal is beating the JavaScript check itself; use the syndicated copies instead
- the release exists only on that one protected page and nowhere else, then ask the site owner
- you need the page's comments or engagement data, which wires do not carry

## Works with
curl 7.x+, python 3.8+ with requests. Wire service pages and EDGAR are stable public endpoints.

## Steps
### 1. Verify the block is the JS check and not a bad URL
```bash
curl -s -o check.html -w "HTTP %{http_code}\n" -A "IntelBriefingBot/1.0" "https://YOUR-site/press-releases"
grep -il "challenge\|captcha\|javascript" check.html || echo "no challenge marker"
```
Expected: HTTP 403 plus a challenge marker in the body. A 404 instead means your archive URL is simply wrong; fix the URL before anything else.

### 2. Look for the site's own feed or sitemap first
```bash
curl -s "https://YOUR-site/press-releases/rss" -o pr.xml && head -3 pr.xml
curl -s "https://YOUR-site/sitemap.xml" | grep -i press | head -10
```
Expected: An RSS feed or sitemap entries for the releases. Many protected archives still publish a feed; subscribe to it instead of scraping the pages.

### 3. Pull the same release from a wire service copy
```bash
curl -s "https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=[cik]&type=8-K" | head -5
```
Expected: Material press releases land in 8-K filings on EDGAR, and most releases are syndicated to wire services with public pages. One canonical copy beats fighting the archive's bot defense.

### 4. Dedupe by release, not by source URL
```python
import hashlib
releases = ["Acme launches widget v2", "Acme launches Widget V2 "]
seen = set()
for r in releases:
    fp = hashlib.md5(r.strip().lower().encode()).hexdigest()
    if fp not in seen:
        seen.add(fp)
        print("new:", r.strip())
```
Expected: One entry per release even when three sources carry it. Syndicated copies normalize to the same key, which keeps the briefing from triple-counting.

## Other ways people phrase this
### 403 on press release archive crawl
Short form. Archives are bot-defended because scrapers hammer them; the releases themselves are public by design and syndicated everywhere.

### javascript check failed scraper press page
The defense fingerprinting the client. Headless or scripted clients fail it by design; that is not a bug you fix, it is a signal to change sources.

### press release fetch blocked but feed works
Common split: the HTML archive is protected while the RSS feed is open. Prefer the feed permanently; it is also faster and cleaner to parse.

## Why it happens
Press archives sit behind JavaScript bot checks because they are scraped aggressively and the HTML pages are expensive to serve. The check fingerprints the client for browser-like behavior, which scripted fetchers fail on purpose. Meanwhile the press release text itself is distributed to wires, regulators, and IR sites by the company, so the defended archive is the worst copy to chase.

## Edge cases
- Some archives protect the listing pages but not individual release pages; test one release URL directly before abandoning the source.
- Wire copies can lag the original by minutes to hours; for market-moving releases, the company's IR site or EDGAR 8-K is the fastest canonical source.
- Deduping by headline fails when wires edit headlines; normalize on release date plus company plus first paragraph instead.
- A few IR sites gate archives behind logins; do not script logins, use their public filings and wire copies.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_KwQQOMwRSlVFnr7URBc5hg
