# PerimeterX is blocking the press release archive crawl

## TL;DR
PerimeterX blocking a press release archive means the publisher put bot protection on pages that scrapers hit hardest, and the timeout is the challenge never completing for an automated client. Press releases are syndicated by design, so the same text lives on wire services, the company's IR site, and EDGAR 8-K filings. Stop waiting on the challenge, treat the archive as a defended source, and pull the releases from those sanctioned copies with deduping.

## The error
```text
HTTP 403 / request timeout
PerimeterX bot-check interstitial on press release archive; crawler times out waiting for the challenge to clear
```

## When this helps
- a press archive crawl times out on a PerimeterX interstitial
- announcement coverage has gaps from one defended publisher
- choosing canonical sources for press release intake
- a briefing agent wastes its time budget waiting on challenges

## When it doesn't
- you want to pass the PerimeterX check; the syndicated copies make that unnecessary
- the release is exclusive to the defended archive; ask the publisher for a feed
- you need the archive's page metadata rather than the release text

## Works with
curl 7.x+, python 3.8+. PerimeterX behavior is vendor-side and changes without notice.

## Steps
### 1. Confirm the PerimeterX interstitial and cap the wait
```bash
curl -s -o px.html --max-time 20 -A "IntelBriefingBot/1.0" "https://YOUR-publisher/press-releases"
grep -il "perimeterx" px.html && echo "perimeterx interstitial confirmed"
```
Expected: Confirmation within 20 seconds. A longer timeout never helps; the interstitial does not clear for scripted clients.

### 2. Check the publisher's own distribution channels
```bash
curl -s "https://YOUR-publisher/press-releases/rss" -o pr.xml && head -3 pr.xml
curl -s "https://YOUR-company/investors/news" -o ir.html -w "HTTP %{http_code}\n" | head -2
```
Expected: An RSS feed or the company's IR news page. Publishers that defend archives usually still push releases to IR sites and wires.

### 3. Pull the release from a wire or EDGAR copy
```bash
curl -s "https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=[cik]&type=8-K&dateb=&owner=include&count=10" | grep -o "8-K" | head -5
```
Expected: Recent 8-K filings carrying the same announcements. Material releases must be filed, which makes EDGAR the canonical backup source.

### 4. Dedupe releases across the syndicated copies
```python
import hashlib
def fp(company, date, headline):
    return hashlib.md5((company.strip().lower() + date + headline.strip().lower()[:60]).encode()).hexdigest()
print(fp("Acme Corp", "2026-10-08", "Acme launches widget v2"))
print("same key for wire, IR, and archive copies of one release")
```
Expected: One dedupe key per release across sources. The briefing counts each announcement once no matter how many sites carry it.

## Other ways people phrase this
### perimeterx blocking crawl timeout fix
The timeout is the challenge not clearing. Shorter timeouts plus a source switch beat longer timeouts.

### press release archive bot protection
Archives are defended because scrapers hammer them. The releases are public by design and syndicated, so defend the intake differently.

### px interstitial press page scraper
Shorthand for the same block. Treat the domain as defended for archive paths and re-source.

## Why it happens
PerimeterX fingerprints the client and serves suspected bots an interstitial that automated clients cannot complete. Press archives attract constant scraping, so publishers defend them while still distributing the actual releases through wires, IR sites, and regulatory filings. The crawler is fighting for the worst copy of public content.

## Edge cases
- Some publishers defend only the archive listing while release pages stay open; test one direct release URL before writing off the domain.
- Wire copies sometimes trim boilerplate; for legal precision cite the IR or EDGAR copy.
- Release timestamps differ across syndications by minutes; normalize on the company's own timestamp.
- If the publisher offers an email alert or API for releases, prefer it over any crawl.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_eegRv0HhfkbsF3OkddksVA
