# Press release PDF attachment fails to download from the IR site

## TL;DR
IR site PDF attachments fail to download when the link is JavaScript-generated, the file moved, or the server blocks non-browser clients. The release itself is almost never exclusive to that PDF: the same text lives on wire services and often in an 8-K exhibit. Diagnose the download with a plain HTTP fetch, and when the IR link resists, pull the release from the wire copy or EDGAR instead of fighting the site.

## The error
```text
(download failure)
press release PDF attachment failed to download from IR site: 404 / 403 / empty file
```

## When this helps
- press release PDFs fail to download from IR sites
- IR attachment links return 404 or 403
- announcement coverage has gaps from one IR site
- choosing canonical press release sources

## When it doesn't
- the PDF downloads but will not parse; that is a PDF parsing problem
- the release is exclusive to the IR PDF; then fix the download, usually a JS link
- you need the IR site's page metadata; the wire copy does not carry it

## Works with
curl 7.x+, python 3.8+. IR site behavior varies; EDGAR and wires are stable.

## Steps
### 1. Fetch the PDF URL directly and read the status
```bash
curl -s -L -A "IntelBriefingBot/1.0" "https://YOUR-company/investors/press-releases/[release].pdf" -o release.pdf -w "HTTP %{http_code}, %{size_download} bytes\n"
file release.pdf | head -2
```
Expected: HTTP 200 with a real PDF size. A 404 means the file moved; a 403 means the server blocks your client; tiny size means an error page saved as PDF.

### 2. Check whether the link is JavaScript-generated
```bash
curl -s -A "IntelBriefingBot/1.0" "https://YOUR-company/investors/press-releases" -o pr.html
grep -o "https\?://[^ \"]*\.pdf" pr.html | sort -u | head -10
```
Expected: Direct PDF URLs in the HTML. If none appear, the links are built by JavaScript and you need the underlying API or the wire copy.

### 3. Pull the same release from the wire or EDGAR
```bash
curl -s -A "IntelBriefingBot/1.0" "https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=[cik]&type=8-K&count=5" -o ek.html -w "HTTP %{http_code}\n"
grep -c "8-K" ek.html
```
Expected: Recent 8-K filings. Material releases are filed as 8-K exhibits, which gives you the same content from a stable URL.

### 4. Dedupe the release against other sources
```python
import hashlib
def rkey(company, date, headline):
    return hashlib.md5((company.strip().lower() + date + headline.strip().lower()[:60]).encode()).hexdigest()
print(rkey("Acme Corp", "2026-10-08", "Acme launches widget v2"))
print("same key for IR pdf, wire, and 8-K copies")
```
Expected: One dedupe key across sources. The briefing counts the announcement once regardless of which copy was fetched.

## Other ways people phrase this
### ir site pdf download failed
Check status first: 404 is a moved file, 403 is client blocking, tiny file is an error page.

### press release attachment 404
IR sites reorganize often. The wire copy or 8-K exhibit is the stable alternative.

### pdf link javascript generated ir site
Find the underlying API or use the syndicated copy; do not automate the JS.

## Why it happens
IR sites are marketing properties with fragile download links: JavaScript-generated URLs, reorganized archives, and bot blocking. The press release content is distributed to wires and regulators by design, so the IR PDF is one of several copies. Download failures are site issues, not content issues, which is why re-sourcing beats debugging the site.

## Edge cases
- Some IR PDFs are scanned images; the wire copy is text and parses better.
- Release dates on IR sites can differ from wire timestamps; normalize on the company's stated date.
- Login-walled IR portals are off-limits for scraping; use public filings.
- Always verify the downloaded bytes are a PDF, not an HTML error page.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_HgI1QT28Tk3t6eCMx9S2UA
