press release pdf attachment fails to download from ir site
This skill fixes press release PDF attachments that fail to download from IR sites. Use it when IR downloads 404 or 403, or when announcement coverage has gaps. It is not for PDFs that download but will not parse; the fix is diagnosing the download status and re-sourcing from wire services or EDGAR 8-K exhibits with deduping.
Press release PDF attachment fails to download from the IR site
TL;DR
IR site PDF attachments fail to download when the link is JavaScript-generated, the file moved, or the server blocks non-browser clients. The release itself is almost never exclusive to that PDF: the same text lives on wire services and often in an 8-K exhibit. Diagnose the download with a plain HTTP fetch, and when the IR link resists, pull the release from the wire copy or EDGAR instead of fighting the site.
The error
(download failure)
press release PDF attachment failed to download from IR site: 404 / 403 / empty fileWhen this helps
- press release PDFs fail to download from IR sites
- IR attachment links return 404 or 403
- announcement coverage has gaps from one IR site
- choosing canonical press release sources
When it doesn't
- the PDF downloads but will not parse; that is a PDF parsing problem
- the release is exclusive to the IR PDF; then fix the download, usually a JS link
- you need the IR site's page metadata; the wire copy does not carry it
Works with
curl 7.x+, python 3.8+. IR site behavior varies; EDGAR and wires are stable.
Steps
1. Fetch the PDF URL directly and read the status
curl -s -L -A "IntelBriefingBot/1.0" "https://YOUR-company/investors/press-releases/[release].pdf" -o release.pdf -w "HTTP %{http_code}, %{size_download} bytes\n"
file release.pdf | head -2Expected: HTTP 200 with a real PDF size. A 404 means the file moved; a 403 means the server blocks your client; tiny size means an error page saved as PDF.
2. Check whether the link is JavaScript-generated
curl -s -A "IntelBriefingBot/1.0" "https://YOUR-company/investors/press-releases" -o pr.html
grep -o "https\?://[^ \"]*\.pdf" pr.html | sort -u | head -10Expected: Direct PDF URLs in the HTML. If none appear, the links are built by JavaScript and you need the underlying API or the wire copy.
3. Pull the same release from the wire or EDGAR
curl -s -A "IntelBriefingBot/1.0" "https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=[cik]&type=8-K&count=5" -o ek.html -w "HTTP %{http_code}\n"
grep -c "8-K" ek.htmlExpected: Recent 8-K filings. Material releases are filed as 8-K exhibits, which gives you the same content from a stable URL.
4. Dedupe the release against other sources
import hashlib
def rkey(company, date, headline):
return hashlib.md5((company.strip().lower() + date + headline.strip().lower()[:60]).encode()).hexdigest()
print(rkey("Acme Corp", "2026-10-08", "Acme launches widget v2"))
print("same key for IR pdf, wire, and 8-K copies")Expected: One dedupe key across sources. The briefing counts the announcement once regardless of which copy was fetched.
Other ways people phrase this
ir site pdf download failed
Check status first: 404 is a moved file, 403 is client blocking, tiny file is an error page.
press release attachment 404
IR sites reorganize often. The wire copy or 8-K exhibit is the stable alternative.
pdf link javascript generated ir site
Find the underlying API or use the syndicated copy; do not automate the JS.
Why it happens
IR sites are marketing properties with fragile download links: JavaScript-generated URLs, reorganized archives, and bot blocking. The press release content is distributed to wires and regulators by design, so the IR PDF is one of several copies. Download failures are site issues, not content issues, which is why re-sourcing beats debugging the site.
Edge cases
- Some IR PDFs are scanned images; the wire copy is text and parses better.
- Release dates on IR sites can differ from wire timestamps; normalize on the company's stated date.
- Login-walled IR portals are off-limits for scraping; use public filings.
- Always verify the downloaded bytes are a PDF, not an HTML error page.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_HgI1QT28Tk3t6eCMx9S2UA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.