# SEC filing exhibit 99.1 PDF fails to parse

## TL;DR
Exhibit 99.1 PDFs, usually press releases and earnings materials attached to 8-Ks, fail to parse for the same reasons any PDF fails: scanned images, broken font maps, or password protection. The exhibit is just a PDF with a fancy name. Download it from the filing index, diagnose it like any PDF with the text-versus-scanned check, and fall back to the 8-K's own text or the wire copy of the release when the PDF resists.

## The error
```text
(parse failure)
Exhibit 99.1 PDF parse error: empty text / garbled output / encrypted document
```

## When this helps
- exhibit 99.1 PDFs fail to parse
- earnings release attachments yield no text
- building 8-K exhibit intake
- deciding whether an exhibit PDF is worth OCR

## When it doesn't
- the exhibit is not a PDF; some are HTM and parse as HTML
- you need the PDF's exact layout; extraction gives content, not layout
- the exhibit is intentionally redacted; no parsing recovers redacted text

## Works with
EDGAR archives as of 2026; python 3.8+ with pdfminer.six and pymupdf.

## Steps
### 1. Download the exhibit from the filing index, not the viewer
```bash
curl -s -A "IntelBriefingBot/1.0" "https://www.sec.gov/Archives/edgar/data/[cik]/[accession]/" -o index.html
grep -o "[a-z0-9_-]*99\.1[a-z0-9_.-]*\.pdf" index.html | sort -u | head -5
```
Expected: The exhibit PDF filename. Viewer downloads sometimes wrap the file; the index link is the clean original.

### 2. Fetch it and check encryption first
```bash
curl -s -A "IntelBriefingBot/1.0" "https://www.sec.gov/Archives/edgar/data/[cik]/[accession]/[exhibit-pdf]" -o ex991.pdf -w "HTTP %{http_code}\n"
python3 -c "import fitz; d=fitz.open("ex991.pdf"); print("encrypted:", d.is_encrypted)"
```
Expected: HTTP 200 and the encryption flag. Encrypted exhibits are rare but unparseable until unlocked; most are simply unprotected.

### 3. Run the text-versus-scanned diagnosis
```python
from pdfminer.high_level import extract_text
text = extract_text("ex991.pdf", maxpages=2)
print("chars:", len(text.strip()))
import fitz
print("images on page 1:", len(fitz.open("ex991.pdf")[0].get_images()))
```
Expected: Character and image counts. Text present means a font or layout issue; images only means OCR is needed.

### 4. Fall back to the 8-K text or wire copy
```bash
curl -s -A "IntelBriefingBot/1.0" "https://www.sec.gov/Archives/edgar/data/[cik]/[accession]/[8k-doc]" -o filing.htm -w "HTTP %{http_code}\n"
python3 -c "print("8-K doc downloaded; its Item 2.02 text usually summarizes the exhibit")"
```
Expected: The 8-K document. Exhibit 99.1 content is typically summarized in the 8-K body and published to wires, so the PDF is rarely the only copy.

## Other ways people phrase this
### exhibit 99.1 pdf empty text
Run the scanned check. Earnings exhibits from small filers are often scanned.

### 8-k exhibit pdf parse error
Same PDF failure modes as any document. Diagnose generically, then fall back to the 8-K text.

### press release exhibit pdf garbled
Font map issues are common in exhibits. Try PyMuPDF, then OCR.

## Why it happens
Exhibit 99.1 is a naming convention, not a file format: it is whatever PDF the company attached, with all the quality variance that implies. Scanned exhibits, broken font maps, and occasional encryption cause the same failures as any PDF. The 8-K body and wire distributions carry the same content in cleaner form, which is why the PDF is the worst source to depend on.

## Edge cases
- Exhibit filenames vary; match on 99.1 in the name rather than a fixed pattern.
- Some exhibits exceed 100 pages of scanned images; OCR cost adds up, so check the 8-K summary first.
- Redacted exhibits parse fine but the redacted regions are gone; note the redaction in the briefing.
- Foreign private issuers attach 6-K exhibits with the same 99.1 convention.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_toag0311kn4yA9WkZejpBw
