# XBRL agent broke on a taxonomy extension parse error

## TL;DR
The XBRL agent breaks on taxonomy extensions for the same reason parsers do: the extension schema was not downloaded with the instance, so references dangle. The agent fix adds a pre-parse step that fetches the filing's full file set from the index before parsing. If the extension stays broken, the agent falls back to the companyfacts API for the same figures.

## The error
```text
(agent parse failed)
xbrl agent broke on taxonomy extension parse error; 0 facts extracted, briefing missing financials
```

## When this helps
- an XBRL agent extracts zero facts
- agent pipelines break on extension taxonomies
- building robust XBRL intake for agents
- deciding parse versus API fallback

## When it doesn't
- the instance itself is corrupt; re-download before blaming the extension
- you need extension calculation relationships; the API flattens those
- the filing predates XBRL; there is nothing to parse

## Works with
python 3.9+ with arelle and requests; data.sec.gov companyfacts API as of 2026.

## Steps
### 1. List the filing's full file set from the index
```bash
curl -s -A "IntelBriefingBot/1.0" "https://www.sec.gov/Archives/edgar/data/[cik]/[accession]/" -o index.html
grep -o "[a-z0-9_.-]*\.xsd\|[a-z0-9_.-]*_cal.xml\|[a-z0-9_.-]*_lab.xml" index.html | sort -u
```
Expected: The extension schema and linkbase filenames. The agent needs this list before it fetches anything.

### 2. Download the instance plus all extension files
```python
import requests
s = requests.Session()
s.headers.update({"User-Agent": "IntelBriefingBot/1.0"})
base = "https://www.sec.gov/Archives/edgar/data/[cik]/[accession]/"
for f in ["[instance].xml", "[cik]-2026.xsd", "[cik]-2026_cal.xml", "[cik]-2026_lab.xml"]:
    r = s.get(base + f, timeout=30)
    print(f, r.status_code)
    open(f, "wb").write(r.content)
```
Expected: All files on disk with 200s. The parse step never runs on a partial file set.

### 3. Parse with arelle and count extension facts
```python
from arelle import Cntlr
cntlr = Cntlr.Cntlr()
model = cntlr.modelManager.modelXbrl
model.load("[instance].xml")
print("facts:", len(model.facts))
print("parse clean; extension resolved locally")
```
Expected: A fact count. Local extension files let arelle resolve every reference.

### 4. Fall back to the companyfacts API on parse failure
```bash
curl -s -A "IntelBriefingBot/1.0" "https://data.sec.gov/api/xbrl/companyfacts/CIK[cik-padded].json" -o facts.json -w "HTTP %{http_code}\n"
python3 -c "import json; print("fallback facts ok:", "us-gaap" in json.load(open("facts.json"))["facts"])"
```
Expected: API facts as fallback. The briefing gets numbers even when the extension defeats the parser.

## Other ways people phrase this
### xbrl agent extension parse error
Fetch the full file set first. The extension lives beside the instance.

### agent 0 facts xbrl filing
Partial downloads cause this. Index, download all, then parse.

### taxonomy extension agent broke
The API fallback keeps the briefing alive when the parse fails.

## Why it happens
Agents that fetch only the instance XML leave the extension schema's relative references dangling, and arelle fails the parse. The extension files live in the same filing directory; the agent must enumerate and download them as a set. This is a fetch-completeness bug, not a parse bug.

## Edge cases
- Some filings reference prior-year schemas; the index listing reveals the real set.
- Amended filings ship new extension files; never mix sets across accessions.
- Large extension sets slow arelle; the API fallback is faster for simple fact needs.
- Log which files were fetched per accession for debugging.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_ZUlv92VNU9AR92SO19LP_Q
