xbrl agent broke on taxonomy extension parse error
This skill fixes XBRL agents that break on taxonomy extension parse errors. Use it when agents extract zero facts or when building XBRL intake. It is not for corrupt instances; the fix is downloading the full file set from the filing index before parsing, with the companyfacts API as fallback.
XBRL agent broke on a taxonomy extension parse error
TL;DR
The XBRL agent breaks on taxonomy extensions for the same reason parsers do: the extension schema was not downloaded with the instance, so references dangle. The agent fix adds a pre-parse step that fetches the filing's full file set from the index before parsing. If the extension stays broken, the agent falls back to the companyfacts API for the same figures.
The error
(agent parse failed)
xbrl agent broke on taxonomy extension parse error; 0 facts extracted, briefing missing financialsWhen this helps
- an XBRL agent extracts zero facts
- agent pipelines break on extension taxonomies
- building robust XBRL intake for agents
- deciding parse versus API fallback
When it doesn't
- the instance itself is corrupt; re-download before blaming the extension
- you need extension calculation relationships; the API flattens those
- the filing predates XBRL; there is nothing to parse
Works with
python 3.9+ with arelle and requests; data.sec.gov companyfacts API as of 2026.
Steps
1. List the filing's full file set from the index
curl -s -A "IntelBriefingBot/1.0" "https://www.sec.gov/Archives/edgar/data/[cik]/[accession]/" -o index.html
grep -o "[a-z0-9_.-]*\.xsd\|[a-z0-9_.-]*_cal.xml\|[a-z0-9_.-]*_lab.xml" index.html | sort -uExpected: The extension schema and linkbase filenames. The agent needs this list before it fetches anything.
2. Download the instance plus all extension files
import requests
s = requests.Session()
s.headers.update({"User-Agent": "IntelBriefingBot/1.0"})
base = "https://www.sec.gov/Archives/edgar/data/[cik]/[accession]/"
for f in ["[instance].xml", "[cik]-2026.xsd", "[cik]-2026_cal.xml", "[cik]-2026_lab.xml"]:
r = s.get(base + f, timeout=30)
print(f, r.status_code)
open(f, "wb").write(r.content)Expected: All files on disk with 200s. The parse step never runs on a partial file set.
3. Parse with arelle and count extension facts
from arelle import Cntlr
cntlr = Cntlr.Cntlr()
model = cntlr.modelManager.modelXbrl
model.load("[instance].xml")
print("facts:", len(model.facts))
print("parse clean; extension resolved locally")Expected: A fact count. Local extension files let arelle resolve every reference.
4. Fall back to the companyfacts API on parse failure
curl -s -A "IntelBriefingBot/1.0" "https://data.sec.gov/api/xbrl/companyfacts/CIK[cik-padded].json" -o facts.json -w "HTTP %{http_code}\n"
python3 -c "import json; print("fallback facts ok:", "us-gaap" in json.load(open("facts.json"))["facts"])"Expected: API facts as fallback. The briefing gets numbers even when the extension defeats the parser.
Other ways people phrase this
xbrl agent extension parse error
Fetch the full file set first. The extension lives beside the instance.
agent 0 facts xbrl filing
Partial downloads cause this. Index, download all, then parse.
taxonomy extension agent broke
The API fallback keeps the briefing alive when the parse fails.
Why it happens
Agents that fetch only the instance XML leave the extension schema's relative references dangling, and arelle fails the parse. The extension files live in the same filing directory; the agent must enumerate and download them as a set. This is a fetch-completeness bug, not a parse bug.
Edge cases
- Some filings reference prior-year schemas; the index listing reveals the real set.
- Amended filings ship new extension files; never mix sets across accessions.
- Large extension sets slow arelle; the API fallback is faster for simple fact needs.
- Log which files were fetched per accession for debugging.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstZUlv92VNU9AR92SO19LPQ