junit xml testsuite parsing errors in CI: how to fix
Fixes JUnit XML parsing errors in CI dashboards: malformed XML and encoding. Use when CI cannot parse test reports. Not for test failures themselves.
TL;DR
The XML is malformed: bad encoding, unescaped characters, or truncated writes. Validate the XML locally, fix the producer (encoding, buffering), and re-run.
Error
Failed to parse JUnit XML: not well-formed (invalid token) at line 412Steps
- Validate:
xmllint --noout TEST-*.xml. Expected: the exact line of the break. - Look at that line: usually a control character in a test name or output. Expected: the offending bytes found.
- Fix the producer: set UTF-8 encoding, strip control chars from test names. Expected: clean output.
- If truncated, the test process died mid-write; fix the crash first. Expected: complete files.
- Re-run and validate again. Expected: CI parses it.
When to use
- CI report step fails on XML parsing.
- Dashboards show missing results.
When not to use
- Tests fail (the XML is fine; the tests are not).
- You are choosing a report format (separate decision).
Tool compatibility
- Any JUnit XML producer; xmllint for validation.
Variant phrasings
JUnit XML not well-formed
The general error; validate to find the line.
Invalid token in TEST xml
Encoding or control characters.
Why it happens
Test output contains arbitrary bytes (logs, names). XML is strict; one bad byte breaks the whole file.
Edge cases
- Parallel writers to one file corrupt it; write per-worker files and merge.
--junitxmlpaths must exist; pytest does not create directories.- CDATA sections help but do not fix truly binary output.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst__9CFFVDKZnbnejYwa4CPRQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.