# RSS feed XML has invalid characters, parser errors

## TL;DR
RSS parsers choke on invalid XML characters because feeds sometimes carry raw control characters or badly encoded text that XML forbids. The feed is not dead; it is dirty. Sanitize the bytes before parsing by stripping the illegal character ranges, and only then hand the cleaned XML to the parser. If the same feed is dirty every poll, report it to the publisher and keep the sanitizer in the pipeline permanently.

## The error
```text
xml.etree.ElementTree.ParseError: not well-formed (invalid token)
feedparser bozo_exception on invalid XML character in RSS feed
```

## When this helps
- RSS parsing fails with invalid character errors
- a feed parses on some polls and fails on others
- hardening a feed pipeline against dirty XML
- deciding whether a feed is broken or just dirty

## When it doesn't
- the feed returns HTTP errors; that is a fetch problem, not XML
- the XML is structurally broken; sanitizing characters will not fix bad nesting
- every feed fails; then the parser setup is wrong, not the feeds

## Works with
python 3.8+ with re and feedparser. XML character rules are spec-stable.

## Steps
### 1. Capture the raw bytes and find the bad characters
```bash
curl -s -A "IntelBriefingBot/1.0" "https://YOUR-publisher/feed" -o raw.xml -w "HTTP %{http_code}\n"
python3 -c "b=open("raw.xml","rb").read(); bad=[c for c in b if c // 32 == 0 and c not in (9,10,13)]; print("illegal bytes:", len(bad), set(bad[:10]))"
```
Expected: A count of illegal bytes. Control characters below 32 except tab, newline, and carriage return are invalid in XML.

### 2. Sanitize the bytes before parsing
```python
import re
raw = open("raw.xml", "rb").read().decode("utf-8", errors="ignore")
clean = re.sub("[^\x09\x0a\x0d\x20-\ud7ff\ue000-\ufffd\U00010000-\U0010ffff]", "", raw)
open("clean.xml", "w").write(clean)
print("cleaned chars:", len(clean))
```
Expected: A cleaned file. The regex keeps every legal XML character range and drops the rest.

### 3. Parse the cleaned feed and verify
```python
import feedparser
d = feedparser.parse("clean.xml")
print("bozo:", d.bozo)
print("entries:", len(d.entries))
```
Expected: No bozo flag and a real entry count. The sanitizer runs before the parser on every poll of this feed.

### 4. Pin the sanitizer to the dirty feed in config
```python
import json
cfg = json.load(open("feeds.json"))
cfg["https://YOUR-publisher/feed"] = {"sanitize_xml": True}
json.dump(cfg, open("feeds.json", "w"), indent=2)
print("sanitizer pinned for this feed")
```
Expected: A config flag. Dirty feeds stay dirty; the pipeline sanitizes them every time without manual intervention.

## Other ways people phrase this
### rss xml invalid character parse error
Control characters in the feed bytes. Strip the illegal ranges, then parse.

### feedparser bozo invalid token
The bozo flag with an invalid-token exception. Sanitize before parsing.

### xml not well-formed rss feed
Character-level dirt versus structural breakage. The byte check distinguishes them.

## Why it happens
XML 1.0 forbids most control characters, but feed generators built on sloppy string handling emit them anyway. Strict parsers reject the whole document on the first illegal byte. The content is fine; the bytes need cleaning, which is a preprocessing step, not a parser setting.

## Edge cases
- Sanitizing cannot fix mismatched tags or truncated feeds; those need the publisher.
- Some feeds declare the wrong encoding; try the declared encoding and UTF-8 before sanitizing.
- Log the byte positions of stripped characters; patterns help the publisher fix the generator.
- Keep the raw bytes for a day in case the sanitizer needs tuning.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_BWJlvrGh0LAjIfgewR4zJA
