# Intel agent failed deduping syndicated press releases

## TL;DR
Syndicated press releases defeat naive deduping because every wire edits the headline slightly and wraps the body differently, so URL or exact-text matching misses the duplicates. Dedupe on normalized content: company plus date plus the first paragraph's fingerprint. Keep one canonical copy per release and record the other URLs as duplicates, not as separate stories.

## The error
```text
(duplicate stories in briefing)
intel agent failed deduping syndicated press releases; same announcement appears 4 times
```

## When this helps
- briefings show the same announcement multiple times
- press release intake pulls from several wires
- building deduping for syndicated content
- validating briefing uniqueness

## When it doesn't
- the duplicates are genuinely different announcements; check dates before merging
- you need every wire's wording; then dedupe is the wrong goal
- releases differ by material updates; version them instead of merging

## Works with
python 3.8+ with hashlib and re. No external dependencies.

## Steps
### 1. Normalize each release to a dedupe key
```python
import hashlib, re
def rkey(company, date, text):
    first = re.sub(r"\s+", " ", text.strip().lower())[:200]
    return hashlib.md5((company.strip().lower() + date + first).encode()).hexdigest()
print(rkey("Acme Corp", "2026-10-08", "Acme launches widget v2 today."))
```
Expected: A stable key per release. Headline edits do not affect the first-paragraph fingerprint.

### 2. Cluster candidate duplicates before the briefing
```python
import hashlib, re
def rkey(company, date, text):
    first = re.sub(r"\s+", " ", text.strip().lower())[:200]
    return hashlib.md5((company.strip().lower() + date + first).encode()).hexdigest()
releases = [("Acme", "2026-10-08", "Acme launches widget v2 today."), ("Acme", "2026-10-08", "Acme launches widget V2 today!")]
seen = {}
for r in releases:
    seen.setdefault(rkey(*r), []).append(r)
print("clusters:", len(seen), "for", len(releases), "releases")
```
Expected: One cluster per real announcement. Near-duplicate wordings collapse to a single key.

### 3. Keep the best copy and record the rest as duplicates
```python
import json
canonical = {"key": "abc123", "url": "https://ir.YOUR-company/release", "dupes": ["YOUR_WIRE_A_HOST/rel", "YOUR_WIRE_B_HOST/rel"]}
open("releases.json", "w").write(json.dumps(canonical, indent=2))
print("one canonical copy; duplicates recorded, not briefed")
```
Expected: A canonical record. The briefing uses the IR or wire copy with the cleanest text.

### 4. Validate deduping on the briefing output
```python
import json
items = [{"key": "a"}, {"key": "a"}, {"key": "b"}]
keys = [i["key"] for i in items]
print("dupes remaining:", len(keys) != len(set(keys)))
```
Expected: A boolean check. Run it before publishing; any True means the deduper missed a variant.

## Other ways people phrase this
### syndicated press releases duplicates briefing
Normalize on company plus date plus first paragraph, not URL or headline.

### press release dedup failed agent
Wires edit headlines; fingerprint the body instead.

### same announcement multiple times intel
Cluster before briefing. One canonical copy per release.

## Why it happens
Press releases are syndicated to dozens of outlets that each tweak the headline and template, so exact-match deduping fails. The release's identity is company plus date plus content, not URL. Content fingerprinting on normalized text catches the variants that string matching misses.

## Edge cases
- Updates to a release (corrected versions) share the fingerprint; version them by wire timestamp.
- Two companies can announce similar things the same day; the company in the key prevents cross-merges.
- Very short releases fingerprint poorly; require a minimum text length.
- Translations of the same release need language-aware normalization.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_AM2YbL8t8QGj8bhL-hp2Yw
