intel agent failed deduping syndicated press releases
This skill fixes intel agents that fail deduping syndicated press releases. Use it when briefings repeat announcements or when building syndication deduping. It is not for genuinely different stories; the fix is content-fingerprint dedupe keys, canonical copies, and output validation.
Intel agent failed deduping syndicated press releases
TL;DR
Syndicated press releases defeat naive deduping because every wire edits the headline slightly and wraps the body differently, so URL or exact-text matching misses the duplicates. Dedupe on normalized content: company plus date plus the first paragraph's fingerprint. Keep one canonical copy per release and record the other URLs as duplicates, not as separate stories.
The error
(duplicate stories in briefing)
intel agent failed deduping syndicated press releases; same announcement appears 4 timesWhen this helps
- briefings show the same announcement multiple times
- press release intake pulls from several wires
- building deduping for syndicated content
- validating briefing uniqueness
When it doesn't
- the duplicates are genuinely different announcements; check dates before merging
- you need every wire's wording; then dedupe is the wrong goal
- releases differ by material updates; version them instead of merging
Works with
python 3.8+ with hashlib and re. No external dependencies.
Steps
1. Normalize each release to a dedupe key
import hashlib, re
def rkey(company, date, text):
first = re.sub(r"\s+", " ", text.strip().lower())[:200]
return hashlib.md5((company.strip().lower() + date + first).encode()).hexdigest()
print(rkey("Acme Corp", "2026-10-08", "Acme launches widget v2 today."))Expected: A stable key per release. Headline edits do not affect the first-paragraph fingerprint.
2. Cluster candidate duplicates before the briefing
import hashlib, re
def rkey(company, date, text):
first = re.sub(r"\s+", " ", text.strip().lower())[:200]
return hashlib.md5((company.strip().lower() + date + first).encode()).hexdigest()
releases = [("Acme", "2026-10-08", "Acme launches widget v2 today."), ("Acme", "2026-10-08", "Acme launches widget V2 today!")]
seen = {}
for r in releases:
seen.setdefault(rkey(*r), []).append(r)
print("clusters:", len(seen), "for", len(releases), "releases")Expected: One cluster per real announcement. Near-duplicate wordings collapse to a single key.
3. Keep the best copy and record the rest as duplicates
import json
canonical = {"key": "abc123", "url": "https://ir.YOUR-company/release", "dupes": ["YOUR_WIRE_A_HOST/rel", "YOUR_WIRE_B_HOST/rel"]}
open("releases.json", "w").write(json.dumps(canonical, indent=2))
print("one canonical copy; duplicates recorded, not briefed")Expected: A canonical record. The briefing uses the IR or wire copy with the cleanest text.
4. Validate deduping on the briefing output
import json
items = [{"key": "a"}, {"key": "a"}, {"key": "b"}]
keys = [i["key"] for i in items]
print("dupes remaining:", len(keys) != len(set(keys)))Expected: A boolean check. Run it before publishing; any True means the deduper missed a variant.
Other ways people phrase this
syndicated press releases duplicates briefing
Normalize on company plus date plus first paragraph, not URL or headline.
press release dedup failed agent
Wires edit headlines; fingerprint the body instead.
same announcement multiple times intel
Cluster before briefing. One canonical copy per release.
Why it happens
Press releases are syndicated to dozens of outlets that each tweak the headline and template, so exact-match deduping fails. The release's identity is company plus date plus content, not URL. Content fingerprinting on normalized text catches the variants that string matching misses.
Edge cases
- Updates to a release (corrected versions) share the fingerprint; version them by wire timestamp.
- Two companies can announce similar things the same day; the company in the key prevents cross-merges.
- Very short releases fingerprint poorly; require a minimum text length.
- Translations of the same release need language-aware normalization.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_AM2YbL8t8QGj8bhL-hp2Yw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.