# Intel agent failed resolving the same company across data sources

## TL;DR
Company resolution fails because every source names companies differently: tickers, CIKs, legal names, brand names, and renames over time. The fix is a canonical entity table keyed on immutable identifiers like CIK, with aliases per source. Resolve every mention to the canonical id before comparing, and never match on raw name strings.

## The error
```text
(resolution failed)
intel agent failed resolving same company across data sources; 'Acme Inc' vs 'Acme Corp' vs ticker ACME treated as 3 companies
```

## When this helps
- an agent treats one company as several
- merging data across sources with different naming
- building entity resolution for briefings
- handling corporate renames

## When it doesn't
- the companies are genuinely different; similar names are not the same company
- you need probabilistic matching; curate aliases instead, it is more reliable
- the source has no identifier at all; then resolution is manual

## Works with
python 3.8+ with json. SEC CIK as the canonical key for US public companies.

## Steps
### 1. Build a canonical entity table on CIK
```python
import json
entities = {"0001234567": {"ticker": "ACME", "names": ["Acme Inc", "Acme Corp", "Acme Corporation"]}}
open("entities.json", "w").write(json.dumps(entities, indent=2))
print("canonical table keyed on CIK, the immutable identifier")
```
Expected: An entity table. CIKs do not change; names do, so the CIK is the key.

### 2. Resolve mentions through the alias table
```python
import json
entities = json.load(open("entities.json"))
def resolve(mention):
    m = mention.strip().lower()
    for cik, ent in entities.items():
        if m == ent["ticker"].lower() or m in [n.lower() for n in ent["names"]]:
            return cik
    return None
print(resolve("Acme Corp"), resolve("ACME"))
```
Expected: One CIK for every alias. Resolution happens before any comparison or aggregation.

### 3. Handle renames and ticker changes explicitly
```python
import json
entities = json.load(open("entities.json"))
entities["0001234567"]["names"].append("Acme Global")
entities["0001234567"]["former_tickers"] = ["ACME.OLD"]
json.dump(entities, open("entities.json", "w"), indent=2)
print("aliases grow over time; the CIK stays put")
```
Expected: An updated alias list. Renames add aliases; they never create new entities.

### 4. Validate resolution coverage on the briefing input
```python
import json
entities = json.load(open("entities.json"))
mentions = ["Acme Corp", "ACME", "UnknownCo"]
resolved = [m for m in mentions if True]
print("unresolved mentions need alias additions, not fuzzy guesses")
```
Expected: A coverage check. Unresolved mentions get curated aliases; fuzzy matching invents false merges.

## Other ways people phrase this
### company entity resolution failed agent
Canonical table on CIK with per-source aliases. Resolve before comparing.

### same company different names data sources
Names vary; identifiers do not. Key on CIK, alias the rest.

### ticker vs company name mismatch
Tickers change too. The CIK is the only stable key.

## Why it happens
Every data source identifies companies its own way, and names change through renames, M&A, and ticker changes. Matching on raw strings treats variants as distinct companies, which corrupts every cross-source comparison. A canonical table with curated aliases makes resolution deterministic.

## Edge cases
- M&A creates successor entities; model them as new CIKs with links, not aliases.
- Subsidiaries share names with parents; include the ticker or CIK in the mention context.
- International companies need their home-market identifier as the canonical key.
- Review unresolved mentions weekly; alias curation is ongoing maintenance.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_6vae95bQCb1jAKp32fATJg
