TL;DR: Hash every input file on ingest and refuse to diff a pair with identical hashes, and store downloads under content-addressed or version-qualified names so collisions cannot happen. The pipeline downloaded v3 over the file still named v3 from the last run, then faithfully reported that a file compared against itself had no changes.

```text
agent compared v3 against v3  -  filename collision in the download folder and the "no changes" report went out
```

1. Confirm the collision. Compute a SHA-256 hash of both input files the diff consumed. Expected: the hashes are identical, proving the "no changes" result came from comparing a file with itself.

2. Fix the download step. Save each downloaded version under a unique name that includes the version identifier and a timestamp or content hash (for example, contract_v3_[hash8].pdf), never the bare filename from the email or portal. Expected: two different versions can never land on the same path again.

3. Add a pre-diff guard. Before any comparison, hash both inputs; if the hashes match, abort with "inputs are identical files" instead of producing a "no changes" report. Expected: a self-comparison fails loudly at the guard instead of shipping a false clean report.

4. Backfill version metadata. For the comparison that already went out, re-download both versions under qualified names, verify the hashes differ, and re-run the diff. Expected: the real changes between the versions appear, and the earlier "no changes" report is retracted or corrected.

5. Make the guard permanent in the pipeline config, not a one-off script. Every future comparison runs the hash check before the diff. Expected: filename collisions become impossible to ship silently, even when a human drops files into the folder by hand.

## Use this when
- a version comparison reports "no changes" but you know the versions differ
- downloads land in a shared folder under their original filenames
- the pipeline takes two file paths and never verifies they are different files
- a "clean" redline report went out and nobody can reproduce how it was produced

## Not for this skill when
- the hashes differ but the diff is still empty (check the diff tool's sensitivity or normalization settings)
- the versions genuinely are identical (the guard firing correctly is the system working)
- the wrong pair was compared but the files are different (that is a file-selection problem, not a collision)

## Variant phrasings
- redline tool diffed a file against itself after a download overwrote it
- filename collision made the agent report no changes between versions
- pipeline compared v3 to v3 because both downloads had the same name
- "no changes" report was wrong, the inputs were the same file

## Why it happens
Download folders are append-by-name: a new file with an old name silently replaces the old bytes. The diff tool downstream has no idea the "two versions" are one file; it does exactly what it is told and reports no differences. The failure is silent because "no changes" is a perfectly valid diff result, so nothing downstream questions it.

## Edge cases
- Some portals serve the same filename for every version ("contract.pdf"). Content-addressed naming on ingest is the only reliable defense.
- Hash the bytes, not the metadata. Two files with different names but identical content should still trip the guard; the comparison would be meaningless either way.
- Keep the original filename in metadata alongside the qualified name, so humans can still find "the v3 file" without relying on the filesystem name being unique.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_tbMtaeu-UngxQHhtcpgh4Q
