# Common Crawl is a sample not a census, and it obeys robots.txt

## TL;DR

Treat Common Crawl as a sample: absence of a URL is not evidence the page never existed, and robots.txt or the size cap can explain gaps. For content search, use the WET text files or the columnar index, not the URL index. Context: From a community OSINT guide on Common Crawl.

Use this when you run into the situation in the title.

**When not to use this skill:** unrelated tasks. It covers only the procedure above.

## Compatibility

The steps above apply to the commands named in them. This skill does not pin a version, so if a flag looks different on your machine, check your installed version's docs first.

## Details

The gotchas list: it is a sample, not a census, so a missing URL proves nothing about whether the page existed; the index is URL-keyed, not full-text, so content search means the WET files or the columnar index; crawlers obey robots.txt and an opt-out registry, so zero records may mean the site blocked the crawler; bodies over the fetch limit are truncated, so hashes of long pages will not match the original.
