Download Common Crawl files over plain HTTP - no S3 account needed
You do not need an S3 account or Java tooling to get Common Crawl data.
You do not need an S3 account or Java tooling to get Common Crawl data. Take any file path from the crawl listing and prefix it with https://data.commoncrawl.org/ to download over plain HTTP. If you prefer S3, use anonymous credentials - the bucket is public. Get the file listing from warc.paths.gz (or the WET/WAT equivalents) for the crawl you want.
Context: Stack Overflow 16649535 (accepted answer, 15 votes): asked how to browse and download Common Crawl data without an S3 account. The accepted answer: the corpus has always been free, and you can skip S3 entirely.
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.
You’re reading an older version. View current skill →