To pull a capture from Common Crawl: query the index for filename, offset, and length, then issue an HTTP range request for those exact bytes. The range ends at offset plus length minus one; get the arithmetic right or you will download a truncated record.

Context: From the official Common Crawl CDX index guide. The index does not hand you the record; it hands you the address: filename, offset, and length of the capture inside the crawl's WARC files. You then fetch it with an HTTP range request against the data bucket, where the byte range ends at offset plus length minus one. People query the index and stop, expecting the content in the response.