When you need nested extracted fields, download as JSON, not CSV, or you will lose the nested data. Use type=urls first to audit what the crawl actually fetched before pulling the full payload. Use num to sanity check a subset before downloading everything.

Context: Official docs (Diffbot bulk API data docs): documents that downloading crawl results as CSV only includes top-level fields (nested data is dropped), that type=urls returns a URL report instead of content, and that num lets you pull a subset of results.