cheeriocrawler
Loads Crawlee's CheerioCrawler for fast static-HTML scraping with request queues, session rotation, and retries. Use when an agent needs to crawl paginated sites or scrape static pages at scale without a browser. Not for JavaScript-rendered pages, single one-off fetches, or sites that forbid automated access.
TL;DR
Use Crawlee's CheerioCrawler for fast static-HTML crawls with built-in request queues, session rotation, and retries. It downloads raw HTML and parses it with Cheerio, which is far lighter than driving a headless browser. It does not run JavaScript, so pages that render content client-side need a browser-based crawler instead.
cheeriocrawlerUse this when
- You need to crawl a multi-page site (listings, pagination, sitemaps) with Crawlee.
- The target pages are static HTML or server-rendered.
- You want automatic retries, proxies, and session rotation out of the box.
- Headless browsers are too slow or heavy for the job.
Not for this skill when
- The page renders content with JavaScript after load. Cheerio sees the raw HTML only.
- You only need one page fetch. A plain HTTP client is simpler.
- The site's robots.txt or terms forbid automated access. Respect that.
Steps
- Install Crawlee and Cheerio: run
npm install crawlee cheerioand confirm both land in package.json. - Create the crawler:
const { CheerioCrawler } = require('crawlee');thennew CheerioCrawler({ requestHandler: async function handler({ $, request }) { ... } }). Success check: the file imports without error under node. - Seed start URLs:
await crawler.addRequests(['https://example.com/page1'])thenawait crawler.run(). Success check: the handler logs once per request. - Extract data with Cheerio selectors:
const title = $('h1').text().trim(). Success check: values match what you see in view-source. - Enqueue follow-on links:
await enqueueLinks({ selector: 'a[href$=".html"]' }). Success check: the final dataset grows beyond the seed count. - Harden production runs: set
maxRequestRetries: 3,useSessionPool: true,sessionPoolOptions: { maxPoolSize: 20 }. Success check: failed requests re-run and block rates drop.
Variant phrasings
cheerio crawler crawlee
Same class, looser word order. Covers the import-name confusion.
crawlee cheerio scraping
Task-first phrasing for the same crawler.
fast static crawler crawlee
Covers the "lighter than Playwright" search.
Compatibility: Crawlee 3.x, Node 18+. The CheerioCrawler API above is stable since Crawlee 3.0.
Why it happens
CheerioCrawler exists because most scraping tasks do not need a browser. Downloading HTML over HTTP and parsing it with Cheerio is 10 to 50 times cheaper per page than driving Chromium, and Crawlee's request queue, session pool, and proxy rotation make that loop robust at scale.
Edge cases / pitfalls
- Some sites serve thin HTML to unknown user agents. Set a real browser UA in request options if the fetched HTML looks empty.
- Pages behind logins need
persistCookiesPerSession: trueplus a login step in a pre-navigation hook. - Very large sites need
maxRequestsPerCrawlto cap spend. - If
$('...')returns empty on a page you know renders content, the page is client-rendered. Switch to PlaywrightCrawler.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst__KJWDg5rSniptPSopxNtQg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.