cheeriocrawler tutorial
Walks an agent through building a fast web scraper with Crawlee's CheerioCrawler: install, requestHandler with Cheerio selectors, link enqueueing, retries, and dataset export. Use it when scraping server-rendered HTML pages where a full browser is overkill. Not for JavaScript-rendered pages that need PlaywrightCrawler or PuppeteerCrawler.
TL;DR
CheerioCrawler is Crawlee's fast HTTP crawler: it fetches pages with plain requests and parses them with Cheerio, the jQuery-style selector library, so no browser ever launches. Install the crawlee package, write a requestHandler that extracts data with dollar-sign selectors, call enqueueLinks to follow more pages, and save rows with pushData. Use it when the target pages are server-rendered HTML; pick PlaywrightCrawler instead when content only appears after JavaScript runs.
cheeriocrawler tutorialSteps
- Install Crawlee and set your project to module mode. Run
npm install crawleein a folder whose package.json has"type": "module".
Expected: the install finishes with no errors and node --version reports v18 or newer.
- Write the smallest possible crawler. Import CheerioCrawler, create it with a requestHandler that receives the page URL and the Cheerio dollar function, and pull one field, like the page title.
Expected: running the script prints the titles of the pages you seeded.
- Save results with Dataset.pushData instead of console logging. Each call appends a JSON row to local storage.
Expected: files appear under storage/datasets/default and every row has the fields you pushed.
- Follow links with enqueueLinks. Point it at a selector for pagination or listing links, and it adds them to the crawl queue with automatic dedup.
Expected: the crawl visits more pages than you seeded and no URL is visited twice.
- Bound the crawl before running it for real. Set maxRequestsPerCrawl as a safety cap and maxRequestRetries for flaky pages, and give slow sites a requestHandlerTimeoutSecs budget.
Expected: the crawler stops on its own at the cap and retries failed pages instead of dying on the first 500.
import { CheerioCrawler, Dataset } from "crawlee";
const crawler = new CheerioCrawler({
maxRequestsPerCrawl: 50,
maxRequestRetries: 3,
requestHandlerTimeoutSecs: 30,
async requestHandler({ request, $, enqueueLinks }) {
const title = $("h1").first().text().trim();
await Dataset.pushData({ url: request.url, title });
await enqueueLinks({ selector: "a.next-page", label: "LIST" });
},
async failedRequestHandler({ request }) {
console.log("Gave up on " + request.url);
},
});
await crawler.run(["https://example.com/products"]);
const dataset = await Dataset.open();
await dataset.exportToCSV("products");Use this when
- You need to scrape server-rendered HTML pages fast without a browser
- A growth agent is gathering listings, prices, or articles from a site
- You want automatic retries, dedup, and dataset storage without building them
- The crawl needs pagination handling through enqueueLinks
- You are choosing which Crawlee crawler class fits the job
Not for this skill when
- The page content renders only after JavaScript executes (use PlaywrightCrawler or PuppeteerCrawler)
- You need to click buttons, fill forms, or scroll (that needs a real browser)
- You are scraping a site whose terms forbid automated access (check first)
- You only need one page fetched once (a plain HTTP request plus Cheerio is simpler)
Variant phrasings
cheerio crawler example
Same class, same pattern: new CheerioCrawler with a requestHandler, Cheerio selectors inside, pushData for rows.
how to use CheerioCrawler in crawlee
This page is the answer: install crawlee, construct CheerioCrawler, implement requestHandler, run with seed URLs.
crawlee tutorial
CheerioCrawler is the recommended starting crawler in the Crawlee docs; this skill covers it. Browser-based crawling (PlaywrightCrawler) is the next step up.
Why it happens
CheerioCrawler exists because most scraping targets are plain HTML and launching a browser for them wastes seconds per page. It pairs a request queue (where to go, with dedup and retries) with Cheerio (what to extract, with jQuery-style selectors). The usual failure is a mismatch: an agent points CheerioCrawler at a JavaScript-rendered page, gets empty selectors back, and assumes the selectors are wrong when the crawler class is wrong.
Edge cases
- Single-page apps return empty results: the HTML has no content until JS runs. Switch to PlaywrightCrawler and wait for the selector.
- Getting blocked after a few dozen requests: add a ProxyConfiguration and lower maxConcurrency; CheerioCrawler is fast enough to trip rate limits.
- Non-UTF8 pages come back garbled: the crawler handles most encodings, but check the raw body in the KeyValueStore when text looks wrong.
- enqueueLinks with no selector enqueues every link on the page: fine for a sitemap-style crawl, dangerous on a huge site. Always set maxRequestsPerCrawl.
- Respect the site: keep concurrency modest, honor robots rules, and stop if the site owner objects. Fast crawlers get noticed.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_abeMZK3tYWOegCWq6TI79A
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.