## TL;DR

CheerioCrawler is Crawlee's fast HTTP crawler: it fetches pages with plain requests and parses them with Cheerio, the jQuery-style selector library, so no browser ever launches. Install the crawlee package, write a requestHandler that extracts data with dollar-sign selectors, call enqueueLinks to follow more pages, and save rows with pushData. Use it when the target pages are server-rendered HTML; pick PlaywrightCrawler instead when content only appears after JavaScript runs.

```text
cheeriocrawler tutorial
```

## Steps

1. Install Crawlee and set your project to module mode. Run `npm install crawlee` in a folder whose package.json has `"type": "module"`.
   Expected: the install finishes with no errors and `node --version` reports v18 or newer.

2. Write the smallest possible crawler. Import CheerioCrawler, create it with a requestHandler that receives the page URL and the Cheerio dollar function, and pull one field, like the page title.
   Expected: running the script prints the titles of the pages you seeded.

3. Save results with Dataset.pushData instead of console logging. Each call appends a JSON row to local storage.
   Expected: files appear under `storage/datasets/default` and every row has the fields you pushed.

4. Follow links with enqueueLinks. Point it at a selector for pagination or listing links, and it adds them to the crawl queue with automatic dedup.
   Expected: the crawl visits more pages than you seeded and no URL is visited twice.

5. Bound the crawl before running it for real. Set maxRequestsPerCrawl as a safety cap and maxRequestRetries for flaky pages, and give slow sites a requestHandlerTimeoutSecs budget.
   Expected: the crawler stops on its own at the cap and retries failed pages instead of dying on the first 500.

```javascript
import { CheerioCrawler, Dataset } from "crawlee";

const crawler = new CheerioCrawler({
  maxRequestsPerCrawl: 50,
  maxRequestRetries: 3,
  requestHandlerTimeoutSecs: 30,
  async requestHandler({ request, $, enqueueLinks }) {
    const title = $("h1").first().text().trim();
    await Dataset.pushData({ url: request.url, title });

    await enqueueLinks({ selector: "a.next-page", label: "LIST" });
  },
  async failedRequestHandler({ request }) {
    console.log("Gave up on " + request.url);
  },
});

await crawler.run(["https://example.com/products"]);

const dataset = await Dataset.open();
await dataset.exportToCSV("products");
```

## Use this when

- You need to scrape server-rendered HTML pages fast without a browser
- A growth agent is gathering listings, prices, or articles from a site
- You want automatic retries, dedup, and dataset storage without building them
- The crawl needs pagination handling through enqueueLinks
- You are choosing which Crawlee crawler class fits the job

## Not for this skill when

- The page content renders only after JavaScript executes (use PlaywrightCrawler or PuppeteerCrawler)
- You need to click buttons, fill forms, or scroll (that needs a real browser)
- You are scraping a site whose terms forbid automated access (check first)
- You only need one page fetched once (a plain HTTP request plus Cheerio is simpler)

## Variant phrasings

### cheerio crawler example
Same class, same pattern: new CheerioCrawler with a requestHandler, Cheerio selectors inside, pushData for rows.

### how to use CheerioCrawler in crawlee
This page is the answer: install crawlee, construct CheerioCrawler, implement requestHandler, run with seed URLs.

### crawlee tutorial
CheerioCrawler is the recommended starting crawler in the Crawlee docs; this skill covers it. Browser-based crawling (PlaywrightCrawler) is the next step up.

## Why it happens

CheerioCrawler exists because most scraping targets are plain HTML and launching a browser for them wastes seconds per page. It pairs a request queue (where to go, with dedup and retries) with Cheerio (what to extract, with jQuery-style selectors). The usual failure is a mismatch: an agent points CheerioCrawler at a JavaScript-rendered page, gets empty selectors back, and assumes the selectors are wrong when the crawler class is wrong.

## Edge cases

- Single-page apps return empty results: the HTML has no content until JS runs. Switch to PlaywrightCrawler and wait for the selector.
- Getting blocked after a few dozen requests: add a ProxyConfiguration and lower maxConcurrency; CheerioCrawler is fast enough to trip rate limits.
- Non-UTF8 pages come back garbled: the crawler handles most encodings, but check the raw body in the KeyValueStore when text looks wrong.
- enqueueLinks with no selector enqueues every link on the page: fine for a sitemap-style crawl, dangerous on a huge site. Always set maxRequestsPerCrawl.
- Respect the site: keep concurrency modest, honor robots rules, and stop if the site owner objects. Fast crawlers get noticed.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_abeMZK3tYWOegCWq6TI79A
