cheerio selector returns empty on javascript-rendered page
Explains why Cheerio selectors return empty on JavaScript-rendered pages and how to fix it: prefer the page's JSON endpoint or switch to PlaywrightCrawler. Use when fetched HTML lacks content the browser shows. Not for wrong selectors or blocked requests.
TL;DR
Cheerio only sees the HTML the server sent. It never runs JavaScript, so anything rendered client-side is invisible to it. Switch the fetch to a real browser (PlaywrightCrawler) or find the JSON endpoint the page's scripts call and hit that directly.
cheerio selector returns empty on javascript-rendered pageUse this when
- Cheerio selectors return empty on a page that shows content in a browser.
- View-source shows a bare shell but devtools shows the content.
- A CheerioCrawler pipeline suddenly returns empty datasets.
Not for this skill when
- The selector is simply wrong. Test it against the fetched HTML first.
- The page blocks the request. That is an access problem.
Steps
- Compare: fetch the page with your crawler and diff it against browser view-source. Verify: the fetched HTML lacks the content the browser shows.
- Open devtools Network tab, reload, and find the XHR or fetch call that returns the data as JSON. Verify: the response contains the fields you want.
- Prefer the JSON endpoint: request it directly with your HTTP client and parse the JSON. Verify: you get structured data with no HTML parsing at all.
- If no clean endpoint exists, switch to PlaywrightCrawler and wait for the selector before extracting. Verify: the selector now returns the content.
- Re-run the pipeline. Verify: the dataset is populated.
Variant phrasings
cheerio empty result dynamic page
Same cause, symptom-first phrasing.
cheerio javascript rendering
The capability question behind the failure.
cheerio vs playwright scraping
The tool-choice search that lands here.
Compatibility: Cheerio 1.x with Crawlee 3.x CheerioCrawler, or standalone. The limitation is architectural, not version-specific.
Why it happens
Cheerio is an HTML parser, not a browser. It parses the exact bytes the HTTP response carried. Modern pages ship an empty shell plus scripts that fill it in; without a JavaScript engine those scripts never run, so the DOM Cheerio sees is the empty shell.
Edge cases / pitfalls
- Some pages render server-side for bots and client-side for browsers. A plain fetch may actually work; compare before switching tools.
- JSON endpoints sometimes need the same cookies or headers as the page. Copy the request headers from devtools.
- Hitting a private JSON endpoint the site did not publish is still their data. Check the terms before building on it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_s8nacCLUVgmL4AV8hAuHkw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.