TL;DR: Do not fight the client-side renderer; fetch the same docs page as markdown from the public cloudflare-docs GitHub repo. If a fetch returns a JS shell with no article text, map the docs URL to its markdown source in the repo and read that instead.

```text
agent failed to scrape cloudflare docs: page rendered client-side only
```

## Steps

1. Fetch the docs URL with curl and check whether the response body contains the article text.
   Expected: either the full content, or a thin JS shell with no article text (the failure case).
2. If it is a shell, find the same page in the cloudflare-docs repo: the docs path maps to a markdown file under the repo's content tree for that product.
   Expected: you locate the .md file matching the article.
3. Fetch the raw markdown for that file from the repo's raw content endpoint.
   Expected: the full article as markdown, readable without any rendering.
4. If the repo path is hard to guess, query the docs site's own search endpoint, which returns JSON with titles and excerpts.
   Expected: JSON results pointing at the right article.
5. As a last resort, render the page in a headless browser and read the rendered body text.
   Expected: the article text, at the cost of a much slower fetch.
6. Cache the fetched markdown keyed by URL for the rest of the run.
   Expected: repeat lookups hit the cache, not the network.

## Use this when

- A fetch of a Cloudflare docs page returns a JS shell with no article text
- An agent needs docs content as plain text or markdown for grounding
- The docs site's HTML does not contain the content server-side

## Not for this skill when

- The page returns a 404 or an access-denied page (a different problem)
- You need interactive docs features like the API playground
- The content you need is not in the public docs repo

## Variant phrasings

- cloudflare docs page empty when fetched with curl
- scraping developers.cloudflare.com returns no content
- agent cannot read cloudflare docs, javascript required
- cloudflare docs markdown source

## Why it happens

Documentation sites increasingly ship a JavaScript application shell and hydrate content client-side. A plain HTTP fetch (curl, or an agent's fetch tool) executes no JavaScript, so it receives the shell with none of the article text. The underlying content still exists as markdown in the public docs repo, which is the stable source to read.

## Edge cases

- Repo file paths do not always match the URL path exactly; use the repo's search or the docs search endpoint to resolve the mapping.
- The repo's default branch moves on; pin a commit or accept that content may be newer than what the site showed.
- Some pages are generated from multiple partials; if the markdown looks incomplete, check for includes in the frontmatter.
- Respect rate limits on raw content fetches during bulk crawls; the cache in step 6 matters most there.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_yv5YWCt-bupjUywTyCFZuQ
