# News-scrape agent timed out on a JavaScript pricing page render

## TL;DR
A JavaScript pricing page that times out the renderer is usually bot-defended or just heavy, and the agent's mistake is treating the render as the only path to the data. Cap the render wait, detect the failure fast, and fall back to the site's API, feed, or sitemap for the pricing facts. Pricing intel does not require winning a render race against a defended page.

## The error
```text
(agent timeout)
news-scrape agent timed out waiting for JavaScript pricing page render; 60s exceeded, no data
```

## When this helps
- a scrape agent times out on JS-heavy pricing pages
- browser rendering eats the agent's time budget
- deciding whether a pricing page is worth rendering
- building fallbacks for render failures

## When it doesn't
- the page is bot-defended; do not try to defeat the defense, re-source instead
- the terms forbid automated pricing collection; skip the source
- you need real-time prices; cached fallbacks cannot give them

## Works with
Any browser automation (Playwright, Puppeteer) plus curl for the API check. Render defenses are site-side.

## Steps
### 1. Cap the render wait and detect the failure fast
```python
import asyncio
async def render_with_cap(url):
    try:
        await asyncio.wait_for(render_page(url), timeout=25)
        return "rendered"
    except asyncio.TimeoutError:
        return "render_timeout"
print("render cap: 25s, then fail fast")
```
Expected: A bounded render. Timing out fast keeps the agent's time budget for sources that work.

### 2. Check for a machine-readable pricing source first
```bash
curl -s -A "IntelBriefingBot/1.0" "https://YOUR-site/sitemap.xml" -o sitemap.xml -w "HTTP %{http_code}\n"
curl -s -A "IntelBriefingBot/1.0" "https://YOUR-site/api/pricing" -o pricing.json -w "api HTTP %{http_code}\n"
```
Expected: A sitemap or pricing API. Many JS-heavy pricing pages are backed by a JSON API the page itself calls.

### 3. Mark defended pages so the agent stops rendering them
```python
import json
reg = json.load(open("sources.json"))
reg["https://YOUR-site/pricing"] = {"status": "render-defended", "fallback": "api or feed"}
json.dump(reg, open("sources.json", "w"), indent=2)
print("page marked; future runs skip the renderer")
```
Expected: A registry entry. The agent learns which pages are render-hostile instead of rediscovering it.

### 4. Fall back to cached pricing with a staleness note
```python
import json, os
cache = "cache/pricing_[site].json"
if os.path.exists(cache):
    print("using cached pricing; briefing notes the as-of date")
else:
    print("no cache; briefing marks pricing as unavailable this cycle")
```
Expected: A graceful fallback. Stale pricing with a date beats a crashed agent.

## Other ways people phrase this
### javascript pricing page render timeout
Cap the wait, check for the backing API, fall back to cache.

### headless browser timeout pricing page
Heavy pages plus bot defenses equal slow renders. The API behind the page is the shortcut.

### scraper timeout js render news agent
Render timeouts should degrade the run, never kill it.

## Why it happens
JavaScript pricing pages are slow for automation: heavy frameworks, bot defenses that stall non-human clients, and render races the agent loses. Agents that wait indefinitely burn their whole budget on one page. The pricing data usually exists in a cleaner form behind the page, and the agent should prefer it.

## Edge cases
- Some pricing APIs require the same session cookies as the page; check the page's network calls.
- Render timeouts vary by time of day; a page that renders at night may not at noon.
- Do not raise the render cap past a minute; the budget is better spent elsewhere.
- If the backing API is undocumented, treat it as unstable and keep the cache.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_x0DqwicvzY886qVljKjA3Q
