# robots.txt is blocking the pricing page crawl

## TL;DR
When robots.txt disallows the pricing path, that is the site owner declining crawler access, and the correct response is to respect it rather than route around it. The debug part is quick: fetch robots.txt, find which rule matches your path, and check whether your crawler is accidentally hitting disallowed URLs it does not need. For the pricing data itself, use the site's official API, a licensed data vendor, the sitemap's allowed pages, or ask the owner for a sanctioned feed.

## The error
```text
(no HTTP error; the crawler refuses to fetch)
robots.txt disallows: /pricing/*
scraper skipped 42 URLs as disallowed
```

## When this helps
- a crawler suddenly skips URLs it used to fetch
- auditing a new source before a briefing agent depends on it
- pricing pages vanish from a crawl while blog pages still work
- deciding whether a competitor pricing source is crawlable at all

## When it doesn't
- you want to ignore robots.txt and fetch anyway; this skill will not help with that
- the path is allowed but returns 403 anyway, that is a bot-protection problem, not robots.txt
- you need data behind a login wall

## Works with
Any crawler: python urllib.robotparser, scrapy, or curl for manual checks. robots.txt semantics are stable across versions.

## Steps
### 1. Read the file your crawler is obeying
```bash
curl -s "https://YOUR-site/robots.txt"
```
Expected: The raw rules. Look for the Disallow lines under the User-agent block that matches your crawler, or under the wildcard block.

### 2. Find which rule matches your pricing URLs
```bash
curl -s "https://YOUR-site/robots.txt" | grep -i -A2 -B2 "pricing"
```
Expected: The specific Disallow line covering your path. If a narrower rule allows part of the section and a broader one blocks it, the most specific match wins per the standard.

### 3. Check your crawler is not wandering into disallowed paths by accident
```python
from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.set_url("https://YOUR-site/robots.txt")
rp.read()
for u in ["https://YOUR-site/pricing", "https://YOUR-site/pricing/enterprise", "https://YOUR-site/blog/pricing-guide"]:
    print(u, "allowed:" , rp.can_fetch("IntelBriefingBot", u))
```
Expected: A per-URL allowed verdict. Blog or docs pages about pricing are often allowed even when the pricing app itself is blocked; fetch only what is allowed.

### 4. Source the pricing data from an allowed channel
```bash
curl -s "https://YOUR-site/sitemap.xml" | head -20
```
Expected: The sitemap's crawlable URLs. Pricing pages that matter for intel often have public summaries, press releases, or a vendor API; if the data truly only exists behind a disallowed path, ask the site owner for a feed.

## Other ways people phrase this
### crawler blocked by robots.txt disallow rule
General form of the same situation. The file is the site owner's stated preference; treat a Disallow as a no.

### pricing page not in crawl, robots.txt check
Pricing sections are the most commonly disallowed paths on SaaS sites. Check for allowed alternates like public pricing summaries or press releases first.

### scraper respects robots.txt but misses data
That is the tool working correctly. The gap is filled with sanctioned sources, not with ignoring the file.

## Why it happens
robots.txt is the site owner's machine-readable statement of where crawlers are welcome. Pricing paths get disallowed because they are high-value, frequently changing, and expensive to serve to bots. A crawler that obeys the file is doing its job; the data gap is a sourcing problem, not a crawler bug.

## Edge cases
- robots.txt is advisory, not access control; obeying it is a norms and ToS matter, and ignoring it is how crawlers get IP-blocked.
- Some sites block all bots but allow specific partner user agents; do not spoof another bot's user agent to qualify.
- A missing robots.txt means everything is technically fetchable, but terms of service and rate etiquette still apply.
- Sitemap URLs can include paths robots.txt disallows; the Disallow still governs crawling even when the sitemap lists the page.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_SazLC5udAbtYND83hj1H1Q
