# Competitive intel scraper IP banned after a bulk crawl

## TL;DR
An IP ban after a bulk crawl means the site decided your crawl pattern was abusive, usually a burst of requests with no pauses, and the ban is the consequence, not a puzzle to route around. Do not rotate IPs or proxies to dodge it; that escalates a rate problem into a ban-evasion problem. Stop crawling the site, contact the owner to explain the accidental burst and ask for the ban to be lifted, and rebuild the collection around polite rates, caching, and sanctioned sources like official APIs or licensed data vendors.

## The error
```text
HTTP 403 Forbidden (all requests)
IP banned after bulk crawl: every request from [client ip range] refused
```

## When this helps
- every request to a site returns 403 after a bulk crawl
- a briefing agent's source goes dark with an IP-level block
- rebuilding a crawl pattern after an abuse incident
- deciding how to collect competitor data without triggering bans

## When it doesn't
- you want to dodge the ban with proxies or IP rotation; that is ban evasion
- the ban followed a deliberate aggressive scrape; own it in the unban request or drop the source
- you need the data faster than the polite rate allows; license it instead

## Works with
Any HTTP client. Ban behavior is entirely site-side; no client setting lifts a ban.

## Steps
### 1. Halt all crawling to the site immediately
```python
import json
reg = json.load(open("sources.json"))
reg["https://YOUR-site/"] = {"status": "ip-banned", "action": "paused pending owner contact"}
json.dump(reg, open("sources.json", "w"), indent=2)
print("crawling paused for", "https://YOUR-site/")
```
Expected: The site marked paused in the registry. Every further request while banned hardens the block and hurts the unban request.

### 2. Write to the site owner explaining the burst and asking for a lift
```bash
printf 'Subject: request to lift IP block after accidental bulk crawl\n\nWe run a market intel briefing agent that crawled your site too aggressively\non [date] and triggered your abuse protection. The burst was unintentional.\nWe have paused all crawling and will keep future fetches under a polite rate\nwith a descriptive user agent. Could you lift the block on our range, or point\nus to an official API or data feed we should use instead?\n' | tee unban_request.txt
cat unban_request.txt
```
Expected: A sent, honest unban request. Owners lift bans for accidental crawlers that commit to polite behavior; they do not lift bans for evaders.

### 3. Rebuild the crawl with a polite rate and caching
```python
import time, requests
s = requests.Session()
s.headers.update({"User-Agent": "IntelBriefingBot/1.0"})
cache = {}
def polite_get(url):
    if url in cache:
        return cache[url]
    time.sleep(10)
    r = s.get(url, timeout=30)
    cache[url] = r.status_code
    return r.status_code
print(polite_get("https://YOUR-site/"))
```
Expected: A crawl pattern with real gaps and a cache so no URL is fetched twice. This is the pattern to promise the site owner.

### 4. Source the backlog from archives and licensed vendors meanwhile
```bash
curl -s "http://archive.org/wayback/available?url=[site]/pricing" | head -c 300; echo
```
Expected: Archived copies covering the banned window. The Wayback Machine and licensed vendors fill the gap without touching the live site.

## Other ways people phrase this
### ip banned after bulk crawl scraper
The classic burst-then-ban sequence. The fix is social (contact the owner) plus technical (polite rebuild), in that order.

### competitor site blocking all requests 403
Blanket refusal is usually IP or range based. Check from a different network once to confirm it is IP-level, then stop probing.

### scraper ip block how to get unbanned
Honest contact plus a credible polite-crawl commitment. There is no header or trick that substitutes for it.

## Why it happens
Bulk crawls with no delays look identical to denial-of-service traffic, so abuse protection bans the source IP. The ban is working as intended: it protects the site's capacity. Rotating IPs to continue the same crawl converts an accident into deliberate evasion, which is why the compliant path runs through the site owner.

## Edge cases
- Shared or datacenter IPs mean someone else's crawl may have caused your ban; mention your user agent and timing in the unban request so the owner can distinguish you.
- Some bans are temporary and lift in 24 to 72 hours; still contact the owner rather than waiting silently with the old crawl pattern.
- Rate limits in robots.txt Crawl-delay are hints some owners set; honoring them prevents most bans.
- If the owner declines to lift the ban, that is final; re-source the data and do not re-approach from new IPs.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_tFB-9-7aziqpACkbaMtxjQ
