ScraperAPI Scrapy: LinkExtractor requests bypass the client - extract links manually and use scrapyGet
ScraperAPI's Scrapy integration only works if every request goes through client.scrapyGet - and a CrawlSpider with LinkExtractor breaks that, because the LinkExtractor's auto-followed requests go out through Scrapy's normal path without the client's proxy/auth, and get blocked.
ScraperAPI's Scrapy integration only works if every request goes through client.scrapyGet - and a CrawlSpider with LinkExtractor breaks that, because the LinkExtractor's auto-followed requests go out through Scrapy's normal path without the client's proxy/auth, and get blocked. The workaround is to skip the CrawlSpider: extract the links in your parse callback and call client.scrapyGet on each URL explicitly. More code, but every request actually routes through ScraperAPI.
Context: Stack Overflow (Scrapy LinkExtractor ScraperApi integration, accepted answer): when using the ScraperAPIClient with Scrapy, every request must go through client.scrapyGet(url=...) for the proxy/auth to apply. With a CrawlSpider plus LinkExtractor, Scrapy sends follow-up requests its usual way, bypassing the client entirely - so those requests get blocked. The fix: extract the links yourself and then call scrapyGet on each one, rather than relying on LinkExtractor's automatic following.
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.