# Paywalled article scraper hits the subscribe wall

## TL;DR
A subscribe wall means the publisher sells that content and scraping around it is not a technical problem but a licensing one. Do not try to bypass the paywall; briefings can cite the article's metadata, use the publisher's licensed API or syndication feed, or quote from licensed aggregators that carry the story. If the story matters, buy the subscription or the API access; that is the legitimate price of the content.

## The error
```text
(no HTTP error; fetch succeeds)
Article body replaced with subscribe-wall markup; scraper returns the paywall text instead of the article
```

## When this helps
- a scraper returns subscribe-wall text instead of articles
- briefing coverage needs stories from paywalled outlets
- deciding whether to license a publisher's content
- hardening a news pipeline against paywall contamination

## When it doesn't
- you want to bypass the paywall technically; that is circumvention of a paid product
- the outlet offers no API or license; cite from metadata and move on
- you need full text for redistribution; only a license grants that

## Works with
Any HTML parser for the guard; publisher APIs vary by outlet. Paywall behavior is site-side.

## Steps
### 1. Detect the paywall in the fetch layer and stop parsing it as an article
```python
def is_paywalled(html):
    t = html.lower()
    markers = ["subscribe to continue", "already a subscriber", "paywall", "sign in to read"]
    return any(m in t for m in markers)
print("paywall guard ready; flag the URL instead of storing wall text")
```
Expected: A guard that flags paywalled fetches. Storing wall text as article content corrupts the briefing silently.

### 2. Capture the metadata, which is usually public
```python
# headline, byline, publish date, and dek are typically in open meta tags
print("extract: headline, author, publish date, canonical URL")
print("cite the piece from metadata; do not quote body text you cannot read")
```
Expected: Usable citation metadata. Briefings can reference a paywalled piece honestly without its full text.

### 3. Find a licensed copy of the story
```bash
K="apiKey"
curl -s "https://newsapi.org/v2/everything?q=[headline keywords]&${K}=${NEWSAPI_KEY}" | head -c 300; echo
```
Expected: The same story from a licensed aggregator or a freely accessible outlet. Major stories are covered by many outlets; one open copy usually exists.

### 4. Buy access where the outlet is the only source
```bash
curl -s "https://YOUR-publisher/api/articles/[id]" -H "your auth header api key]" -o article.json -w "HTTP %{http_code}\n"
```
Expected: HTTP 200 with licensed article text. Publishers with APIs sell exactly this; a subscription plus API access is the durable fix for must-have outlets.

## Other ways people phrase this
### scraper hits paywall fetch not working
The fetch works; the content is just not free. Detect the wall and re-source instead of debugging the fetcher.

### subscribe wall blocks article scraper
Soft paywalls sometimes allow a few free articles; do not build a pipeline on the free-article loophole.

### paywalled news article parse failed
The parse fails because there is no article body in the HTML. The fix is licensing, not parsing.

## Why it happens
Paywalls are the publisher's business model: the HTML deliberately omits the article body for non-subscribers. Scrapers that expect article markup get wall markup instead. Bypassing it would be defeating a paid access control, so the legitimate options are metadata citation, licensed copies, and paid API access.

## Edge cases
- Metered paywalls that allow a few free views are not a data source; build on the API instead.
- AMP or text-only versions sometimes lack the wall, but using them to dodge the paywall is still circumvention.
- Briefings should mark paywalled citations clearly so readers know the full text needs a subscription.
- Some publishers license through aggregators more cheaply than direct API access; compare both.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Fp4wC6dkhm_GSv8zQqTGVg
