# Supported Paywall Domains and Bypass Strategies in qiaomu-anything-to-notebooklm

> Discover supported paywall domains and bypass strategies. This script extracts full text from 77+ publications like WSJ, FT, and NYT, overcoming paywall restrictions.

- Repository: [向阳乔木/qiaomu-anything-to-notebooklm](https://github.com/joeseesun/qiaomu-anything-to-notebooklm)
- Tags: how-to-guide
- Published: 2026-05-16

---

**The [`fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/fetch_url.sh) script maintains five specialized domain lists and executes a hierarchical six-level bypass system to extract full-text content from over 77 paywall-protected publications, including The Wall Street Journal, Financial Times, and The New York Times.**

The `qiaomu-anything-to-notebooklm` repository provides a robust shell-based pipeline for converting web articles into notebook-ready formats. At its core, the [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh) utility implements sophisticated **supported paywall domains and bypass strategies** that programmatically retrieve subscription content by rotating through search-engine user agents, AMP redirects, and archive fallbacks.

## The Five Domain Classification Lists

The script organizes paywall behavior into five distinct arrays defined at the top of [`fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/fetch_url.sh). These lists determine which bypass tactics are attempted for a given URL.

### Googlebot-Whitelisted Domains

Sites in the **`GOOGLEBOT_DOMAINS`** array serve full content to Google's crawler for SEO indexing. This list includes `wsj.com`, `barrons.com`, `ft.com`, `economist.com`, and `theaustralian.com.au` (lines 24‑25). When a URL matches this set, the script immediately attempts Level 2 Googlebot spoofing.

### Bingbot-Whitelisted Domains

The **`BINGBOT_DOMAINS`** array contains publishers that allow Bingbot access, such as `haaretz.com`, `nzherald.co.nz`, `stratfor.com`, and `themarker.com` (lines 27‑28). These trigger alternative Microsoft crawler impersonation.

### Referer-Based Bypass Domains

The **`FACEBOOK_REF_DOMAINS`** list identifies sites that unlock content when the HTTP referer mimics social traffic, including `law.com`, `ftm.nl`, `law360.com`, and `sloanreview.mit.edu` (lines 30‑31).

### AMP-Enabled Publishers

The **`AMP_DOMAINS`** array tracks publishers with lightweight AMP versions that often bypass paywalls, covering `wsj.com`, `bostonglobe.com`, `latimes.com`, `chicagotribune.com`, and `seattletimes.com` (lines 33‑34).

### Master Paywall Registry

The comprehensive **`PAYWALL_DOMAINS`** list contains approximately 77 known subscription sites—including `nytimes.com`, `bloomberg.com`, `medium.com`, and `businessinsider.com`—that trigger the full hierarchy of generic bypass attempts (lines 35‑36).

## The Six-Tier Bypass Hierarchy

The script processes requests through a cascading failure system, stopping immediately when `_has_content` returns true and `_is_paywall_content` returns false.

### Level 1: Proxy Services

Before attempting direct spoofing, the script queries public extraction proxies. It calls `https://r.jina.ai/<URL>` for wide-coverage Markdown conversion and `https://defuddle.md/<URL>` for structured YAML front-matter output (lines 28‑34). These services act as pre-processing layers that often strip paywall JavaScript entirely.

### Level 2: Site-Specific Search-Engine Spoofing

For URLs matching `GOOGLEBOT_DOMAINS` or `BINGBOT_DOMAINS`, the script executes targeted user-agent deception. 

- **Googlebot path**: Sets UA to Googlebot, injects `X-Forwarded-For: 66.249.66.1`, and appends a Google referer (lines 38‑66).
- **Bingbot path**: Applies analogous Microsoft crawler headers (lines 66‑88).

The `_domain_matches` helper function (line 38) performs the list lookup before these branches execute.

### Level 3: Generic Paywall Bypass Suite

When a domain appears in `PAYWALL_DOMAINS` but not in earlier specific lists, the script iterates through six sub-strategies:

1. **Googlebot UA with X‑Forwarded‑For** – Applies the same SEO-whitelist spoofing to generic domains.
2. **Bingbot UA** – Alternate search-engine identity.
3. **Facebook referer** – Sets referer to `facebook.com` for domains in `FACEBOOK_REF_DOMAINS`.
4. **Twitter referer** – Uses `t.co` as the HTTP referer.
5. **AMP redirect attempts** – Tries common patterns including `/amp`, `?outputType=amp`, and [`.amp.html`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/.amp.html), falling back to regex-based URL transformation.
6. **EU IP simulation** – Injects random European Union IP ranges with a standard browser UA, exploiting regional content variations.

This suite runs from line 92 through line 124, utilizing the `_is_paywall_content` check after each attempt to detect soft-paywall blocks.

### Level 4: Archive.today Fallback

If direct fetching fails, the script requests `https://archive.today/newest/<URL>` (lines 126‑151). When archive.today returns a CAPTCHA challenge, the script emits exit code **75**, allowing calling programs to prompt users for manual intervention.

### Level 5: Google Cache

The script attempts to retrieve the page from Google's cached copy at `webcache.googleusercontent.com` (lines 152‑166), bypassing live paywall checks entirely.

### Level 6: Local Agent Fetch

As a final fallback, if `npx` is available on the system, the script executes `agent-fetch` (lines 167‑174), a local Node-based scraping utility.

## Content Extraction and Validation

Regardless of which bypass level succeeds, the script applies two critical validation functions before returning content.

The `_has_content` function verifies that the response contains substantial text, while `_is_paywall_content` inspects for subscription banners and paywall keywords. If both checks pass, the script attempts `_extract_jsonld_article` to pull structured article data from JSON‑LD schemas embedded in the HTML. When JSON‑LD is unavailable, `_html_to_text` converts the remaining markup to clean Markdown or plain text for downstream processing.

## Practical Usage Examples

The following patterns demonstrate how to invoke [`fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/fetch_url.sh) for various paywall scenarios:

```bash

# Basic fetch from a paywalled site

./scripts/fetch_url.sh "https://www.nytimes.com/2024/04/01/technology/ai-news.html"

```

```bash

# Using a corporate proxy for all requests

./scripts/fetch_url.sh "https://www.wsj.com/articles/example" "http://proxy.company.com:3128"

```

```bash

# Handling CAPTCHA requirements programmatically

url="https://www.economist.com/leaders/2024/01/01/article"
./scripts/fetch_url.sh "$url"
if [ $? -eq 75 ]; then
  echo "Archive.today requires CAPTCHA. Complete manually at:"
  echo "https://archive.today/newest/$url"
fi

```

```bash

# Capturing article text into a variable for LLM processing

article_text=$(./scripts/fetch_url.sh "https://www.ft.com/content/abc123")
if [ $? -eq 0 ] && [ -n "$article_text" ]; then
  echo "Successfully extracted $(wc -l <<< "$article_text") lines"
fi

```

## Summary

- **Five domain lists** categorize paywall behavior: `GOOGLEBOT_DOMAINS`, `BINGBOT_DOMAINS`, `FACEBOOK_REF_DOMAINS`, `AMP_DOMAINS`, and the master `PAYWALL_DOMAINS` registry.
- **Six hierarchical levels** attempt extraction: proxy services, site-specific bot spoofing, generic multi-vector bypass, archive.today, Google cache, and local agent-fetch.
- **Validation functions** (`_has_content`, `_is_paywall_content`) ensure only readable articles are returned, not paywall blocks.
- **Exit code 75** signals when archive.today requires manual CAPTCHA completion.
- **Source location**: All logic resides in [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh) (lines 24‑174), with domain definitions starting at line 24.

## Frequently Asked Questions

### How does the script determine which bypass strategy to use first?

The script checks URL patterns against the five domain lists using the `_domain_matches` helper function. It prioritizes Level 1 proxy services for all requests, then escalates to Level 2 Googlebot/Bingbot spoofing only if the domain appears in those specific whitelists. Generic paywall domains receive the full Level 3 multi-vector treatment before falling back to archives and caches.

### What happens when a site is not in any predefined domain list?

If a URL does not match `GOOGLEBOT_DOMAINS`, `BINGBOT_DOMAINS`, `FACEBOOK_REF_DOMAINS`, or `AMP_DOMAINS`, but the user suspects a paywall, the script can still attempt extraction through the proxy services (Level 1) or, if manually classified, the generic paywall bypass suite (Level 3). The master `PAYWALL_DOMAINS` list covers approximately 77 major publications, but the Level 3 strategies execute for any domain flagged as paywalled.

### Why does the script use exit code 75 specifically for CAPTCHA detection?

Exit code 75 (`EX_TEMPFAIL` in sysexits.h) signals a temporary failure requiring user intervention. When [`fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/fetch_url.sh) detects that archive.today has returned a CAPTCHA challenge instead of article content, it exits with code 75. This convention allows wrapper scripts and automation pipelines to distinguish between permanent fetch failures and cases where manual browser authentication can resolve the block.

### Can the bypass strategies be extended to new paywall domains?

Yes. The domain lists are defined as bash arrays near the top of [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh) (lines 24‑36). Adding a new domain requires appending it to the appropriate array—such as `GOOGLEBOT_DOMAINS` if the site serves content to Googlebot, or `PAYWALL_DOMAINS` for generic handling. The hierarchical logic automatically includes new entries in the relevant bypass paths without modifying the core execution flow.