Supported Paywall Domains and Bypass Strategies in qiaomu-anything-to-notebooklm
The fetch_url.sh script maintains five specialized domain lists and executes a hierarchical six-level bypass system to extract full-text content from over 77 paywall-protected publications, including The Wall Street Journal, Financial Times, and The New York Times.
The qiaomu-anything-to-notebooklm repository provides a robust shell-based pipeline for converting web articles into notebook-ready formats. At its core, the scripts/fetch_url.sh utility implements sophisticated supported paywall domains and bypass strategies that programmatically retrieve subscription content by rotating through search-engine user agents, AMP redirects, and archive fallbacks.
The Five Domain Classification Lists
The script organizes paywall behavior into five distinct arrays defined at the top of fetch_url.sh. These lists determine which bypass tactics are attempted for a given URL.
Googlebot-Whitelisted Domains
Sites in the GOOGLEBOT_DOMAINS array serve full content to Google's crawler for SEO indexing. This list includes wsj.com, barrons.com, ft.com, economist.com, and theaustralian.com.au (lines 24‑25). When a URL matches this set, the script immediately attempts Level 2 Googlebot spoofing.
Bingbot-Whitelisted Domains
The BINGBOT_DOMAINS array contains publishers that allow Bingbot access, such as haaretz.com, nzherald.co.nz, stratfor.com, and themarker.com (lines 27‑28). These trigger alternative Microsoft crawler impersonation.
Referer-Based Bypass Domains
The FACEBOOK_REF_DOMAINS list identifies sites that unlock content when the HTTP referer mimics social traffic, including law.com, ftm.nl, law360.com, and sloanreview.mit.edu (lines 30‑31).
AMP-Enabled Publishers
The AMP_DOMAINS array tracks publishers with lightweight AMP versions that often bypass paywalls, covering wsj.com, bostonglobe.com, latimes.com, chicagotribune.com, and seattletimes.com (lines 33‑34).
Master Paywall Registry
The comprehensive PAYWALL_DOMAINS list contains approximately 77 known subscription sites—including nytimes.com, bloomberg.com, medium.com, and businessinsider.com—that trigger the full hierarchy of generic bypass attempts (lines 35‑36).
The Six-Tier Bypass Hierarchy
The script processes requests through a cascading failure system, stopping immediately when _has_content returns true and _is_paywall_content returns false.
Level 1: Proxy Services
Before attempting direct spoofing, the script queries public extraction proxies. It calls https://r.jina.ai/<URL> for wide-coverage Markdown conversion and https://defuddle.md/<URL> for structured YAML front-matter output (lines 28‑34). These services act as pre-processing layers that often strip paywall JavaScript entirely.
Level 2: Site-Specific Search-Engine Spoofing
For URLs matching GOOGLEBOT_DOMAINS or BINGBOT_DOMAINS, the script executes targeted user-agent deception.
- Googlebot path: Sets UA to Googlebot, injects
X-Forwarded-For: 66.249.66.1, and appends a Google referer (lines 38‑66). - Bingbot path: Applies analogous Microsoft crawler headers (lines 66‑88).
The _domain_matches helper function (line 38) performs the list lookup before these branches execute.
Level 3: Generic Paywall Bypass Suite
When a domain appears in PAYWALL_DOMAINS but not in earlier specific lists, the script iterates through six sub-strategies:
- Googlebot UA with X‑Forwarded‑For – Applies the same SEO-whitelist spoofing to generic domains.
- Bingbot UA – Alternate search-engine identity.
- Facebook referer – Sets referer to
facebook.comfor domains inFACEBOOK_REF_DOMAINS. - Twitter referer – Uses
t.coas the HTTP referer. - AMP redirect attempts – Tries common patterns including
/amp,?outputType=amp, and.amp.html, falling back to regex-based URL transformation. - EU IP simulation – Injects random European Union IP ranges with a standard browser UA, exploiting regional content variations.
This suite runs from line 92 through line 124, utilizing the _is_paywall_content check after each attempt to detect soft-paywall blocks.
Level 4: Archive.today Fallback
If direct fetching fails, the script requests https://archive.today/newest/<URL> (lines 126‑151). When archive.today returns a CAPTCHA challenge, the script emits exit code 75, allowing calling programs to prompt users for manual intervention.
Level 5: Google Cache
The script attempts to retrieve the page from Google's cached copy at webcache.googleusercontent.com (lines 152‑166), bypassing live paywall checks entirely.
Level 6: Local Agent Fetch
As a final fallback, if npx is available on the system, the script executes agent-fetch (lines 167‑174), a local Node-based scraping utility.
Content Extraction and Validation
Regardless of which bypass level succeeds, the script applies two critical validation functions before returning content.
The _has_content function verifies that the response contains substantial text, while _is_paywall_content inspects for subscription banners and paywall keywords. If both checks pass, the script attempts _extract_jsonld_article to pull structured article data from JSON‑LD schemas embedded in the HTML. When JSON‑LD is unavailable, _html_to_text converts the remaining markup to clean Markdown or plain text for downstream processing.
Practical Usage Examples
The following patterns demonstrate how to invoke fetch_url.sh for various paywall scenarios:
# Basic fetch from a paywalled site
./scripts/fetch_url.sh "https://www.nytimes.com/2024/04/01/technology/ai-news.html"
# Using a corporate proxy for all requests
./scripts/fetch_url.sh "https://www.wsj.com/articles/example" "http://proxy.company.com:3128"
# Handling CAPTCHA requirements programmatically
url="https://www.economist.com/leaders/2024/01/01/article"
./scripts/fetch_url.sh "$url"
if [ $? -eq 75 ]; then
echo "Archive.today requires CAPTCHA. Complete manually at:"
echo "https://archive.today/newest/$url"
fi
# Capturing article text into a variable for LLM processing
article_text=$(./scripts/fetch_url.sh "https://www.ft.com/content/abc123")
if [ $? -eq 0 ] && [ -n "$article_text" ]; then
echo "Successfully extracted $(wc -l <<< "$article_text") lines"
fi
Summary
- Five domain lists categorize paywall behavior:
GOOGLEBOT_DOMAINS,BINGBOT_DOMAINS,FACEBOOK_REF_DOMAINS,AMP_DOMAINS, and the masterPAYWALL_DOMAINSregistry. - Six hierarchical levels attempt extraction: proxy services, site-specific bot spoofing, generic multi-vector bypass, archive.today, Google cache, and local agent-fetch.
- Validation functions (
_has_content,_is_paywall_content) ensure only readable articles are returned, not paywall blocks. - Exit code 75 signals when archive.today requires manual CAPTCHA completion.
- Source location: All logic resides in
scripts/fetch_url.sh(lines 24‑174), with domain definitions starting at line 24.
Frequently Asked Questions
How does the script determine which bypass strategy to use first?
The script checks URL patterns against the five domain lists using the _domain_matches helper function. It prioritizes Level 1 proxy services for all requests, then escalates to Level 2 Googlebot/Bingbot spoofing only if the domain appears in those specific whitelists. Generic paywall domains receive the full Level 3 multi-vector treatment before falling back to archives and caches.
What happens when a site is not in any predefined domain list?
If a URL does not match GOOGLEBOT_DOMAINS, BINGBOT_DOMAINS, FACEBOOK_REF_DOMAINS, or AMP_DOMAINS, but the user suspects a paywall, the script can still attempt extraction through the proxy services (Level 1) or, if manually classified, the generic paywall bypass suite (Level 3). The master PAYWALL_DOMAINS list covers approximately 77 major publications, but the Level 3 strategies execute for any domain flagged as paywalled.
Why does the script use exit code 75 specifically for CAPTCHA detection?
Exit code 75 (EX_TEMPFAIL in sysexits.h) signals a temporary failure requiring user intervention. When fetch_url.sh detects that archive.today has returned a CAPTCHA challenge instead of article content, it exits with code 75. This convention allows wrapper scripts and automation pipelines to distinguish between permanent fetch failures and cases where manual browser authentication can resolve the block.
Can the bypass strategies be extended to new paywall domains?
Yes. The domain lists are defined as bash arrays near the top of scripts/fetch_url.sh (lines 24‑36). Adding a new domain requires appending it to the appropriate array—such as GOOGLEBOT_DOMAINS if the site serves content to Googlebot, or PAYWALL_DOMAINS for generic handling. The hierarchical logic automatically includes new entries in the relevant bypass paths without modifying the core execution flow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →