How the Paywall Bypass Cascade in fetch_url.sh Works: A 6-Level Technical Breakdown

The fetch_url.sh script implements a six-level fallback cascade that attempts increasingly sophisticated techniques—from public content extraction proxies to search engine bot impersonation and archive services—to retrieve full article text without triggering paywalls.

The fetch_url.sh utility serves as the core content acquisition engine for the joeseesun/qiaomu-anything-to-notebooklm open-source project. This self-contained bash script executes a paywall bypass cascade that systematically escalates through multiple retrieval strategies, validating each response before proceeding to the next level. The architecture prioritizes speed and reliability, starting with lightweight proxy services before resorting to bot User-Agent spoofing, social referrer tricks, and archived snapshots.

Level 1: Public Content Extraction Proxies

The cascade begins with the fastest retrieval methods: public content extraction services that specialize in cleaning HTML and returning readable Markdown or text.

At line 29, the script queries https://r.jina.ai/<URL> to fetch a processed version of the article. The response passes through _has_content (lines 48‑60) to verify it contains more than eight lines and 500 characters, and _is_paywall_content (lines 77‑81) to check for subscription-block indicators. If validation succeeds, the script outputs the content and exits immediately.

If the Jina AI proxy fails or returns paywalled content, the script falls back to https://defuddle.md/<URL> at line 33, repeating the same validation sequence. These public proxies handle the heavy lifting of HTML parsing and paywall detection for standard news sites, making them the preferred first-line solution.

Level 2: Site-Specific Bot User-Agent Spoofing

When public proxies fail, the cascade escalates to impersonating legitimate search engine crawlers for specific domains known to serve full content to bots.

For domains listed in the GOOGLEBOT_DOMAINS array (line 24), the script constructs a curl request using the Googlebot User-Agent, adds X‑Forwarded‑For: 66.249.66.1 to mimic Google's IP range, and sets a Google referer (lines 39‑45). The response processing first attempts _extract_jsonld_article (line 92) to parse structured JSON‑LD article data. If successful, it formats a Markdown headline and prints the article (lines 49‑58). Otherwise, it falls back to _html_to_text (line 99) for raw HTML stripping.

The same logic applies to BINGBOT_DOMAINS (line 27) using the Bingbot User-Agent (lines 66‑73). This targeted approach exploits the fact that many publishers serve complete article markup to search engine crawlers to maintain SEO visibility while restricting human readers.

Level 3: Generic Paywall Bypass Suite

For URLs matching the PAYWALL_DOMAINS pattern (line 36), the script unleashes a comprehensive toolkit of bypass techniques in rapid succession:

  • Googlebot UA + X‑Forwarded‑For (lines 94‑101): Applies the Level 2 Googlebot strategy generically to any paywalled domain.
  • Bingbot UA (lines 118‑125): Repeats the request with Bing's crawler signature.
  • Social Referrer Spoofing: Sets Facebook referers (lines 140‑147) and Twitter referers (lines 150‑158) to exploit platforms that grant limited free views to social media traffic.
  • AMP Page Probing: Appends common AMP suffixes (/amp, ?outputType=amp, etc.) to the URL (lines 162‑184) and attempts .amp.html or /amp path rewrites (lines 186‑197) to retrieve lightweight mobile versions often exempt from paywalls.
  • EU-Style X‑Forwarded‑For: Randomizes European IP addresses (lines 200‑207) to bypass GDPR-related content restrictions that allow EU visitors sample access.

Each sub-step extracts JSON‑LD first via _extract_jsonld_article, falls back to _html_to_text for plain text conversion, and validates the result through _try_output (lines 16‑24).

Level 4: Archive.today Snapshot Retrieval

When active bypass techniques fail, the cascade switches to passive retrieval using archived snapshots. The script constructs https://archive.today/newest/<URL> at line 128 and fetches the archived version (line 130).

Before outputting, it checks for CAPTCHA pages using _is_captcha_page (line 86). If detected, the script exits with code 75 and prints a user-facing message (lines 145‑151) rather than returning garbage data. Otherwise, the HTML passes through _html_to_text (lines 136‑140) for clean text extraction.

Level 5: Google Cache Fallback

If no archive exists, the script queries Google's cached copy at https://webcache.googleusercontent.com/search?q=cache:<URL> (line 55). When the cache contains readable text passing the content validation checks, the script outputs it (lines 60‑64). This level captures content that may have been recently indexed but not yet archived by third-party services.

Level 6: Agent-Fetch Last Resort

The final level executes the local agent-fetch Node.js tool via npx --yes agent-fetch "$URL" --json (line 70), provided npx is available in the environment. This tool prints the JSON result (lines 70‑74) as a last-ditch effort to retrieve content using headless browser automation techniques.

Content Validation Functions

The cascade relies on three critical helper functions defined in scripts/fetch_url.sh to determine success:

  • _has_content (lines 48‑60): Validates that output contains more than eight lines and 500 characters, filtering out error pages and empty responses.
  • _is_paywall_content (lines 77‑81): Scans text for subscription indicators like "subscribe now" or "paywall" keywords.
  • _try_output (lines 16‑24): Orchestrates the validation sequence, printing content and exiting with status 0 immediately when both checks pass, preventing unnecessary cascade progression.

Usage Examples

Run the cascade against any news article URL:

./scripts/fetch_url.sh "https://www.nytimes.com/2024/05/01/technology/example-article.html"

Route requests through an HTTP proxy when behind corporate firewalls:

./scripts/fetch_url.sh "https://www.wsj.com/finance/example" "http://myproxy:3128"

Capture output into a variable for downstream processing:

article_md=$(./scripts/fetch_url.sh "https://www.theguardian.com/science/example")
echo "$article_md"

Summary

  • The paywall bypass cascade in fetch_url.sh implements six distinct retrieval levels, stopping at the first successful extraction.
  • Level 1 prioritizes public proxies (r.jina.ai, defuddle.md) for speed and minimal infrastructure.
  • Levels 2‑3 employ search engine bot impersonation (GOOGLEBOT_DOMAINS, BINGBOT_DOMAINS) and referrer spoofing to exploit publisher SEO exceptions.
  • Level 4 leverages archive.today snapshots while Level 5 uses Google Cache for historical retrieval.
  • Validation functions (_has_content, _is_paywall_content) ensure only clean, non-paywalled text reaches stdout.

Frequently Asked Questions

How does fetch_url.sh determine if content is actually accessible or trapped behind a paywall?

The script uses the _is_paywall_content function (lines 77‑81) to scan retrieved text for common subscription indicators like "subscribe," "paywall," or "create an account." Additionally, _has_content (lines 48‑60) ensures the response exceeds eight lines and 500 characters, filtering out soft paywalls that return truncated teaser text or login prompts.

Why does the cascade try multiple User-Agent strings instead of just using Googlebot for everything?

Different publishers maintain varying allowlists for search engine crawlers. While most sites whitelist Googlebot (defined in GOOGLEBOT_DOMAINS at line 24), others specifically check for Bingbot (BINGBOT_DOMAINS at line 27). The generic paywall bypass level (lines 94‑125) attempts both sequentially because some sites serve complete markup only to specific crawlers or block Googlebot requests lacking proper IP headers while allowing Bingbot.

What happens when archive.today returns a CAPTCHA page instead of the article?

The _is_captcha_page function (line 86) detects CAPTCHA challenges in the archive response. When triggered, the script exits with code 75 and prints a helpful message (lines 145‑151) rather than returning unreadable text, allowing calling processes to handle the failure gracefully or prompt for manual intervention.

Can the script bypass all types of paywalls, including hard paywalls requiring login credentials?

No. According to the source code in scripts/fetch_url.sh, the cascade targets "soft" paywalls that allow search engines, social media referrals, or AMP versions to access full content. Hard paywalls requiring authentication credentials cannot be bypassed by the techniques implemented in levels 1‑6, as the script does not handle authentication flows or session management.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →