How JSON-LD Article Body Extraction Works in fetch_url.sh

The fetch_url.sh script extracts full article text from embedded JSON-LD structured data using the _extract_jsonld_article function, enabling reliable paywall bypass before falling back to HTML-to-text conversion.

The fetch_url.sh utility in the joeseesun/qiaomu-anything-to-notebooklm repository orchestrates web page retrieval through multiple paywall-bypass strategies. JSON-LD article body extraction serves as the primary method for obtaining clean article content, leveraging standardized structured data that publishers embed for search engine optimization.

The _extract_jsonld_article Function

The core extraction logic resides in the _extract_jsonld_article function at lines 90‑95 of scripts/fetch_url.sh. This helper function receives the raw HTML response from a fetched URL and attempts to locate article content within <script type="application/ld+json"> blocks, where publishers commonly store the full text of paywalled articles for SEO purposes.

The function executes a targeted grep operation to isolate the articleBody property:

grep -o '"articleBody":"[^"]*"'

This pattern extracts only the first occurrence of the articleBody field as a quoted string, capturing everything between the property name and the closing quote.

Step-by-Step Processing Pipeline

The extraction process follows a precise sequence to transform raw JSON-LD into usable Markdown:

  1. HTML Acquisition: The script fetches the target URL using curl with various user-agent strategies (Googlebot, Bingbot, or generic paywall bypass), storing the raw response in a variable.

  2. Field Isolation: The grep command extracts the "articleBody":"content here" substring from the JSON-LD script block.

  3. Text Cleaning: Two sed operations at lines 94‑95 handle formatting:

    • Strip the leading "articleBody":" and trailing " delimiters
    • Un-escape JSON sequences: convert \n to literal newlines, \" to literal quotes, and \\ to backslashes
  4. Validation: If the resulting text exceeds 200 bytes, the script treats the extraction as successful (lines 48‑58) and outputs the content immediately, bypassing slower HTML parsing methods.

Integration with Paywall Bypass Strategies

JSON-LD extraction operates at multiple points in the fetch cascade. The script attempts this method when fetching via:

  • Googlebot user-agent simulation
  • Bingbot user-agent simulation
  • Generic paywall bypass techniques
  • AMP page handling

This multi-point strategy maximizes the probability of obtaining structured data before resorting to fallback HTML-to-text conversion, which is more prone to navigation elements and advertisement noise.

Implementation Reference

The extraction chain relies on standard Unix utilities available in the base environment. The complete processing logic appears as follows in the source:


# Inside scripts/fetch_url.sh (lines 90-95)

_extract_jsonld_article() {
  echo "$1" | grep -o '"articleBody":"[^"]*"' | \
    sed 's/"articleBody":"//;s/"$//' | \
    sed 's/\\n/\n/g;s/\\"/"/g;s/\\\\/\\/g'
}

When successful, the script formats the output with the page title (extracted from the <title> tag) and source URL:


# Page Title Here

Source: https://example.com/article-url

Full article text extracted from JSON-LD appears here...

Practical Usage Examples

Execute the script against any URL containing JSON-LD article markup:


# Direct fetch with automatic JSON-LD extraction

./scripts/fetch_url.sh "https://example.com/news/2024/important-report"

To route traffic through a proxy while maintaining extraction capabilities:

PROXY="http://127.0.0.1:3128"
./scripts/fetch_url.sh "https://another-site.com/story" "$PROXY"

Both commands return Markdown-formatted output when the article body is successfully extracted from JSON-LD data, or fall back to alternative parsing methods if the structured data is insufficient.

Summary

  • scripts/fetch_url.sh implements JSON-LD extraction via the _extract_jsonld_article function at lines 90‑95
  • The grep pattern "articleBody":"[^"]*" isolates the content field from structured data
  • sed commands strip JSON delimiters and un-escape sequences (\n, \", \\) to produce clean text
  • A 200-byte minimum threshold validates extraction success before outputting Markdown
  • This method takes priority over HTML parsing in the paywall-bypass cascade for multiple user-agent strategies

Frequently Asked Questions

What is JSON-LD and why does fetch_url.sh use it for article extraction?

JSON-LD (JavaScript Object Notation for Linked Data) is a structured data format publishers embed in <script type="application/ld+json"> tags to help search engines understand page content. The fetch_url.sh script targets JSON-LD because it often contains the complete article text in a clean, machine-readable format, allowing extraction without parsing complex HTML or executing JavaScript.

How does the script handle escaped characters in JSON-LD data?

The function uses two sed transformations at lines 94‑95 to normalize the extracted string: first removing the "articleBody":" prefix and trailing ", then replacing escaped sequences with literal characters—specifically converting \n to newlines, \" to quotation marks, and \\ to single backslashes.

What happens if JSON-LD extraction fails or returns insufficient content?

If the extracted text is shorter than 200 bytes or if the articleBody property is missing, the script continues its cascade to alternative methods. These include fetching via different user agents (Googlebot, Bingbot), attempting AMP versions of pages, or falling back to direct HTML-to-text conversion using readability algorithms.

Which user agents does fetch_url.sh employ when attempting JSON-LD extraction?

The script cycles through multiple user-agent strings to maximize access to JSON-LD data behind soft paywalls. According to the source implementation, it attempts extraction using Googlebot and Bingbot user agents, along with generic paywall-bypass configurations, retrying the JSON-LD parsing at each stage before proceeding to the next strategy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →