Content Processing Pipeline in qiaomu‑anything‑to‑NotebookLM: From Input to Upload Explained

The qiaomu‑anything‑to‑NotebookLM CLI transforms any supported source—URLs, PDFs, EPUBs, podcasts, or social media links—into a NotebookLM notebook through a four-stage automated pipeline that optionally generates AI-driven deep analysis.

The repository joeseesun/qiaomu-anything-to-notebooklm provides a single-entry command-line interface that orchestrates the entire content processing pipeline from raw input detection to final NotebookLM upload. This pipeline handles diverse content types, bypasses pay-walls when necessary, and produces structured text files compatible with Google's NotebookLM. The core logic resides in main.py, which delegates acquisition tasks to specialized scripts and validates the environment through check_env.py before execution.

Stage 1: Input Detection and Classification

The pipeline begins with input classification to determine the appropriate downstream handler. The detect_input_type function (located at line 16 of main.py) analyzes the raw CLI argument to categorize it as a URL, EPUB, PDF, podcast link, X/Twitter post, WeChat article, or local file. This classification dictates which acquisition module the pipeline invokes and whether the content requires fetching, transcription, or direct processing.

Stage 2: Content Acquisition and Normalization

Once classified, the input enters the acquisition phase where all handlers converge on producing a standardized local text file (*.txt) suitable for NotebookLM ingestion.

Web URLs and Pay-Wall Bypass

For web-based content, the pipeline executes scripts/fetch_url.sh, which implements a six-level cascade to retrieve article text:

  1. Proxy services – Attempts extraction via r.jina.ai and defuddle.md
  2. Search-engine spoofing – Uses Googlebot or Bingbot user agents with spoofed X-Forwarded-For and Referer headers
  3. Social referrer tricks – Applies generic user agents with social-referrer spoofing and EU IP simulation
  4. AMP redirects – Tries AMP versions via …/amp or ?amp endpoints
  5. Google cache – Fetches from webcache.googleusercontent.com
  6. Local agent-fetch – Executes npx agent-fetch as final fallback

The script validates content through _has_content() and _is_paywall_content() checks, ensuring only substantive text proceeds to the next stage.

Podcast and Video Transcription

Audio and video inputs trigger scripts/get_podcast_transcript.py, which interfaces with the GetNote API to generate transcript files. This script converts temporal media into searchable plain text, enabling NotebookLM to process podcast episodes and video content as structured sources.

EPUB and Document Extraction

For EPUB files, the pipeline calls extract_epub_to_txt (defined at lines 50-69 of main.py) to extract and flatten the book structure into a single text document. PDFs and other plain documents pass through directly when they require no format conversion.

Stage 3: NotebookLM Notebook Creation and Upload

With the normalized text file ready, the upload_to_notebooklm function (line 71 of main.py) executes the NotebookLM CLI commands. It first creates a notebook via notebooklm create <title>, then attaches the content as a source using notebooklm source add <file> --title <title>. This function returns a boolean status that determines whether the pipeline proceeds to optional analysis or terminates with an error.

Stage 4: Deep Analysis and Export (Optional)

When the --deep-analysis flag is present, the pipeline enters the enrichment phase via deep_analysis (lines 65-96 of main.py). This stage executes three sequential steps:

  • Question generation – generate_questions_progressive() builds three rounds of increasingly specific questions tailored to the content type
  • AI interrogation – ask_notebooklm() (via ask_round()) iterates through question sets by spawning notebooklm ask "<question>" subprocess calls
  • Export formatting – Results are compiled into JSON at /tmp/<title>_analysis.json or formatted as Markdown and sent to Feishu via format_feishu_markdown and create_feishu_doc

This stage operates independently of the basic upload, allowing users to generate supplementary insights without re-uploading source material.

Complete Pipeline Examples

Execute the full pipeline with deep analysis for a local PDF:

python main.py ./paper.pdf --deep-analysis

Fetch a pay-walled article, upload to NotebookLM, and export analysis to Feishu:

python main.py "https://www.wsj.com/articles/example" --deep-analysis --to-feishu

Both commands trigger the detection → acquisition → upload → analysis sequence, with check_env.py validating required binaries (notebooklm, lark-cli, etc.) before execution.

Summary

  • The content processing pipeline consists of four stages: input detection, content acquisition, NotebookLM upload, and optional deep analysis
  • detect_input_type in main.py routes inputs to specialized handlers including fetch_url.sh for web content and get_podcast_transcript.py for audio
  • scripts/fetch_url.sh implements a six-level pay-wall bypass cascade using proxies, bot spoofing, and cache services
  • All acquisition methods converge on producing a local .txt file consumed by upload_to_notebooklm
  • The deep analysis stage generates progressive question sets and exports results to JSON or Feishu via MCP integration

Frequently Asked Questions

How does the pipeline handle pay-walled content?

The pipeline delegates URL fetching to scripts/fetch_url.sh, which attempts six distinct extraction strategies ranging from content proxies (r.jina.ai) to search-engine cache fallbacks. Each level validates retrieved content against pay-wall indicators before proceeding to the next tier, ensuring maximal retrieval success for restricted articles.

What input formats does qiaomu-anything-to-notebooklm support?

The repository accepts URLs (including X/Twitter and WeChat articles), EPUB e-books, PDF documents, podcast links, and local text files. The detect_input_type function classifies these inputs and routes them to appropriate acquisition modules, such as the GetNote API for podcasts or the MCP server for WeChat content.

Where is the deep analysis output stored?

By default, deep analysis results are saved to /tmp/<title>_analysis.json when running locally. If the --to-feishu flag is specified, the pipeline instead formats the Q&A content as Markdown and creates a document in your Feishu workspace using format_feishu_markdown and create_feishu_doc functions.

How does the script validate the environment before running?

Prior to pipeline execution, check_env.py verifies the presence of required binaries including notebooklm and lark-cli. This pre-flight check ensures all external dependencies—the NotebookLM CLI, Feishu integration tools, and optional transcription services—are available before attempting content processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →