# Content Processing Pipeline in qiaomu‑anything‑to‑NotebookLM: From Input to Upload Explained

> Discover the qiaomu-anything-to-NotebookLM content processing pipeline. Learn how this automated system transforms inputs like URLs, PDFs, and podcasts into NotebookLM notebooks with optional AI analysis.

- Repository: [向阳乔木/qiaomu-anything-to-notebooklm](https://github.com/joeseesun/qiaomu-anything-to-notebooklm)
- Tags: deep-dive
- Published: 2026-05-16

---

**The qiaomu‑anything‑to‑NotebookLM CLI transforms any supported source—URLs, PDFs, EPUBs, podcasts, or social media links—into a NotebookLM notebook through a four-stage automated pipeline that optionally generates AI-driven deep analysis.**

The repository `joeseesun/qiaomu-anything-to-notebooklm` provides a single-entry command-line interface that orchestrates the entire **content processing pipeline** from raw input detection to final NotebookLM upload. This pipeline handles diverse content types, bypasses pay-walls when necessary, and produces structured text files compatible with Google's NotebookLM. The core logic resides in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py), which delegates acquisition tasks to specialized scripts and validates the environment through [`check_env.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/check_env.py) before execution.

## Stage 1: Input Detection and Classification

The pipeline begins with **input classification** to determine the appropriate downstream handler. The `detect_input_type` function (located at line 16 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)) analyzes the raw CLI argument to categorize it as a URL, EPUB, PDF, podcast link, X/Twitter post, WeChat article, or local file. This classification dictates which acquisition module the pipeline invokes and whether the content requires fetching, transcription, or direct processing.

## Stage 2: Content Acquisition and Normalization

Once classified, the input enters the acquisition phase where all handlers converge on producing a standardized local text file (`*.txt`) suitable for NotebookLM ingestion.

### Web URLs and Pay-Wall Bypass

For web-based content, the pipeline executes [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh), which implements a **six-level cascade** to retrieve article text:

1. **Proxy services** – Attempts extraction via `r.jina.ai` and [`defuddle.md`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/defuddle.md)
2. **Search-engine spoofing** – Uses Googlebot or Bingbot user agents with spoofed `X-Forwarded-For` and `Referer` headers
3. **Social referrer tricks** – Applies generic user agents with social-referrer spoofing and EU IP simulation
4. **AMP redirects** – Tries AMP versions via `…/amp` or `?amp` endpoints
5. **Google cache** – Fetches from `webcache.googleusercontent.com`
6. **Local agent-fetch** – Executes `npx agent-fetch` as final fallback

The script validates content through `_has_content()` and `_is_paywall_content()` checks, ensuring only substantive text proceeds to the next stage.

### Podcast and Video Transcription

Audio and video inputs trigger [`scripts/get_podcast_transcript.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/get_podcast_transcript.py), which interfaces with the GetNote API to generate transcript files. This script converts temporal media into searchable plain text, enabling NotebookLM to process podcast episodes and video content as structured sources.

### EPUB and Document Extraction

For EPUB files, the pipeline calls `extract_epub_to_txt` (defined at lines 50-69 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)) to extract and flatten the book structure into a single text document. PDFs and other plain documents pass through directly when they require no format conversion.

## Stage 3: NotebookLM Notebook Creation and Upload

With the normalized text file ready, the `upload_to_notebooklm` function (line 71 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)) executes the NotebookLM CLI commands. It first creates a notebook via `notebooklm create <title>`, then attaches the content as a source using `notebooklm source add <file> --title <title>`. This function returns a boolean status that determines whether the pipeline proceeds to optional analysis or terminates with an error.

## Stage 4: Deep Analysis and Export (Optional)

When the `--deep-analysis` flag is present, the pipeline enters the enrichment phase via `deep_analysis` (lines 65-96 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)). This stage executes three sequential steps:

- **Question generation** – `generate_questions_progressive()` builds three rounds of increasingly specific questions tailored to the content type
- **AI interrogation** – `ask_notebooklm()` (via `ask_round()`) iterates through question sets by spawning `notebooklm ask "<question>"` subprocess calls
- **Export formatting** – Results are compiled into JSON at `/tmp/<title>_analysis.json` or formatted as Markdown and sent to Feishu via `format_feishu_markdown` and `create_feishu_doc`

This stage operates independently of the basic upload, allowing users to generate supplementary insights without re-uploading source material.

## Complete Pipeline Examples

Execute the full pipeline with deep analysis for a local PDF:

```python
python main.py ./paper.pdf --deep-analysis

```

Fetch a pay-walled article, upload to NotebookLM, and export analysis to Feishu:

```bash
python main.py "https://www.wsj.com/articles/example" --deep-analysis --to-feishu

```

Both commands trigger the detection → acquisition → upload → analysis sequence, with [`check_env.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/check_env.py) validating required binaries (`notebooklm`, `lark-cli`, etc.) before execution.

## Summary

- The **content processing pipeline** consists of four stages: input detection, content acquisition, NotebookLM upload, and optional deep analysis
- `detect_input_type` in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py) routes inputs to specialized handlers including [`fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/fetch_url.sh) for web content and [`get_podcast_transcript.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/get_podcast_transcript.py) for audio
- [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh) implements a six-level pay-wall bypass cascade using proxies, bot spoofing, and cache services
- All acquisition methods converge on producing a local `.txt` file consumed by `upload_to_notebooklm`
- The deep analysis stage generates progressive question sets and exports results to JSON or Feishu via MCP integration

## Frequently Asked Questions

### How does the pipeline handle pay-walled content?

The pipeline delegates URL fetching to [`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh), which attempts six distinct extraction strategies ranging from content proxies (`r.jina.ai`) to search-engine cache fallbacks. Each level validates retrieved content against pay-wall indicators before proceeding to the next tier, ensuring maximal retrieval success for restricted articles.

### What input formats does qiaomu-anything-to-notebooklm support?

The repository accepts URLs (including X/Twitter and WeChat articles), EPUB e-books, PDF documents, podcast links, and local text files. The `detect_input_type` function classifies these inputs and routes them to appropriate acquisition modules, such as the GetNote API for podcasts or the MCP server for WeChat content.

### Where is the deep analysis output stored?

By default, deep analysis results are saved to `/tmp/<title>_analysis.json` when running locally. If the `--to-feishu` flag is specified, the pipeline instead formats the Q&A content as Markdown and creates a document in your Feishu workspace using `format_feishu_markdown` and `create_feishu_doc` functions.

### How does the script validate the environment before running?

Prior to pipeline execution, [`check_env.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/check_env.py) verifies the presence of required binaries including `notebooklm` and `lark-cli`. This pre-flight check ensures all external dependencies—the NotebookLM CLI, Feishu integration tools, and optional transcription services—are available before attempting content processing.