# How the System Handles Multiple Content Sources for Mixed Output

> Learn how qiaomu-anything-to-notebooklm merges PDFs, web pages, podcasts & social media into one NotebookLM context for AI-driven insights across all your sources.

- Repository: [向阳乔木/qiaomu-anything-to-notebooklm](https://github.com/joeseesun/qiaomu-anything-to-notebooklm)
- Tags: internals
- Published: 2026-05-16

---

**The `qiaomu-anything-to-notebooklm` repository consolidates PDFs, web pages, podcasts, and social media posts into a single NotebookLM conversation context, enabling AI-generated answers that synthesize information across all registered sources.**

This system solves the challenge of fragmented research by providing a unified ingestion pipeline. Through [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py), users feed heterogeneous inputs into NotebookLM's persistent context, allowing the model to perform cross-source reasoning and produce cohesive mixed outputs that reference every uploaded material.

## Input Type Detection and Classification

The workflow begins with intelligent input classification. In [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py), the `detect_input_type()` function (lines 16‑48) examines the supplied argument and categorizes it as `epub`, `document`, `podcast`, `x_twitter`, or `url`.

For local files, the function validates the file suffix. For remote resources, it checks whether the string starts with `http` and matches known domain patterns (lines 16‑28). This detection ensures each content type follows the appropriate extraction and registration path before entering the unified NotebookLM context.

## Source Registration Workflow

Once classified, each source type undergoes specialized preprocessing before registration. The system leverages NotebookLM's native `source add` CLI command to build a cumulative knowledge base.

### Local Files and Documents

For **PDF, EPUB, and DOCX files**, the system calls `upload_to_notebooklm()`, which executes `notebooklm source add <file> --title <title>` (lines 84‑88). EPUB files undergo preliminary text extraction before upload, while other document formats pass directly to NotebookLM as binary sources.

### Web URLs

**HTTP and HTTPS links** bypass local file handling. The system registers these directly via `notebooklm source add <url>` (lines 55‑58), allowing NotebookLM to crawl and index the webpage content without intermediate storage.

### Podcasts and Social Media Content

For **podcast episodes** and **X/Twitter posts**, the system first transforms remote content into local text artifacts. The helper script [`scripts/get_podcast_transcript.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/get_podcast_transcript.py) generates a `.txt` transcript file, while [`fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/fetch_url.sh) processes social media URLs. These extracted files then flow through the same `upload_to_notebooklm()` logic (see the `podcast` block, lines 65‑71, and the `x_twitter` block, lines 30‑36), ensuring uniform treatment of derived text content.

```bash

# Example 2 – Adding a podcast transcript and a tweet before asking

# 1️⃣ Get transcript (script returns a .txt file)

$ python scripts/get_podcast_transcript.py https://podcast.com/episode123

# 2️⃣ Register the transcript as a source

$ notebooklm source add /tmp/podcast_123.txt --title "Podcast 123"

# 3️⃣ Register a tweet URL as another source

$ notebooklm source add https://x.com/user/status/456789

# 4️⃣ Ask a mixed‑source question

$ notebooklm ask "Summarize the main arguments from the podcast and the tweet."

```

## Cross-Source Analysis and Mixed Output Generation

NotebookLM maintains a **persistent conversation context** across all added sources. This architectural feature enables the mixed-output capability, where the model references previously uploaded materials when answering questions about newly added content.

### Progressive Three-Round Questioning

The `deep_analysis()` routine (lines 66‑94) orchestrates the synthesis phase. After uploading a new source, the system invokes `generate_questions_progressive()` (lines 80‑84) to create tailored questions that reference the current content type via `label_for(content_type)` (lines 98‑110). Because NotebookLM retains earlier sources in context, the resulting answers naturally incorporate insights from all previously registered materials—regardless of whether they originated as PDFs, podcasts, or tweets.

### Context Accumulation Across Sources

Each `source add` operation expands the available knowledge graph. When processing multiple inputs sequentially, the system does not isolate analyses. Instead, the progressive questioning rounds treat the accumulated sources as a unified corpus, allowing queries to draw simultaneously on disparate input types.

```bash

# Example 1 – Add a local PDF and a web URL, then run deep analysis

$ python main.py /path/to/report.pdf --deep-analysis

# main.py detects 'document', uploads the PDF as a source, and starts the rounds.

$ python main.py https://example.com/article.html --deep-analysis

# The URL is added as a source (notebooklm source add <url>) and the same rounds run,

# now the model can answer using both the PDF and the webpage content.

```

## Result Aggregation and Storage

The system captures all interactions in a structured format. During the `deep_analysis()` execution, questions and answers populate a `result` dictionary (lines 98‑106) which serializes to `/tmp/<title>_analysis.json`. This JSON output includes the complete question set, generated answers, round counts, and metadata flags indicating answer completeness.

For teams using Feishu integration, the `format_feishu_markdown()` function converts results into webhook-compatible formats, enabling automated distribution of mixed-source analyses to collaboration channels via [`feishu-read-mcp/src/server.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/feishu-read-mcp/src/server.py).

## Summary

- **Unified ingestion**: The `detect_input_type()` function in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py) (lines 16‑48) automatically classifies inputs ranging from local documents to remote URLs and social media posts.
- **Flexible preprocessing**: Podcasts and X/Twitter content convert to text via dedicated helper scripts before registration through `upload_to_notebooklm()`.
- **Contextual accumulation**: NotebookLM's persistent conversation state allows each new source to enrich the collective knowledge base without overwriting previous uploads.
- **Progressive analysis**: The `deep_analysis()` routine (lines 66‑94) generates cross-source insights through three-round questioning cycles powered by `generate_questions_progressive()`.
- **Structured output**: Results serialize to JSON at `/tmp/<title>_analysis.json`, with optional Feishu markdown formatting for team distribution.

## Frequently Asked Questions

### How does the system distinguish between a local file and a URL?

The `detect_input_type()` function checks whether the input string starts with `http` to identify remote resources (lines 16‑28). For local paths, it examines file suffixes to classify documents, EPUBs, images, and audio files accordingly.

### Can I analyze a PDF and a podcast together in the same session?

Yes. Run `python main.py /path/to/document.pdf --deep-analysis` followed by `python main.py https://podcast.com/episode --deep-analysis`. Both sources register to the same NotebookLM conversation context, allowing the progressive questioning rounds to synthesize answers using content from both inputs.

### What happens to the analysis results after processing?

The system stores questions and answers in a JSON file at `/tmp/<title>_analysis.json` (lines 98‑106). This file contains the complete dialogue history, question labels generated by `label_for()`, and metadata about the number of rounds completed.

### Where does the content extraction happen for podcasts and X posts?

Specialized helper scripts handle preprocessing. Podcast URLs process through [`scripts/get_podcast_transcript.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/get_podcast_transcript.py), while X/Twitter content fetches via [`fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/fetch_url.sh) (referenced in lines 30‑36 and 65‑71 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)). These scripts produce local `.txt` files that the system then uploads to NotebookLM as standard document sources.