# How qiaomu‑anything‑to‑NotebookLM Automatically Detects and Handles Different Content Types

> Learn how qiaomu-anything automatically detects and handles diverse content types. Discover its URL, file, and search term processing for seamless data integration into NotebookLM.

- Repository: [向阳乔木/qiaomu-anything-to-notebooklm](https://github.com/joeseesun/qiaomu-anything-to-notebooklm)
- Tags: how-to-guide
- Published: 2026-05-16

---

**The** `qiaomu‑anything` **system uses the** `detect_input_type` **function in** [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py) **to classify inputs as URLs, local files, or search terms, then routes each type through specialized extraction pipelines including EPUB text extraction, podcast transcription, and direct document upload.**

The `joeseesun/qiaomu-anything-to-notebooklm` repository provides an automated pipeline for ingesting diverse content formats into NotebookLM without manual configuration. By combining string pattern matching with filesystem inspection, the tool eliminates format-specific setup and automatically determines how to process each input. Understanding how the system detects and handles different content types automatically reveals the architecture behind its seamless file-to-notebook workflow.

## Input Detection Logic in main.py

At the core of the classification system sits the `detect_input_type` function defined in **[`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)** (lines 16‑48). This utility examines the input argument through a cascading logic structure that first tests for URL patterns, then falls back to filesystem analysis.

The function operates through two primary detection pathways:

### URL Pattern Recognition

When the input string starts with `http`, the system executes a series of substring checks to categorize web resources into platform-specific types:

- **WeChat articles** — Identified by `mp.weixin.qq.com` domains, tagged as `weixin`
- **YouTube content** — Detected via `youtube.com` or `youtu.be`, marked as `youtube`
- **Podcast platforms** — URLs containing `xiaoyuzhoufm.com`, `ximalaya.com`, or `bilibili.com` receive the `podcast` classification
- **X/Twitter links** — Matched against `x.com` or `twitter.com` domains as `x_twitter`
- **Generic URLs** — Any other HTTP address defaults to the `url` type

### Filesystem and Extension Mapping

If the argument lacks an HTTP prefix, the system converts it to a `Path` object and verifies existence. Non-existent paths are treated as **search** keywords. For existing files, the extension determines the content type:

- **EPUB ebooks** — `.epub` files map directly to the `epub` type
- **Documents** — `.pdf`, `.txt`, and `.md` files classify as `document`
- **Office formats** — Microsoft Word, PowerPoint, and Excel files (`.docx`, `.pptx`, `.xlsx`) receive the `office` designation
- **Images** — Visual media including `.jpg`, `.jpeg`, `.png`, `.gif`, and `.webp` fall under `image`
- **Audio files** — `.mp3` and `.wav` extensions trigger the `audio` type
- **Archives** — `.zip` files are identified separately for potential batch processing
- **Unknown types** — Any unrecognized extension returns `unknown`

## Routing Inputs to Processing Pipelines

Following detection, the **`main()`** routine (lines approximately 24‑50 in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)) implements a branching structure that dispatches each content type to its appropriate handler.

### Document and Ebook Processing

For **EPUB** files, the system invokes **`extract_epub_to_txt`** to strip formatting and extract plain text before upload. **Document** types (PDF, TXT, Markdown) bypass conversion and upload directly to NotebookLM.

### Multimedia and Web Content Extraction

**Podcast** and **YouTube** inputs trigger external transcription workflows. The system calls **[`scripts/get_podcast_transcript.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/get_podcast_transcript.py)** to retrieve audio streams and convert speech to text. For **X/Twitter** links, the pipeline executes **[`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh)** to capture tweet content. Generic **URL** types receive direct integration as NotebookLM sources without intermediate processing.

## Deep Analysis Mode and Content-Specific Workflows

When users supply the **`--deep-analysis`** flag, the detected `content_type` parameter drives customized analytical pipelines. The **`deep_analysis`** function generates progressive prompt sets tailored to specific formats—distinct question frameworks for `epub` narratives versus `youtube` video transcripts, for example—ensuring relevant AI-assisted summarization regardless of input medium.

## Practical Implementation Examples

The following patterns demonstrate practical usage of the detection system:

```python
from pathlib import Path
from main import detect_input_type

# YouTube link detection

url = "https://www.youtube.com/watch?v=abc123"
print(detect_input_type(url))  # Output: 'youtube'

# Local document classification

pdf_path = "/home/user/report.pdf"
print(detect_input_type(pdf_path))  # Output: 'document'

# Search keyword fallback for non-existent paths

print(detect_input_type("machine learning papers"))  # Output: 'search'

```

For conditional processing based on supported types:

```python
import sys
from main import detect_input_type, deep_analysis

input_arg = sys.argv[1]
content_type = detect_input_type(input_arg)

supported = {"epub", "document", "podcast", "x_twitter", "youtube", "url"}

if content_type in supported:
    deep_analysis(
        file_path=input_arg,
        title="Analysis",
        content_type=content_type,
        to_feishu=False
    )
else:
    print(f"Unsupported content type: {content_type}")

```

## Summary

- The **`detect_input_type`** function in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py) (lines 16‑48) serves as the central classification engine, distinguishing between URLs and filesystem paths through pattern matching and extension analysis.
- **URL detection** identifies specific platforms including WeChat, YouTube, podcasts, and X/Twitter through domain substring matching, defaulting to generic web URLs when no specific pattern matches.
- **File detection** maps extensions to content categories: EPUBs extract to text, documents upload directly, images/audio receive media handling, and unknown types trigger fallback logic.
- The **`main()`** function routes classified inputs to specialized pipelines: transcription scripts for multimedia, extraction utilities for ebooks, and direct API integration for web sources.
- **Deep analysis mode** leverages the detected content type to apply format-specific AI prompting strategies, optimizing NotebookLM output for each input medium.

## Frequently Asked Questions

### How does the system differentiate between a podcast URL and a regular YouTube video?

The `detect_input_type` function checks for specific domain substrings. URLs containing `xiaoyuzhoufm.com`, `ximalaya.com`, or `bilibili.com` receive the `podcast` classification, while `youtube.com` or `youtu.be` domains map to `youtube`. According to the [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py) source code, both types ultimately trigger transcription workflows, though the distinct typing allows for platform-specific handling.

### What happens when the system encounters an unsupported file format?

Files with extensions not explicitly mapped (anything other than `.epub`, `.pdf`, `.txt`, `.md`, `.docx`, `.pptx`, `.xlsx`, image formats, audio formats, or `.zip`) return the `unknown` type from `detect_input_type`. The routing logic typically skips processing for unknown types or treats them as search keywords if the path does not exist on the filesystem.

### Can the detection system handle relative file paths or only absolute paths?

The function utilizes `Path(arg)` and `path.exists()` checks, which resolve both relative and absolute paths according to Python's `pathlib` semantics. As implemented in `joeseesun/qiaomu-anything-to-notebooklm`, the extension-based detection triggers for any existing file path regardless of whether it is relative to the working directory or fully qualified.

### Where does the transcription logic for audio and video content reside?

Multimedia transcription implementations live in **[`scripts/get_podcast_transcript.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/get_podcast_transcript.py)**, which the main routine calls when processing `podcast` or `youtube` content types. For X/Twitter content retrieval, the system executes **[`scripts/fetch_url.sh`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/scripts/fetch_url.sh)** as a shell subprocess to fetch the tweet text before ingestion into NotebookLM.