How qiaomu‑anything‑to‑NotebookLM Automatically Detects and Handles Different Content Types

The qiaomu‑anything system uses the detect_input_type function in main.py to classify inputs as URLs, local files, or search terms, then routes each type through specialized extraction pipelines including EPUB text extraction, podcast transcription, and direct document upload.

The joeseesun/qiaomu-anything-to-notebooklm repository provides an automated pipeline for ingesting diverse content formats into NotebookLM without manual configuration. By combining string pattern matching with filesystem inspection, the tool eliminates format-specific setup and automatically determines how to process each input. Understanding how the system detects and handles different content types automatically reveals the architecture behind its seamless file-to-notebook workflow.

Input Detection Logic in main.py

At the core of the classification system sits the detect_input_type function defined in main.py (lines 16‑48). This utility examines the input argument through a cascading logic structure that first tests for URL patterns, then falls back to filesystem analysis.

The function operates through two primary detection pathways:

URL Pattern Recognition

When the input string starts with http, the system executes a series of substring checks to categorize web resources into platform-specific types:

  • WeChat articles — Identified by mp.weixin.qq.com domains, tagged as weixin
  • YouTube content — Detected via youtube.com or youtu.be, marked as youtube
  • Podcast platforms — URLs containing xiaoyuzhoufm.com, ximalaya.com, or bilibili.com receive the podcast classification
  • X/Twitter links — Matched against x.com or twitter.com domains as x_twitter
  • Generic URLs — Any other HTTP address defaults to the url type

Filesystem and Extension Mapping

If the argument lacks an HTTP prefix, the system converts it to a Path object and verifies existence. Non-existent paths are treated as search keywords. For existing files, the extension determines the content type:

  • EPUB ebooks — .epub files map directly to the epub type
  • Documents — .pdf, .txt, and .md files classify as document
  • Office formats — Microsoft Word, PowerPoint, and Excel files (.docx, .pptx, .xlsx) receive the office designation
  • Images — Visual media including .jpg, .jpeg, .png, .gif, and .webp fall under image
  • Audio files — .mp3 and .wav extensions trigger the audio type
  • Archives — .zip files are identified separately for potential batch processing
  • Unknown types — Any unrecognized extension returns unknown

Routing Inputs to Processing Pipelines

Following detection, the main() routine (lines approximately 24‑50 in main.py) implements a branching structure that dispatches each content type to its appropriate handler.

Document and Ebook Processing

For EPUB files, the system invokes extract_epub_to_txt to strip formatting and extract plain text before upload. Document types (PDF, TXT, Markdown) bypass conversion and upload directly to NotebookLM.

Multimedia and Web Content Extraction

Podcast and YouTube inputs trigger external transcription workflows. The system calls scripts/get_podcast_transcript.py to retrieve audio streams and convert speech to text. For X/Twitter links, the pipeline executes scripts/fetch_url.sh to capture tweet content. Generic URL types receive direct integration as NotebookLM sources without intermediate processing.

Deep Analysis Mode and Content-Specific Workflows

When users supply the --deep-analysis flag, the detected content_type parameter drives customized analytical pipelines. The deep_analysis function generates progressive prompt sets tailored to specific formats—distinct question frameworks for epub narratives versus youtube video transcripts, for example—ensuring relevant AI-assisted summarization regardless of input medium.

Practical Implementation Examples

The following patterns demonstrate practical usage of the detection system:

from pathlib import Path
from main import detect_input_type

# YouTube link detection

url = "https://www.youtube.com/watch?v=abc123"
print(detect_input_type(url))  # Output: 'youtube'

# Local document classification

pdf_path = "/home/user/report.pdf"
print(detect_input_type(pdf_path))  # Output: 'document'

# Search keyword fallback for non-existent paths

print(detect_input_type("machine learning papers"))  # Output: 'search'

For conditional processing based on supported types:

import sys
from main import detect_input_type, deep_analysis

input_arg = sys.argv[1]
content_type = detect_input_type(input_arg)

supported = {"epub", "document", "podcast", "x_twitter", "youtube", "url"}

if content_type in supported:
    deep_analysis(
        file_path=input_arg,
        title="Analysis",
        content_type=content_type,
        to_feishu=False
    )
else:
    print(f"Unsupported content type: {content_type}")

Summary

  • The detect_input_type function in main.py (lines 16‑48) serves as the central classification engine, distinguishing between URLs and filesystem paths through pattern matching and extension analysis.
  • URL detection identifies specific platforms including WeChat, YouTube, podcasts, and X/Twitter through domain substring matching, defaulting to generic web URLs when no specific pattern matches.
  • File detection maps extensions to content categories: EPUBs extract to text, documents upload directly, images/audio receive media handling, and unknown types trigger fallback logic.
  • The main() function routes classified inputs to specialized pipelines: transcription scripts for multimedia, extraction utilities for ebooks, and direct API integration for web sources.
  • Deep analysis mode leverages the detected content type to apply format-specific AI prompting strategies, optimizing NotebookLM output for each input medium.

Frequently Asked Questions

How does the system differentiate between a podcast URL and a regular YouTube video?

The detect_input_type function checks for specific domain substrings. URLs containing xiaoyuzhoufm.com, ximalaya.com, or bilibili.com receive the podcast classification, while youtube.com or youtu.be domains map to youtube. According to the main.py source code, both types ultimately trigger transcription workflows, though the distinct typing allows for platform-specific handling.

What happens when the system encounters an unsupported file format?

Files with extensions not explicitly mapped (anything other than .epub, .pdf, .txt, .md, .docx, .pptx, .xlsx, image formats, audio formats, or .zip) return the unknown type from detect_input_type. The routing logic typically skips processing for unknown types or treats them as search keywords if the path does not exist on the filesystem.

Can the detection system handle relative file paths or only absolute paths?

The function utilizes Path(arg) and path.exists() checks, which resolve both relative and absolute paths according to Python's pathlib semantics. As implemented in joeseesun/qiaomu-anything-to-notebooklm, the extension-based detection triggers for any existing file path regardless of whether it is relative to the working directory or fully qualified.

Where does the transcription logic for audio and video content reside?

Multimedia transcription implementations live in scripts/get_podcast_transcript.py, which the main routine calls when processing podcast or youtube content types. For X/Twitter content retrieval, the system executes scripts/fetch_url.sh as a shell subprocess to fetch the tweet text before ingestion into NotebookLM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →