What Input Types Does the detect_input_type Function Support?

The detect_input_type function supports six distinct input categories: PDF documents, Microsoft Word files (.doc/.docx), plain text (.txt), Markdown (.md), HTML (.html/.htm), and web URLs (http/https), routing each to specialized processing pipelines based on extension or URL pattern analysis.

The detect_input_type function in the joeseesun/qiaomu-anything-to-notebooklm repository serves as the content ingestion gateway, automatically classifying incoming files and URLs to determine appropriate processing strategies. Located in main.py, this utility inspects file extensions and URL schemes to categorize resources before they enter the extraction pipeline.

How detect_input_type Classifies Content

The function operates by inspecting the input string for specific file extensions or URL schemes. According to the source code in main.py, it applies pattern matching to categorize resources into document-specific processing queues.

File-Based Document Types

For local file paths, the function recognizes the following extensions:

  • PDF: Files ending with .pdf
  • Word Documents: Both legacy .doc and modern .docx formats
  • Plain Text: Files using the .txt extension
  • Markdown: Documentation files with .md extensions
  • HTML: Web archives and markup files using .html or .htm

Web Resources

When the input string begins with http:// or https://, the function classifies it as a URL type. This triggers web scraping protocols rather than local file parsing, enabling the ingestion of remote content.

Fallback Handling

Any input that fails to match recognized extensions or URL patterns is categorized as unknown. This fallback mechanism prevents pipeline crashes when encountering unsupported formats.

Implementation Details in main.py

The classification logic is implemented at line 16 of main.py, where detect_input_type parses the input path to determine the resource type. The function's output directly influences processing decisions at line 324, determining which extraction library—such as PyPDF2 for PDFs or python-docx for Word files—handles the subsequent content processing.

Practical Usage Examples


# Detecting a PDF file

input_path = "reports/annual_report.pdf"
input_type = detect_input_type(input_path)
print(input_type)      # Output: "pdf"

# Detecting a Markdown document

input_path = "notes/project_overview.md"
input_type = detect_input_type(input_path)
print(input_type)      # Output: "markdown"

# Detecting a web URL

input_path = "https://github.com/joeseesun/qiaomu-anything-to-notebooklm"
input_type = detect_input_type(input_path)
print(input_type)      # Output: "url"

Summary

  • The detect_input_type function in joeseesun/qiaomu-anything-to-notebooklm supports six input categories: PDF, Word (.doc/.docx), plain text, Markdown, HTML, and URLs.
  • Located at line 16 of main.py, the function inspects file extensions and URL schemes to determine processing routes.
  • Supported extensions include .pdf, .doc, .docx, .txt, .md, .html, and .htm.
  • URL detection triggers when inputs start with http:// or https://.
  • Unmatched inputs fall back to an unknown category for robust error handling.

Frequently Asked Questions

What file extensions does detect_input_type recognize?

The function recognizes .pdf, .doc, .docx, .txt, .md, .html, and .htm. It maps these to specific document types to ensure the correct parsing library is invoked for content extraction.

How does the function handle unsupported file types?

When an input lacks a recognized extension or valid URL prefix, detect_input_type returns an unknown classification. This allows the application to gracefully handle edge cases without causing pipeline failures.

Where is detect_input_type defined in the codebase?

The function is defined at line 16 of main.py in the joeseesun/qiaomu-anything-to-notebooklm repository. Its classification results are consumed downstream at line 324, where they determine which content extraction strategy executes.

Can detect_input_type distinguish between .doc and .docx formats?

Yes, the function recognizes both .doc and .docx extensions and categorizes both as Word document types. While the specific extension is preserved in the input path, the internal classification groups them together for processing by Word-compatible extraction libraries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →