# What Input Types Does the detect_input_type Function Support?

> Discover the input types the detect_input_type function supports: PDF, Word, TXT, MD, HTML, and web URLs. Learn how it processes various file formats for notebook integration.

- Repository: [向阳乔木/qiaomu-anything-to-notebooklm](https://github.com/joeseesun/qiaomu-anything-to-notebooklm)
- Tags: api-reference
- Published: 2026-05-16

---

**The `detect_input_type` function supports six distinct input categories: PDF documents, Microsoft Word files (.doc/.docx), plain text (.txt), Markdown (.md), HTML (.html/.htm), and web URLs (http/https), routing each to specialized processing pipelines based on extension or URL pattern analysis.**

The `detect_input_type` function in the `joeseesun/qiaomu-anything-to-notebooklm` repository serves as the content ingestion gateway, automatically classifying incoming files and URLs to determine appropriate processing strategies. Located in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py), this utility inspects file extensions and URL schemes to categorize resources before they enter the extraction pipeline.

## How detect_input_type Classifies Content

The function operates by inspecting the input string for specific file extensions or URL schemes. According to the source code in [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py), it applies pattern matching to categorize resources into document-specific processing queues.

### File-Based Document Types

For local file paths, the function recognizes the following extensions:

- **PDF**: Files ending with `.pdf`
- **Word Documents**: Both legacy `.doc` and modern `.docx` formats
- **Plain Text**: Files using the `.txt` extension
- **Markdown**: Documentation files with `.md` extensions
- **HTML**: Web archives and markup files using `.html` or `.htm`

### Web Resources

When the input string begins with `http://` or `https://`, the function classifies it as a **URL** type. This triggers web scraping protocols rather than local file parsing, enabling the ingestion of remote content.

### Fallback Handling

Any input that fails to match recognized extensions or URL patterns is categorized as **unknown**. This fallback mechanism prevents pipeline crashes when encountering unsupported formats.

## Implementation Details in main.py

The classification logic is implemented at **line 16 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)**, where `detect_input_type` parses the input path to determine the resource type. The function's output directly influences processing decisions at **line 324**, determining which extraction library—such as `PyPDF2` for PDFs or `python-docx` for Word files—handles the subsequent content processing.

## Practical Usage Examples

```python

# Detecting a PDF file

input_path = "reports/annual_report.pdf"
input_type = detect_input_type(input_path)
print(input_type)      # Output: "pdf"

```

```python

# Detecting a Markdown document

input_path = "notes/project_overview.md"
input_type = detect_input_type(input_path)
print(input_type)      # Output: "markdown"

```

```python

# Detecting a web URL

input_path = "https://github.com/joeseesun/qiaomu-anything-to-notebooklm"
input_type = detect_input_type(input_path)
print(input_type)      # Output: "url"

```

## Summary

- The `detect_input_type` function in `joeseesun/qiaomu-anything-to-notebooklm` supports **six input categories**: PDF, Word (.doc/.docx), plain text, Markdown, HTML, and URLs.
- Located at **line 16 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)**, the function inspects file extensions and URL schemes to determine processing routes.
- Supported extensions include `.pdf`, `.doc`, `.docx`, `.txt`, `.md`, `.html`, and `.htm`.
- URL detection triggers when inputs start with `http://` or `https://`.
- Unmatched inputs fall back to an **unknown** category for robust error handling.

## Frequently Asked Questions

### What file extensions does detect_input_type recognize?

The function recognizes `.pdf`, `.doc`, `.docx`, `.txt`, `.md`, `.html`, and `.htm`. It maps these to specific document types to ensure the correct parsing library is invoked for content extraction.

### How does the function handle unsupported file types?

When an input lacks a recognized extension or valid URL prefix, `detect_input_type` returns an **unknown** classification. This allows the application to gracefully handle edge cases without causing pipeline failures.

### Where is detect_input_type defined in the codebase?

The function is defined at **line 16 of [`main.py`](https://github.com/joeseesun/qiaomu-anything-to-notebooklm/blob/main/main.py)** in the `joeseesun/qiaomu-anything-to-notebooklm` repository. Its classification results are consumed downstream at line 324, where they determine which content extraction strategy executes.

### Can detect_input_type distinguish between .doc and .docx formats?

Yes, the function recognizes both `.doc` and `.docx` extensions and categorizes both as Word document types. While the specific extension is preserved in the input path, the internal classification groups them together for processing by Word-compatible extraction libraries.