How Markitdown Integration Converts Different File Formats in qiaomu-anything-to-notebooklm

The markitdown integration converts PDFs, Office documents, images, and audio files into plain text by invoking the Microsoft markitdown CLI with format-specific flags, then normalizing the output to TXT files ready for NotebookLM ingestion.

The qiaomu-anything-to-notebooklm skill leverages Microsoft's markitdown tool to transform diverse file formats into NotebookLM-compatible text. This integration routes local file paths through a decision table that selects the appropriate conversion workflow based on file type, handling everything from structured Office documents to multimedia files through a unified command-line interface.

File Type Routing and Conversion Modes

The integration determines how to process each file using the decision table defined in SKILL.md at lines 191-199. When a user supplies a local filesystem path, the skill maps the file extension to one of three markitdown sub-capabilities:

  • Structured documents (PDF, DOCX, PPTX, XLSX): Converted via markitdown … -o <out>.md → Markdown → plain text
  • Images (JPG, PNG, etc.): Processed with automatic OCR to extract text content
  • Audio files (MP3, WAV, etc.): Processed with automatic transcription to generate speech-to-text output

The skill treats the resulting Markdown or extracted text as the intermediate format, ultimately producing clean TXT files that NotebookLM can ingest directly.

The Conversion Pipeline

The markitdown integration operates through a four-stage pipeline that handles file detection, CLI execution, and output normalization.

File Detection and Routing

When parsing user queries, the skill identifies local filesystem paths and consults the mapping table in SKILL.md (lines 191-199) to determine the appropriate workflow. This routing happens before any conversion begins, ensuring the correct markitdown mode is invoked for the specific file type.

CLI Invocation

The skill constructs and executes the markitdown command using the pattern documented in README.md at line 242. The generic template follows this structure:

markitdown /path/to/file.docx -o /tmp/converted.md

For image files, markitdown automatically detects the image format and applies OCR. For audio files, it triggers the transcription engine. The command writes output to a temporary location (typically /tmp) before further processing.

Markdown-to-TXT Normalization

After markitdown generates the Markdown file, the skill strips Markdown syntax or treats the Markdown as plain text, storing the final content as a TXT file. This normalization step ensures compatibility with NotebookLM's ingestion requirements, whether the source was a Word document, PowerPoint presentation, or Excel spreadsheet.

Batch ZIP Processing

For ZIP archives, the skill first extracts the contents, then iterates over each supported file inside. According to SKILL.md at lines 259-262, the skill invokes markitdown on each entry individually and concatenates the results into a single TXT file or multiple NotebookLM sources:


# Extract archive

unzip /path/to/file.zip -d /tmp/unzipped

# Process each supported file

for f in $(find /tmp/unzipped -type f \( -iname '*.pdf' -o -iname '*.docx' \)); do
    markitdown "$f" -o "${f}.md"
done

Environment Setup and Validation

The integration requires the markitdown CLI binary to be present on the system. The check_env.py script (lines 150-171) verifies this dependency during environment checks, ensuring the tool is available before attempting conversions. The install.sh script (lines 66-69) handles installation and reports successful setup of the markitdown tool.

The Python dependency is declared in requirements.txt at line 8 as markitdown[all], which installs the full feature set including OCR and audio transcription capabilities.

Practical Code Examples

Convert a Word document to NotebookLM-ready text:


# Convert DOCX to Markdown

markitdown /path/to/file.docx -o /tmp/file.md

# Optional: Normalize to pure plain text

pandoc /tmp/file.md -t plain -o /tmp/file.txt

Extract text from an image using OCR:

markitdown /path/to/image.jpg -o /tmp/image.txt

Transcribe an audio file to text:

markitdown /path/to/audio.mp3 -o /tmp/audio.txt

Batch process a ZIP archive of documents:


# Extract and convert all supported files

unzip /path/to/archive.zip -d /tmp/unzipped
for file in $(find /tmp/unzipped -type f \( -iname '*.pdf' -o -iname '*.docx' -o -iname '*.pptx' -o -iname '*.xlsx' \)); do
    markitdown "$file" -o "${file}.txt"
done

Summary

  • The markitdown integration in qiaomu-anything-to-notebooklm provides universal file conversion through Microsoft's markitdown CLI tool.
  • File type detection occurs via the decision table in SKILL.md (lines 191-199), routing documents, images, and audio to appropriate conversion modes.
  • OCR and transcription happen automatically when markitdown detects image or audio inputs, requiring no additional flags.
  • Batch processing of ZIP archives extracts and converts multiple files sequentially, concatenating results for NotebookLM ingestion.
  • Environment validation in check_env.py (lines 150-171) ensures the markitdown binary is installed before execution.

Frequently Asked Questions

What file formats does markitdown support in this integration?

The integration supports PDF, DOCX, PPTX, and XLSX for structured document conversion. It also handles common image formats (JPG, PNG) through automatic OCR and audio formats (MP3, WAV) through automatic transcription. ZIP archives containing these file types are extracted and processed batch-wise.

How does the integration handle ZIP archives containing multiple files?

When a ZIP archive is provided, the skill extracts the contents to a temporary directory, then iterates over each supported file inside. As documented in SKILL.md at lines 259-262, markitdown is invoked on each file individually, and the resulting text is concatenated into a single output or split into multiple NotebookLM sources depending on configuration.

Is markitdown installed automatically with the skill?

The install.sh script (lines 66-69) attempts to install markitdown and reports successful setup. The check_env.py script (lines 150-171) verifies the binary is present in the environment before allowing operations. The dependency is also declared in requirements.txt (line 8) as markitdown[all] for Python environment consistency.

Why convert Markdown to TXT instead of using Markdown directly?

While markitdown outputs Markdown format, the skill normalizes to TXT to ensure consistent ingestion by NotebookLM. This normalization strips Markdown syntax or treats the content as plain text, creating a uniform text format regardless of whether the source was a PDF, Word document, or extracted image text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →