What is pymupdf_rag.py? The Hiring Agent's PDF-to-Markdown Engine Explained

The pymupdf_rag.py module is the PDF-to-Markdown conversion engine that bridges raw PDF documents and LLM processing pipelines in the Hiring Agent project, parsing documents with PyMuPDF to produce structured Markdown output.

The pymupdf_rag.py module sits at the heart of the interviewstreet/hiring-agent repository's document ingestion pipeline. It transforms static PDF files into clean, structured Markdown that downstream LLM components can process and evaluate. According to the repository's architecture documentation, this conversion represents the critical first step before any AI-driven analysis occurs.

Core Responsibilities of pymupdf_rag.py

Document Parsing and Structure Detection

The module leverages PyMuPDF (v1.25.5+) to extract text while preserving document hierarchy. The IdentifyHeaders class inspects font sizes to map larger fonts to Markdown header tags (#...), while TocHeaders uses the PDF's table of contents for heading detection. This ensures that resumes, cover letters, and technical documents retain their logical structure when converted.

Rich Content Extraction

Beyond plain text, pymupdf_rag.py handles complex PDF elements:

  • Images: Extracts raster graphics and optionally embeds them as Base-64 data URIs using the format GRAPHICS_TEXT = "\n![](%s)\n"
  • Tables: Invokes page.find_tables via pymupdf4llm and renders them as valid Markdown tables
  • Formatting: Preserves bold, italic, monospaced code-like text, and strikethrough styling
  • Links: Converts PDF link annotations to standard Markdown link syntax
  • OCR Detection: Identifies OCR-only pages and background colors, with options to ignore invisible text

Command-Line and Programmatic Interfaces

The module functions as both a standalone CLI tool and a Python library, offering flexibility for different workflows.

How to Convert PDFs Using pymupdf_rag.py

Command-Line Usage

The simplest way to use the module is via the command line, processing specific page ranges:

python pymupdf_rag.py résumé.pdf -pages 1-5,10-N

This creates résumé.md containing only the selected pages from the input PDF.

Programmatic Conversion with to_markdown

For integration into Python applications, import the to_markdown function from the module:

from pymupdf_rag import to_markdown

# Convert PDF with embedded images and table extraction

markdown_text = to_markdown(
    "résumé.pdf",
    write_images=False,
    embed_images=True,
    ignore_images=False,
    ignore_graphics=False,
    table_strategy="lines_strict",
)

with open("résumé.md", "w", encoding="utf-8") as f:
    f.write(markdown_text)

Selective Page Processing

Process specific pages by passing 0-based page numbers to the pages parameter:

pages = [0, 2, 4]  # First, third, and fifth pages

markdown = to_markdown("contract.pdf", pages=pages)

Custom Header Detection

Use the table of contents for heading detection instead of font heuristics:

from pymupdf_rag import TocHeaders

hdr = TocHeaders("report.pdf")  # Use PDF TOC for headings

markdown = to_markdown("report.pdf", hdr_info=hdr)

Integration with the Hiring Agent Architecture

According to the repository's architecture documentation, pymupdf_rag.py operates as the first stage in a two-step pipeline:

  1. pymupdf_rag.py converts PDF pages to Markdown-like text
  2. pdf.py consumes this Markdown and calls the LLM per section using Jinja templates

The pdf.py module orchestrates section-wise LLM calls, while the Markdown output from pymupdf_rag.py feeds directly into the Jinja templates located in prompts/templates/. This separation of concerns ensures that the LLM receives clean, structured text rather than raw binary PDF data.

Configuration Parameters for LLM Optimization

The to_markdown function accepts numerous keyword arguments to tailor output for downstream LLM processing:

  • write_images: Save extracted images to disk
  • embed_images: Embed images as Base-64 data URIs within the Markdown
  • ignore_images: Skip image extraction entirely
  • table_strategy: Control table detection algorithm (e.g., "lines_strict")
  • ignore_graphics: Skip vector graphic processing
  • hdr_info: Custom header detector instance (IdentifyHeaders or TocHeaders)

Summary

  • The pymupdf_rag.py module is the core PDF-to-Markdown conversion engine in the interviewstreet/hiring-agent repository.
  • It uses PyMuPDF (v1.25.5+) to extract text, tables, images, and formatting while preserving document structure through IdentifyHeaders and TocHeaders.
  • The module provides both CLI and programmatic interfaces via the to_markdown function, supporting selective page processing and custom header detection.
  • It serves as the first step in the Hiring Agent pipeline, feeding structured Markdown to pdf.py for LLM evaluation via Jinja templates.
  • Configuration options like embed_images, table_strategy, and hdr_info allow fine-tuning for specific LLM processing requirements.

Frequently Asked Questions

What does pymupdf_rag.py do in the Hiring Agent project?

The pymupdf_rag.py module converts PDF documents into Markdown format using PyMuPDF. It serves as the document ingestion layer that transforms binary PDF files into structured text that the Hiring Agent's LLM pipeline can process and evaluate, handling headers, tables, images, and text formatting.

How do I use pymupdf_rag.py to convert specific pages of a PDF?

Pass a list of 0-based page numbers to the pages parameter in the to_markdown function, or use the -pages flag in the command-line interface with ranges like 1-5,10-N to select specific sections of the document rather than processing the entire file.

Can pymupdf_rag.py handle tables and images in PDFs?

Yes, the module extracts tables using page.find_tables via pymupdf4llm and renders them as Markdown tables. For images, it can either save them to disk using write_images=True or embed them as Base-64 data URIs within the Markdown using embed_images=True.

What is the relationship between pymupdf_rag.py and pdf.py?

The pymupdf_rag.py module performs the initial PDF-to-Markdown conversion, while pdf.py orchestrates the LLM evaluation phase. According to the repository architecture, pdf.py consumes the Markdown output from pymupdf_rag.py and processes it through Jinja templates in prompts/templates/ to generate AI-driven evaluations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →