What is pymupdf_rag.py? The Hiring Agent's PDF-to-Markdown Engine Explained
The pymupdf_rag.py module is the PDF-to-Markdown conversion engine that bridges raw PDF documents and LLM processing pipelines in the Hiring Agent project, parsing documents with PyMuPDF to produce structured Markdown output.
The pymupdf_rag.py module sits at the heart of the interviewstreet/hiring-agent repository's document ingestion pipeline. It transforms static PDF files into clean, structured Markdown that downstream LLM components can process and evaluate. According to the repository's architecture documentation, this conversion represents the critical first step before any AI-driven analysis occurs.
Core Responsibilities of pymupdf_rag.py
Document Parsing and Structure Detection
The module leverages PyMuPDF (v1.25.5+) to extract text while preserving document hierarchy. The IdentifyHeaders class inspects font sizes to map larger fonts to Markdown header tags (#...), while TocHeaders uses the PDF's table of contents for heading detection. This ensures that resumes, cover letters, and technical documents retain their logical structure when converted.
Rich Content Extraction
Beyond plain text, pymupdf_rag.py handles complex PDF elements:
- Images: Extracts raster graphics and optionally embeds them as Base-64 data URIs using the format
GRAPHICS_TEXT = "\n\n" - Tables: Invokes
page.find_tablesviapymupdf4llmand renders them as valid Markdown tables - Formatting: Preserves bold, italic, monospaced code-like text, and strikethrough styling
- Links: Converts PDF link annotations to standard Markdown link syntax
- OCR Detection: Identifies OCR-only pages and background colors, with options to ignore invisible text
Command-Line and Programmatic Interfaces
The module functions as both a standalone CLI tool and a Python library, offering flexibility for different workflows.
How to Convert PDFs Using pymupdf_rag.py
Command-Line Usage
The simplest way to use the module is via the command line, processing specific page ranges:
python pymupdf_rag.py résumé.pdf -pages 1-5,10-N
This creates résumé.md containing only the selected pages from the input PDF.
Programmatic Conversion with to_markdown
For integration into Python applications, import the to_markdown function from the module:
from pymupdf_rag import to_markdown
# Convert PDF with embedded images and table extraction
markdown_text = to_markdown(
"résumé.pdf",
write_images=False,
embed_images=True,
ignore_images=False,
ignore_graphics=False,
table_strategy="lines_strict",
)
with open("résumé.md", "w", encoding="utf-8") as f:
f.write(markdown_text)
Selective Page Processing
Process specific pages by passing 0-based page numbers to the pages parameter:
pages = [0, 2, 4] # First, third, and fifth pages
markdown = to_markdown("contract.pdf", pages=pages)
Custom Header Detection
Use the table of contents for heading detection instead of font heuristics:
from pymupdf_rag import TocHeaders
hdr = TocHeaders("report.pdf") # Use PDF TOC for headings
markdown = to_markdown("report.pdf", hdr_info=hdr)
Integration with the Hiring Agent Architecture
According to the repository's architecture documentation, pymupdf_rag.py operates as the first stage in a two-step pipeline:
pymupdf_rag.pyconverts PDF pages to Markdown-like textpdf.pyconsumes this Markdown and calls the LLM per section using Jinja templates
The pdf.py module orchestrates section-wise LLM calls, while the Markdown output from pymupdf_rag.py feeds directly into the Jinja templates located in prompts/templates/. This separation of concerns ensures that the LLM receives clean, structured text rather than raw binary PDF data.
Configuration Parameters for LLM Optimization
The to_markdown function accepts numerous keyword arguments to tailor output for downstream LLM processing:
write_images: Save extracted images to diskembed_images: Embed images as Base-64 data URIs within the Markdownignore_images: Skip image extraction entirelytable_strategy: Control table detection algorithm (e.g.,"lines_strict")ignore_graphics: Skip vector graphic processinghdr_info: Custom header detector instance (IdentifyHeadersorTocHeaders)
Summary
- The
pymupdf_rag.pymodule is the core PDF-to-Markdown conversion engine in the interviewstreet/hiring-agent repository. - It uses PyMuPDF (v1.25.5+) to extract text, tables, images, and formatting while preserving document structure through
IdentifyHeadersandTocHeaders. - The module provides both CLI and programmatic interfaces via the
to_markdownfunction, supporting selective page processing and custom header detection. - It serves as the first step in the Hiring Agent pipeline, feeding structured Markdown to
pdf.pyfor LLM evaluation via Jinja templates. - Configuration options like
embed_images,table_strategy, andhdr_infoallow fine-tuning for specific LLM processing requirements.
Frequently Asked Questions
What does pymupdf_rag.py do in the Hiring Agent project?
The pymupdf_rag.py module converts PDF documents into Markdown format using PyMuPDF. It serves as the document ingestion layer that transforms binary PDF files into structured text that the Hiring Agent's LLM pipeline can process and evaluate, handling headers, tables, images, and text formatting.
How do I use pymupdf_rag.py to convert specific pages of a PDF?
Pass a list of 0-based page numbers to the pages parameter in the to_markdown function, or use the -pages flag in the command-line interface with ranges like 1-5,10-N to select specific sections of the document rather than processing the entire file.
Can pymupdf_rag.py handle tables and images in PDFs?
Yes, the module extracts tables using page.find_tables via pymupdf4llm and renders them as Markdown tables. For images, it can either save them to disk using write_images=True or embed them as Base-64 data URIs within the Markdown using embed_images=True.
What is the relationship between pymupdf_rag.py and pdf.py?
The pymupdf_rag.py module performs the initial PDF-to-Markdown conversion, while pdf.py orchestrates the LLM evaluation phase. According to the repository architecture, pdf.py consumes the Markdown output from pymupdf_rag.py and processes it through Jinja templates in prompts/templates/ to generate AI-driven evaluations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →