# What is pymupdf_rag.py? The Hiring Agent's PDF-to-Markdown Engine Explained

> Discover what pymupdf_rag.py is: the PDF-to-Markdown conversion engine for Hiring Agent. It prepares PDFs for LLM processing by converting them to structured Markdown.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-07-09

---

**The [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) module is the PDF-to-Markdown conversion engine that bridges raw PDF documents and LLM processing pipelines in the Hiring Agent project, parsing documents with PyMuPDF to produce structured Markdown output.**

The [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) module sits at the heart of the **interviewstreet/hiring-agent** repository's document ingestion pipeline. It transforms static PDF files into clean, structured Markdown that downstream LLM components can process and evaluate. According to the repository's architecture documentation, this conversion represents the critical first step before any AI-driven analysis occurs.

## Core Responsibilities of pymupdf_rag.py

### Document Parsing and Structure Detection

The module leverages **PyMuPDF (v1.25.5+)** to extract text while preserving document hierarchy. The `IdentifyHeaders` class inspects font sizes to map larger fonts to Markdown header tags (`#...`), while `TocHeaders` uses the PDF's table of contents for heading detection. This ensures that resumes, cover letters, and technical documents retain their logical structure when converted.

### Rich Content Extraction

Beyond plain text, [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) handles complex PDF elements:

- **Images**: Extracts raster graphics and optionally embeds them as Base-64 data URIs using the format `GRAPHICS_TEXT = "\n![](%s)\n"`
- **Tables**: Invokes `page.find_tables` via `pymupdf4llm` and renders them as valid Markdown tables
- **Formatting**: Preserves bold, italic, monospaced code-like text, and strikethrough styling
- **Links**: Converts PDF link annotations to standard Markdown link syntax
- **OCR Detection**: Identifies OCR-only pages and background colors, with options to ignore invisible text

### Command-Line and Programmatic Interfaces

The module functions as both a standalone CLI tool and a Python library, offering flexibility for different workflows.

## How to Convert PDFs Using pymupdf_rag.py

### Command-Line Usage

The simplest way to use the module is via the command line, processing specific page ranges:

```bash
python pymupdf_rag.py résumé.pdf -pages 1-5,10-N

```

This creates `résumé.md` containing only the selected pages from the input PDF.

### Programmatic Conversion with to_markdown

For integration into Python applications, import the `to_markdown` function from the module:

```python
from pymupdf_rag import to_markdown

# Convert PDF with embedded images and table extraction

markdown_text = to_markdown(
    "résumé.pdf",
    write_images=False,
    embed_images=True,
    ignore_images=False,
    ignore_graphics=False,
    table_strategy="lines_strict",
)

with open("résumé.md", "w", encoding="utf-8") as f:
    f.write(markdown_text)

```

### Selective Page Processing

Process specific pages by passing 0-based page numbers to the `pages` parameter:

```python
pages = [0, 2, 4]  # First, third, and fifth pages

markdown = to_markdown("contract.pdf", pages=pages)

```

### Custom Header Detection

Use the table of contents for heading detection instead of font heuristics:

```python
from pymupdf_rag import TocHeaders

hdr = TocHeaders("report.pdf")  # Use PDF TOC for headings

markdown = to_markdown("report.pdf", hdr_info=hdr)

```

## Integration with the Hiring Agent Architecture

According to the repository's architecture documentation, [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) operates as the first stage in a two-step pipeline:

1. [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) converts PDF pages to Markdown-like text
2. [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) consumes this Markdown and calls the LLM per section using Jinja templates

The [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) module orchestrates section-wise LLM calls, while the Markdown output from [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) feeds directly into the Jinja templates located in `prompts/templates/`. This separation of concerns ensures that the LLM receives clean, structured text rather than raw binary PDF data.

## Configuration Parameters for LLM Optimization

The `to_markdown` function accepts numerous keyword arguments to tailor output for downstream LLM processing:

- **`write_images`**: Save extracted images to disk
- **`embed_images`**: Embed images as Base-64 data URIs within the Markdown
- **`ignore_images`**: Skip image extraction entirely
- **`table_strategy`**: Control table detection algorithm (e.g., `"lines_strict"`)
- **`ignore_graphics`**: Skip vector graphic processing
- **`hdr_info`**: Custom header detector instance (`IdentifyHeaders` or `TocHeaders`)

## Summary

- **The [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) module** is the core PDF-to-Markdown conversion engine in the interviewstreet/hiring-agent repository.
- It uses **PyMuPDF (v1.25.5+)** to extract text, tables, images, and formatting while preserving document structure through `IdentifyHeaders` and `TocHeaders`.
- The module provides both **CLI and programmatic interfaces** via the `to_markdown` function, supporting selective page processing and custom header detection.
- It serves as the **first step in the Hiring Agent pipeline**, feeding structured Markdown to [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) for LLM evaluation via Jinja templates.
- **Configuration options** like `embed_images`, `table_strategy`, and `hdr_info` allow fine-tuning for specific LLM processing requirements.

## Frequently Asked Questions

### What does pymupdf_rag.py do in the Hiring Agent project?

The [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) module converts PDF documents into Markdown format using PyMuPDF. It serves as the document ingestion layer that transforms binary PDF files into structured text that the Hiring Agent's LLM pipeline can process and evaluate, handling headers, tables, images, and text formatting.

### How do I use pymupdf_rag.py to convert specific pages of a PDF?

Pass a list of 0-based page numbers to the `pages` parameter in the `to_markdown` function, or use the `-pages` flag in the command-line interface with ranges like `1-5,10-N` to select specific sections of the document rather than processing the entire file.

### Can pymupdf_rag.py handle tables and images in PDFs?

Yes, the module extracts tables using `page.find_tables` via `pymupdf4llm` and renders them as Markdown tables. For images, it can either save them to disk using `write_images=True` or embed them as Base-64 data URIs within the Markdown using `embed_images=True`.

### What is the relationship between pymupdf_rag.py and pdf.py?

The [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) module performs the initial PDF-to-Markdown conversion, while [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) orchestrates the LLM evaluation phase. According to the repository architecture, [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) consumes the Markdown output from [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) and processes it through Jinja templates in `prompts/templates/` to generate AI-driven evaluations.