# What Libraries Are Used for PDF to Markdown Conversion in Hiring Agent?

> Discover how Hiring Agent converts PDFs to Markdown using PyMuPDF and pymupdf4llm. Learn about the libraries enabling efficient resume parsing and structured text generation.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-02

---

**Hiring Agent uses PyMuPDF (pymupdf) version 1.26.3 for low-level PDF parsing and pymupdf4llm version 0.0.27 for high-level Markdown generation, working together in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) and [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) to convert resume PDFs into clean, structured Markdown.**

The interviewstreet/hiring-agent repository processes candidate resumes by converting PDF documents into Markdown format suitable for LLM prompting. This conversion relies on a two-stage pipeline that separates raw document parsing from intelligent content formatting. Understanding these libraries and their interaction is essential for anyone modifying the resume ingestion workflow or troubleshooting extraction quality.

## The Two-Library Architecture for PDF Processing

Hiring Agent employs a layered approach to PDF conversion, delegating specific responsibilities to specialized libraries.

### PyMuPDF (pymupdf) — Low-Level PDF Parsing

**PyMuPDF** serves as the foundational engine for document access. In [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt), the project pins this dependency to version `1.26.3`. This library handles the critical first step of opening PDF files and extracting raw page content, including text, images, vector graphics, and metadata. The `PDFHandler.extract_text_from_pdf` method in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) initiates the process by calling `pymupdf.open()` to load the document into memory, creating a document object that subsequent processing layers consume.

### pymupdf4llm — High-Level Markdown Generation

**pymupdf4llm** (version `0.0.27` as specified in [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt)) provides the intelligent formatting layer. Rather than returning unstructured text, this library analyzes the visual layout using PyMuPDF's underlying data structures to recognize structural elements. The core routine `to_markdown`, defined in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py), orchestrates this transformation. It employs helpers such as `get_raw_lines` for text extraction, `column_boxes` for layout detection, and `IdentifyHeaders` to compute header levels from font size variations.

## How the Conversion Pipeline Works in Code

The actual execution flow spans two primary source files, demonstrating a clear separation between interface and implementation.

### Entry Point: PDFHandler.extract_text_from_pdf

The `PDFHandler` class in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) provides the high-level interface that the rest of the application calls. When processing a resume, `extract_text_from_pdf` opens the target document using PyMuPDF, selects the relevant pages (supporting partial document processing), and passes the document object along with page ranges to the conversion layer. This method abstracts the complexity of the underlying libraries, offering a simple API for PDF ingestion.

### Core Logic: to_markdown in pymupdf_rag.py

The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) executes the heavy lifting of format conversion. This function receives the PyMuPDF document object and iterates through specified pages, applying several analytical steps:

- **Layout Analysis**: Uses `column_boxes` to detect text columns and distinguish body text from tables or graphics.
- **Header Detection**: Leverages `IdentifyHeaders` to calculate hierarchical levels based on font size deltas across the document.
- **Content Rendering**: Calls `write_text` to emit Markdown syntax, with specialized handlers for `output_tables` and image processing.

The function returns a complete Markdown string that preserves the document's structural hierarchy, making resumes machine-readable for downstream LLM processing.

## Practical Implementation Example

The following code demonstrates the exact pattern used within Hiring Agent's PDF processing workflow:

```python
from pymupdf import open as open_pdf          # PyMuPDF v1.26.3

from pymupdf_rag import to_markdown            # pymupdf4llm v0.0.27

pdf_path = "candidate_resume.pdf"

# Stage 1: Open PDF with PyMuPDF (equivalent to PDFHandler.extract_text_from_pdf)

doc = open_pdf(pdf_path)

# Stage 2: Convert to Markdown (as implemented in pymupdf_rag.py)

markdown_output = to_markdown(
    doc,
    pages=range(doc.page_count),   # Process all pages; accepts list for subsets

    hdr_info=None,                 # Auto-detect headers from font metrics

    write_images=False,            # Optional: save images to filesystem

    embed_images=False,            # Optional: base64-encode images inline

    ignore_images=False,
    ignore_graphics=False,
    detect_bg_color=True,
)

print(markdown_output[:1000])      # Preview first 1000 characters

```

This implementation mirrors the production code in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), showing how the application combines PyMuPDF's document handling with pymupdf4llm's formatting intelligence.

## Key Files in the Conversion Pipeline

Three critical files define Hiring Agent's PDF processing capabilities:

- **[`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)**: Contains the `PDFHandler` class with `extract_text_from_pdf`, serving as the primary entry point for resume ingestion.
- **[`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)**: Houses the `to_markdown` function and all supporting logic for Markdown generation, including header detection and table formatting.
- **[`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt)**: Pins the exact library versions (`pymupdf==1.26.3` and `pymupdf4llm==0.0.27`) ensuring reproducible PDF processing behavior across deployments.

## Summary

- **Hiring Agent uses two complementary libraries**: PyMuPDF for raw PDF parsing and pymupdf4llm for intelligent Markdown generation.
- **Version pinning matters**: The project specifically requires PyMuPDF `1.26.3` and pymupdf4llm `0.0.27` as defined in [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt).
- **Processing occurs across two files**: [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) handles high-level document management while [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) contains the core `to_markdown` logic.
- **Pipeline preserves document structure**: The conversion detects headers, tables, and columns using font metrics and layout analysis rather than simple text extraction.

## Frequently Asked Questions

### What version of PyMuPDF does Hiring Agent use?

Hiring Agent pins PyMuPDF to version `1.26.3` in its [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt) file. This specific version ensures stable behavior for the low-level PDF parsing operations that feed into the Markdown conversion pipeline.

### How does Hiring Agent handle tables and images during PDF conversion?

The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) includes specialized logic for both elements. It uses `output_tables` to render tabular data as Markdown tables and provides parameters like `write_images` and `embed_images` to control whether images are saved as external files, embedded as base64 strings, or ignored entirely based on the configuration passed from `PDFHandler`.

### What is the difference between PyMuPDF and pymupdf4llm?

**PyMuPDF** is a general-purpose PDF library that provides access to document content at the structural level. **pymupdf4llm** is a specialized wrapper that builds on PyMuPDF to specifically convert PDF content into clean Markdown format, adding capabilities like header detection (`IdentifyHeaders`) and column analysis (`column_boxes`) that raw PyMuPDF does not provide.

### Can the PDF to Markdown conversion be configured to ignore specific elements?

Yes. The `to_markdown` function accepts boolean flags including `ignore_images` and `ignore_graphics` that allow the caller to skip specific content types. The `PDFHandler` in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) can pass these parameters based on application requirements, enabling flexible processing for different resume formats or downstream LLM contexts.