# What is PyMuPDF Used for in the Hiring Agent? PDF Parsing and Résumé Extraction Explained

> Discover how PyMuPDF powers the Hiring Agent by parsing PDF résumés and extracting key information for LLM pipelines. Understand the core PDF parsing engine.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-06-28

---

**PyMuPDF serves as the core PDF parsing engine in the Hiring Agent, converting raw résumé PDFs into structured Markdown that feeds into LLM-based extraction pipelines.**

The interviewstreet/hiring-agent repository is a résumé-parsing tool that transforms PDF documents into structured JSON objects. At the heart of this system lies **PyMuPDF** (imported as `pymupdf`), which handles all low-level PDF processing operations according to the source code. The library bridges the gap between binary PDF files and the LLM-based extraction pipeline by converting document content into clean, normalized Markdown format.

## Core PDF Processing Operations

The Hiring Agent relies on PyMuPDF for fundamental document operations that power the entire extraction workflow, from opening files to extracting text in reading order.

### Opening PDF Documents

In [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 52-57), PyMuPDF initializes document objects through `pymupdf.open(pdf_path)`. This creates a `Document` object representing the entire PDF file, which serves as the entry point for all subsequent parsing operations.

```python

# From pdf.py:52-57 - Document initialization

doc = pymupdf.open(pdf_path)

```

### Extracting Text in Reading Order

The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 33-41) orchestrates text extraction by walking through each page, collecting text blocks, and identifying structural elements like headings, tables, and code sections. This helper returns a normalized Markdown string that preserves document semantics while enabling downstream LLM processing.

```python

# Conceptual flow from pymupdf_rag.py:33-41

markdown_content = to_markdown(doc)

```

## Advanced Document Analysis Features

Beyond basic text extraction, PyMuPDF enables sophisticated document analysis capabilities critical for parsing complex résumé layouts.

### Table Detection and Conversion

PyMuPDF identifies tabular data through `page.find_tables()` as implemented in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 191-196). The Hiring Agent converts detected tables into Markdown table format, preserving structured data found in candidate résumés such as skills matrices and employment timelines.

### Image and Graphics Handling

The system detects embedded media using `page.get_image_info()` and `page.get_links()` (found in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py), lines 111-118). The `save_image` helper optionally writes images to disk or embeds them as Base-64 data URIs, ensuring visual content from portfolios or certifications remains accessible to the extraction pipeline.

### Header Detection Strategies

Two distinct header identification strategies operate within [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 70-84):

- **IdentifyHeaders**: Analyzes font sizes to detect hierarchical headings
- **TocHeaders**: Leverages the PDF's Table of Contents to generate Markdown header markers (`#`)

## OCR and Special Case Handling

PyMuPDF handles scanned documents through OCR detection logic in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 1000-1009). The `page_is_ocr` function checks for "ignore-text" glyphs generated by OCR engines, determining whether to treat a page as text-only or image-based content that requires alternative processing.

## Integration with the LLM Pipeline

The `PDFHandler.extract_text_from_pdf` method in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 47-58) serves as the integration point between PyMuPDF and the LLM. This method:
1. Opens the document using `pymupdf.open()`
2. Invokes `to_markdown` to convert content to Markdown
3. Returns raw Markdown that feeds LLM-based section extractors defined in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)

```python
from pdf import PDFHandler

# Initialize the handler

handler = PDFHandler()

# Convert PDF to markdown (PyMuPDF does the heavy lifting)

markdown = handler.extract_text_from_pdf("resume.pdf")

# Full pipeline: PDF → markdown → LLM → JSONResume

json_resume = handler.extract_json_from_pdf("resume.pdf")

```

The dependency is declared in [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt) as `pymupdf`, cementing PyMuPDF as the backbone of the Hiring Agent's PDF handling capabilities.

## Summary

- **PyMuPDF** provides the foundational PDF parsing layer for the interviewstreet/hiring-agent repository according to the source code
- The `pymupdf.open()` function in `pdf.py:52-57` initializes document processing
- **Text extraction** occurs through `to_markdown` in `pymupdf_rag.py:33-41`, which handles reading order, headers, and code blocks
- **Table detection** uses `page.find_tables()` in `pymupdf_rag.py:191-196` to preserve structured data
- **Image handling** via `page.get_image_info()` and `save_image` in `pymupdf_rag.py:111-118` processes visual content
- **OCR detection** through `page_is_ocr` in `pymupdf_rag.py:1000-1009` handles scanned documents
- The `PDFHandler` class in `pdf.py:47-58` bridges PyMuPDF output with LLM-based JSON extraction

## Frequently Asked Questions

### How does PyMuPDF handle scanned PDFs in the Hiring Agent?

PyMuPDF detects OCR-generated content through the `page_is_ocr` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 1000-1009). This helper checks for "ignore-text" glyphs that indicate OCR processing, allowing the system to determine whether to extract text directly or process the page as an image-based document requiring special handling.

### What is the difference between [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) and [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) in the Hiring Agent?

[`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) contains the high-level `PDFHandler` class that orchestrates the extraction pipeline and integrates with the LLM. [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) implements the core PyMuPDF logic, including the `to_markdown` function, header detection strategies (`IdentifyHeaders` and `TocHeaders`), table extraction, and image handling. The former manages workflow while the latter performs the actual PDF parsing operations.

### Can PyMuPDF extract tables from résumés in the Hiring Agent?

Yes, PyMuPDF identifies tables through `page.find_tables()` as implemented in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 191-196). The Hiring Agent converts these tables into Markdown format, preserving the structured layout of skills matrices, employment histories, and other tabular data found in candidate documents.

### What format does PyMuPDF output in the Hiring Agent?

PyMuPDF outputs extracted content as **Markdown** through the `to_markdown` helper function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py). This normalized format preserves document structure—including headers, tables, and code blocks—making it suitable for downstream LLM processing that ultimately generates structured JSON résumé objects.