# PDF Text Extraction in Hiring Agent: PyMuPDF Implementation Guide

> Learn how the Hiring Agent uses PyMuPDF for efficient PDF text extraction. This guide details the Python implementation for processing résumés with LLMs.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-09

---

**The Hiring Agent uses PyMuPDF (the `pymupdf` Python package) to extract text from PDF résumés and convert them into structured Markdown for LLM processing.**

The Hiring Agent project by interviewstreet relies on a robust PDF text extraction pipeline to parse résumé uploads. The system leverages **PyMuPDF** as its core extraction engine, transforming binary PDF documents into clean Markdown text that feeds downstream language models. This architecture ensures accurate preservation of document structure including headers, tables, and images.

## Core Extraction Architecture

The PDF processing layer is built on two primary components that handle document ingestion and text conversion.

### The PDFHandler Interface (pdf.py)

The entry point for PDF processing resides in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), which defines the `PDFHandler` class. According to the interviewstreet/hiring-agent source code, this class opens PDF documents using `pymupdf.open(...)` at lines 52‑55, creating a document object that serves as the input for further processing.

### The to_markdown Conversion Layer (pymupdf_rag.py)

Once opened, the raw pages flow to the `to_markdown` helper function implemented in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 52‑66). This function walks the document structure, identifying headers, tables, and images, then returns a Markdown‑formatted string suitable for consumption by the LLM.

## Implementation Flow

The extraction pipeline follows a three‑stage process:

1. **Document Opening**: `pymupdf.open()` creates a Document instance from the binary PDF data.
2. **Markdown Conversion**: The `to_markdown` function processes the document pages, extracting structured text.
3. **JSON Structuring**: The resulting Markdown text is consumed by the LLM to build a structured résumé JSON object.

## Code Examples

### High‑Level PDFHandler Usage

```python
from pdf import PDFHandler

handler = PDFHandler()
markdown = handler.extract_text_from_pdf("resume.pdf")
print(markdown)          # Markdown representation of the PDF content

```

### Direct PyMuPDF Access

```python
import pymupdf

doc = pymupdf.open("resume.pdf")
pages = range(doc.page_count)            # all pages

md_text = to_markdown(doc, pages=pages)  # from pymupdf_rag.py

print(md_text)

```

## Summary

- The Hiring Agent relies on **PyMuPDF** as its primary PDF text extraction library.
- The `PDFHandler` class in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) orchestrates document opening via `pymupdf.open()` at lines 52‑55.
- The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 52‑66) converts PyMuPDF documents into structured Markdown.
- This pipeline preserves document semantics (headers, tables, images) for downstream LLM processing.

## Frequently Asked Questions

### What Python library does Hiring Agent use to read PDF files?

The Hiring Agent uses **PyMuPDF** (imported as `pymupdf`) to open and read PDF files. This library provides the low‑level document parsing capabilities that power the extraction pipeline, specifically through the `pymupdf.open()` method called in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py).

### How does Hiring Agent convert PDF content to Markdown?

The conversion happens in the `to_markdown` function within [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py). This function takes a PyMuPDF Document object and a page range, then walks through the document to identify structural elements like headers and tables, outputting clean Markdown text suitable for LLM consumption.

### Where is the PDF extraction logic located in the repository?

The extraction logic is split between two files: [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) contains the `PDFHandler` class that manages document opening, while [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) contains the `to_markdown` function that handles the actual text extraction and formatting. These files constitute the complete PDF text extraction layer of the Hiring Agent.

### Why does Hiring Agent use PyMuPDF for PDF text extraction?

According to the implementation in the interviewstreet/hiring-agent repository, PyMuPDF provides the granular control needed to preserve document structure during extraction. The `to_markdown` helper can identify specific elements like tables and headers, which is essential for building structured résumé data for LLM consumption.