# How PDF Text Extraction Using PyMuPDF Works in the Hiring-Agent Repository

> Learn how the hiring-agent repository extracts PDF text with PyMuPDF. Discover the PDFHandler class and to_markdown function for structured Markdown conversion.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-05

---

**The hiring-agent repository implements PDF text extraction using PyMuPDF by combining a high-level `PDFHandler` class with a specialized `to_markdown` conversion function that transforms document pages into structured Markdown while automatically detecting headers, tables, and graphics.**

The `interviewstreet/hiring-agent` project processes resume documents through a dedicated pipeline for **PDF text extraction using PyMuPDF** (the Python binding for MuPDF). This system converts binary PDF files into clean Markdown format, enabling downstream LLM-based parsing of candidate information. The architecture separates file handling concerns from document layout analysis, utilizing PyMuPDF's `fitz` interface for low-level page access.

## Core Architecture Components

The extraction system relies on two primary components that work sequentially to process documents.

### PDFHandler High-Level API

The **`PDFHandler`** class in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) provides the main entry point for document processing. Its `extract_text_from_pdf` method (lines 47-58) orchestrates the opening of PDF files and delegates conversion to the markdown generator. This wrapper handles file I/O context management and page range selection before invoking the core transformation logic.

### to_markdown Conversion Engine

The **`to_markdown`** function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 302-336) performs the heavy lifting of content extraction. It creates a `Parameters` dataclass to store page-level metadata, iterates through specified pages using `page.get_textpage()`, and applies layout detection algorithms including `IdentifyHeaders` for heading hierarchy and `page.find_tables()` for tabular data recognition.

## The PDF Text Extraction Pipeline

The process follows a four-stage pipeline from file ingestion to structured text output.

### 1. Document Opening and Validation

PyMuPDF opens the file stream through its `pymupdf.open()` context manager. As implemented in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 52-53), this creates a `Document` object that provides access to page count, metadata, and content streams while ensuring proper resource cleanup.

```python
with pymupdf.open(pdf_path) as doc:
    pages = range(doc.page_count)

```

### 2. Page Selection and Preparation

The handler extracts all pages by default using `range(doc.page_count)`, though the `to_markdown` signature accepts a customizable `pages` parameter for selective extraction (lines 53-56 in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)). This allows processing of specific page ranges without loading unnecessary content into memory.

### 3. Content Processing and Layout Detection

The `to_markdown` function processes each page through several specialized subsystems:

- **Text Block Extraction**: Uses `get_raw_lines` and `column_boxes` helpers to extract raw text while preserving reading order
- **Header Detection**: The `IdentifyHeaders` class analyzes font size hierarchies to automatically assign Markdown heading levels (H1-H6)
- **Table Recognition**: Calls `page.find_tables()` to detect tabular structures and render them as Markdown tables
- **Image Handling**: Processes graphics either by saving to disk (`write_images=True`), embedding as Base64 (`embed_images=True`), or skipping entirely (`ignore_images=True`)

### 4. Markdown Assembly and Return

The function aggregates processed blocks into a single Markdown string, which `extract_text_from_pdf` returns to callers for ingestion by the LLM resume parser defined in subsequent pipeline stages.

## Implementation Examples

### Using the PDFHandler Wrapper

For standard use cases, instantiate `PDFHandler` to leverage automatic resource management and logging.

```python
from pdf import PDFHandler

handler = PDFHandler()
markdown_text = handler.extract_text_from_pdf("samples/resume.pdf")

if markdown_text:
    print("Extracted Markdown (first 200 chars):")
    print(markdown_text[:200])

```

This approach handles the `pymupdf.open()` context, page enumeration, and error handling automatically.

### Direct PyMuPDF Integration

For fine-grained control over page ranges or conversion parameters, invoke `to_markdown` directly.

```python
import pymupdf
from pymupdf_rag import to_markdown

doc = pymupdf.open("samples/resume.pdf")

# Extract only pages 0-4 (first five pages)

markdown = to_markdown(doc, pages=range(5))

print(markdown)

```

### Optimizing for Text-Only Extraction

When processing speed takes priority over image content, disable graphic extraction to reduce I/O overhead.

```python
markdown = to_markdown(
    doc,
    pages=None,                     # all pages

    write_images=False,
    embed_images=False,
    ignore_images=True,             # skip image extraction

    force_text=True,                # still try OCR on image-only pages

)

```

Setting `ignore_images=True` bypasses the `write_images` and `embed_images` processing paths, significantly accelerating throughput for text-heavy documents.

## Key Source Files and Responsibilities

| File | Primary Concern |
|------|-----------------|
| [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) | **PDFHandler** class and high-level orchestration of LLM extraction |
| [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) | Core conversion logic including `to_markdown`, header detection, and table processing |
| [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) | Pydantic models receiving parsed JSON after text extraction |
| [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) | Helper utilities converting LLM JSON output to `JSONResume` objects |
| [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) | Downstream LLM interaction utilities consuming the extracted Markdown |

These modules collectively form the complete processing chain: **PDF → PyMuPDF → Markdown → LLM → Structured Resume JSON**.

## Summary

- The **`PDFHandler.extract_text_from_pdf`** method in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) serves as the primary interface for **PDF text extraction using PyMuPDF**, handling file opening and page selection.
- The **`to_markdown`** function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) performs content conversion using `page.get_textpage()`, `IdentifyHeaders`, and `page.find_tables()`.
- **Layout detection** automatically distinguishes between headers, paragraphs, tables, and graphics without manual template configuration.
- ** Performance optimization** is achievable through parameters like `ignore_images=True` and selective page ranging.
- The extracted Markdown feeds directly into the repository's LLM pipeline defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) and [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py).

## Frequently Asked Questions

### How does the system handle PDF files with complex multi-column layouts?

The `to_markdown` function utilizes `column_boxes` and `get_raw_lines` helpers to analyze page geometry and preserve proper reading order across columns. PyMuPDF's `TextPage` extraction captures positional data that the conversion logic uses to reconstruct linear text flow, ensuring that content reads correctly even when source documents use sophisticated print-style layouts.

### What is the difference between the PDFHandler class and the to_markdown function?

**`PDFHandler`** (in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)) provides a high-level abstraction that manages file contexts, page enumeration, and error handling, making it ideal for standard batch processing. **`to_markdown`** (in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)) is the low-level conversion engine that accepts an already-opened PyMuPDF `Document` object and exposes granular parameters for image handling, page selection, and layout detection, suited for custom integration scenarios.

### Can I extract text from specific pages rather than the entire document?

Yes. While `PDFHandler` defaults to `range(doc.page_count)`, you can pass a custom `pages` parameter directly to `to_markdown`. Supply any iterable of page indices (e.g., `range(5)` for the first five pages or `[0, 2, 4]` for non-sequential pages) to process only specific sections of large PDF files.

### How are images and graphics processed during text extraction?

The system supports three image handling modes controlled via boolean flags: **`write_images=True`** saves extracted images to disk with file references in the Markdown; **`embed_images=True`** encodes graphics as Base64 data URIs within the document; and **`ignore_images=True`** skips image extraction entirely for faster text-only processing. When `force_text=True` is set alongside image flags, the system attempts OCR on image-only pages using PyMuPDF's text extraction capabilities.