# PDF to Markdown Extraction Process Using PyMuPDF in the Hiring Agent

> Learn how Hiring Agent uses PyMuPDF to convert PDF résumés to Markdown. Discover text extraction, header detection, and layout preservation techniques.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-06-27

---

**The Hiring Agent converts résumé PDFs into structured Markdown by opening documents with PyMuPDF, detecting headers via font-size analysis, extracting text while preserving layout columns, and processing embedded images and tables before assembling the final output.**

The interviewstreet/hiring-agent repository implements a sophisticated document parsing pipeline that transforms unstructured PDF résumés into clean Markdown suitable for large language model (LLM) consumption. At the core of this system lies a PyMuPDF-based extraction engine that preserves semantic structure while filtering visual noise. This article explains the complete PDF to Markdown extraction process using PyMuPDF, referencing the actual implementation in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) and [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py).

## High-Level Architecture

The extraction pipeline is orchestrated by `PDFHandler` in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), which delegates the heavy lifting to the `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py). This separation of concerns allows the handler to manage document lifecycle and LLM interactions while the conversion engine focuses on content extraction.

### Document Initialization

When `PDFHandler.extract_text_from_pdf` receives a file path, it initializes a PyMuPDF `Document` object using `pymupdf.open(pdf_path)` (lines 52-56 of [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)). The handler immediately establishes a page range iterator to support partial document processing.

```python
with pymupdf.open(pdf_path) as doc:  # pdf.py → line 52-56

    pages = range(doc.page_count)

```

The document and page range are then passed to `to_markdown` (lines 54-57).

```python
resume_text = to_markdown(doc, pages=pages)  # pdf.py → line 54-57

```

## The Conversion Pipeline

The `to_markdown` function executes an eight-stage pipeline to transform raw PDF content into semantic Markdown.

### Header Detection via Font Analysis

Before extracting text, the system analyzes typographic hierarchy using the `IdentifyHeaders` class (lines 84-90 and 153-164 of [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)). This utility scans every page to count characters per font size, designates the most frequent size as body text, and maps larger sizes to Markdown header prefixes (`#` through `######`). For documents with existing table of contents structures, the alternative `TocHeaders` class provides mapping based on document metadata rather than font metrics. Header mapping is applied later via `get_header_id`.

### Layout Normalization

For re-flowable PDFs, the engine forces content into a single tall page layout using `doc.layout` (around lines 88-107). This stage normalizes margins, clears page rotation metadata, and optionally detects background colors to distinguish content regions from decorative elements.

### Visual Element Extraction

The system identifies non-textual content through three specialized collectors:

- **Images**: Gathered via `page.get_image_info()` and filtered by dimensional thresholds. Images are either persisted to disk using `write_images` or embedded as base64 data URIs via `embed_images` (lines 131-165).
- **Vector Graphics**: Collected through `page.get_drawings()`, spatially clustered using `cluster_drawings`, and filtered for significance via `is_significant` (lines 182-210).
- **Tables**: Detected using `page.find_tables()` and rendered to Markdown syntax through `Table.to_markdown` (lines 230-255).

### Text Extraction and Formatting

Text processing begins with `page.get_textpage` for high-performance access to raw content. The `column_boxes` function computes reading-order column boundaries while excluding rectangles occupied by images, graphics, and tables. The `write_text` helper (starting at line 298) iterates line-by-line through text rectangles, applying inline formatting (bold, italic, code, strike-through) and inserting appropriate header prefixes via `get_header_id`. Hyperlinks within text spans are resolved to Markdown link syntax using `resolve_links`.

### Final Assembly

In the concluding stages (lines 984-1000 of [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)), the engine interleaves table and image markdown with processed text while respecting page order and separator directives. The system cleans duplicate whitespace and removes stray control characters before returning the finalized Markdown string to `PDFHandler`.

## Implementation Examples

You can extract Markdown from a résumé PDF using the high-level `PDFHandler` wrapper:

```python
from pdf import PDFHandler

handler = PDFHandler()

# Convert PDF → Markdown

markdown = handler.extract_text_from_pdf("resume.pdf")
print("--- Markdown preview ---")
print(markdown[:500])  # first 500 characters

# Convert Markdown → JSON resume (uses LLM behind the scenes)

json_resume = handler.extract_json_from_pdf("resume.pdf")
print("\n--- Structured resume ---")
print(json_resume)  # pydantic model representation

```

For direct access to the conversion engine without the LLM pipeline:

```python
import pymupdf
from pymupdf_rag import to_markdown

doc = pymupdf.open("resume.pdf")
md = to_markdown(doc)  # full document

print(md)

```

## Key Components and Responsibilities

The extraction system spans five primary files:

- **[`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)**: Contains `PDFHandler`, the high-level wrapper that opens PDFs, invokes `to_markdown`, and orchestrates LLM section extraction.
- **[`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)**: Houses the core PDF-to-Markdown engine, including `to_markdown`, `IdentifyHeaders`, and handlers for images, tables, and vector graphics.
- **[`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py)**: Normalizes LLM JSON output into the internal `JSONResume` model after Markdown extraction completes.
- **[`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)**: Defines Pydantic schemas for `JSONResume`, `Basics`, `Work`, and other structured data types.
- **[`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py)**: Manages system-message and user-prompt templates delivered to the LLM after Markdown extraction.

## Summary

- The **PDF to Markdown extraction process using PyMuPDF** begins with `PDFHandler` opening documents via `pymupdf.open()` and delegating to `to_markdown`.
- **Header detection** relies on `IdentifyHeaders` to map font sizes to Markdown header levels (`#` through `######`).
- **Visual elements** including images, vector graphics, and tables are extracted via `page.get_image_info()`, `page.get_drawings()`, and `page.find_tables()` respectively.
- **Text extraction** uses `page.get_textpage` and `column_boxes` to respect multi-column layouts while avoiding overlaid graphics.
- The final output is assembled in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 984-1000) and returned to `PDFHandler` for downstream LLM processing.

## Frequently Asked Questions

### How does the system determine which text should be formatted as headers?

The `IdentifyHeaders` class in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) analyzes the entire document to count character frequencies per font size, establishing the most common size as body text. Any text rendered in larger font sizes is mapped to Markdown header prefixes (`#` through `######`) via the `get_header_id` method (lines 153-164). Alternatively, documents with existing table of contents metadata can use the `TocHeaders` class for structure detection.

### What happens to images embedded in the PDF?

Images are extracted using `page.get_image_info()` and filtered by size thresholds to exclude icons or decorative elements. Depending on configuration, the system either saves images to disk via `write_images` or embeds them directly as base64 data URIs using `embed_images` (lines 131-165 of [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)).

### How does the extraction handle complex table layouts?

Tables are detected using PyMuPDF's `page.find_tables()` method and converted to Markdown syntax through `Table.to_markdown` (lines 230-255). During text extraction, table regions are excluded from column boundary calculations to prevent text interference, ensuring clean separation between tabular data and flowing prose.

### Can the extraction process handle rotated or multi-column PDFs?

Yes. The pipeline explicitly clears rotation metadata during layout normalization (around lines 88-107). For multi-column documents, the `column_boxes` function computes reading-order boundaries while avoiding rectangles occupied by images, graphics, and tables, preserving the logical text flow regardless of physical layout complexity.