# How the Hiring Agent Converts PDF to Markdown: A Technical Deep Dive

> Discover how the Hiring Agent converts PDF to Markdown using PyMuPDF and a custom function for semantic extraction. Learn the technical details of this efficient process.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-06-28

---

**The Hiring Agent converts PDF résumés to clean Markdown through a two-step pipeline that uses PyMuPDF for document loading and a custom `to_markdown` function for semantic extraction and formatting.**

The `interviewstreet/hiring-agent` repository implements a sophisticated document processing pipeline that transforms static PDF résumés into structured Markdown text. This conversion enables downstream LLM parsers to extract candidate information section-by-section. Understanding how this system converts PDF to Markdown reveals the intricate logic behind modern document parsing architectures.

## The Two-Step PDF to Markdown Pipeline

The conversion process orchestrates two distinct phases: document ingestion and semantic Markdown generation.

### Step 1: PDF Loading and Page Selection with PDFHandler

The `PDFHandler` class in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) initiates the process by opening the PDF using **PyMuPDF** (`pymupdf.open`). According to the source code at [`PDFHandler.extract_text_from_pdf`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L47-L58), this method builds a list of page numbers to process and validates the document structure before extraction.

### Step 2: Markdown Conversion via pymupdf_rag.py

After loading, the selected document and page range pass to the `to_markdown` function defined in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py). This implementation ([`to_markdown`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py#L32-L85)) serves as the primary entry point for converting PDF content to Markdown format.

## Technical Implementation Details

The conversion logic extends beyond simple text extraction to preserve document semantics.

### Header Hierarchy Detection via Font Size Analysis

The `IdentifyHeaders` class analyzes font sizes across the document to detect hierarchical structure. By comparing text metrics, the system distinguishes between H1, H2, and H3 equivalents in the original PDF, mapping them to appropriate Markdown heading syntax.

### Table, Image, and Vector Graphics Extraction

The `to_markdown` function extracts tables, images, and vector graphics during the page walk. It renders these elements into Markdown-compatible formats while maintaining their positional relationships within the document flow.

### Text Formatting and Hyperlink Resolution

As the parser walks the page column-by-column, it recognizes formatting patterns including bullet lists, code blocks, bold, italic, and strike-through styling. The `write_text` utility resolves hyperlinks into proper Markdown syntax (`[text](url)`) and assembles extracted pieces into a clean Markdown string, stripping artifacts and normalizing whitespace.

## Code Examples

Hiring Agent exposes both high-level and low-level APIs for PDF conversion.

### Using PDFHandler for High-Level Conversion

```python

# Example: Convert a PDF résumé to Markdown using the Hiring Agent

from pdf import PDFHandler

pdf_path = "candidate_resume.pdf"
handler = PDFHandler()

# Step 1 – extract raw Markdown from the PDF

markdown_resume = handler.extract_text_from_pdf(pdf_path)

print(markdown_resume[:500])   # show the first 500 characters

```

### Direct Low-Level Conversion with to_markdown

```python

# Direct use of the low-level conversion utility

import pymupdf
from pymupdf_rag import to_markdown

doc = pymupdf.open("candidate_resume.pdf")

# Convert the whole document (pages are zero-based)

md = to_markdown(doc, pages=range(doc.page_count))

print(md)   # full Markdown representation of the PDF

```

## Summary

- The Hiring Agent converts PDF to Markdown using a two-step pipeline involving `PDFHandler` and `to_markdown`
- PyMuPDF (`pymupdf.open`) handles document loading and page selection in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)
- The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) handles header detection, table extraction, and formatting preservation
- Supporting classes like `IdentifyHeaders` and `write_text` manage semantic analysis and text assembly
- The resulting Markdown feeds into LLM parsers for structured resume extraction in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py)

## Frequently Asked Questions

### What library does the Hiring Agent use to open PDF files?

The Hiring Agent uses **PyMuPDF** (`pymupdf.open`) to open and process PDF documents. This library provides the low-level document manipulation capabilities required for extracting text, fonts, and layout information.

### How does the Hiring Agent detect headings in PDF documents?

The system uses the `IdentifyHeaders` class to analyze font sizes throughout the document. By comparing text metrics and identifying size hierarchies, it maps PDF text styles to appropriate Markdown heading levels (H1, H2, H3).

### Can the Hiring Agent extract tables and images from PDFs?

Yes. The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) extracts and renders tables, images, and vector graphics during the conversion process. It preserves these elements within the Markdown output while maintaining their document context.

### Where does the converted Markdown go after PDF processing?

After conversion, the Markdown text passes to [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) for post-processing. This module parses the Markdown into structured JSON resume objects, enabling section-wise extraction of candidate information such as work history, education, and contact details.