How the Hiring Agent Converts PDF to Markdown: A Technical Deep Dive

The Hiring Agent converts PDF résumés to clean Markdown through a two-step pipeline that uses PyMuPDF for document loading and a custom to_markdown function for semantic extraction and formatting.

The interviewstreet/hiring-agent repository implements a sophisticated document processing pipeline that transforms static PDF résumés into structured Markdown text. This conversion enables downstream LLM parsers to extract candidate information section-by-section. Understanding how this system converts PDF to Markdown reveals the intricate logic behind modern document parsing architectures.

The Two-Step PDF to Markdown Pipeline

The conversion process orchestrates two distinct phases: document ingestion and semantic Markdown generation.

Step 1: PDF Loading and Page Selection with PDFHandler

The PDFHandler class in pdf.py initiates the process by opening the PDF using PyMuPDF (pymupdf.open). According to the source code at PDFHandler.extract_text_from_pdf, this method builds a list of page numbers to process and validates the document structure before extraction.

Step 2: Markdown Conversion via pymupdf_rag.py

After loading, the selected document and page range pass to the to_markdown function defined in pymupdf_rag.py. This implementation (to_markdown) serves as the primary entry point for converting PDF content to Markdown format.

Technical Implementation Details

The conversion logic extends beyond simple text extraction to preserve document semantics.

Header Hierarchy Detection via Font Size Analysis

The IdentifyHeaders class analyzes font sizes across the document to detect hierarchical structure. By comparing text metrics, the system distinguishes between H1, H2, and H3 equivalents in the original PDF, mapping them to appropriate Markdown heading syntax.

Table, Image, and Vector Graphics Extraction

The to_markdown function extracts tables, images, and vector graphics during the page walk. It renders these elements into Markdown-compatible formats while maintaining their positional relationships within the document flow.

As the parser walks the page column-by-column, it recognizes formatting patterns including bullet lists, code blocks, bold, italic, and strike-through styling. The write_text utility resolves hyperlinks into proper Markdown syntax ([text](url)) and assembles extracted pieces into a clean Markdown string, stripping artifacts and normalizing whitespace.

Code Examples

Hiring Agent exposes both high-level and low-level APIs for PDF conversion.

Using PDFHandler for High-Level Conversion


# Example: Convert a PDF résumé to Markdown using the Hiring Agent

from pdf import PDFHandler

pdf_path = "candidate_resume.pdf"
handler = PDFHandler()

# Step 1 – extract raw Markdown from the PDF

markdown_resume = handler.extract_text_from_pdf(pdf_path)

print(markdown_resume[:500])   # show the first 500 characters

Direct Low-Level Conversion with to_markdown


# Direct use of the low-level conversion utility

import pymupdf
from pymupdf_rag import to_markdown

doc = pymupdf.open("candidate_resume.pdf")

# Convert the whole document (pages are zero-based)

md = to_markdown(doc, pages=range(doc.page_count))

print(md)   # full Markdown representation of the PDF

Summary

  • The Hiring Agent converts PDF to Markdown using a two-step pipeline involving PDFHandler and to_markdown
  • PyMuPDF (pymupdf.open) handles document loading and page selection in pdf.py
  • The to_markdown function in pymupdf_rag.py handles header detection, table extraction, and formatting preservation
  • Supporting classes like IdentifyHeaders and write_text manage semantic analysis and text assembly
  • The resulting Markdown feeds into LLM parsers for structured resume extraction in transform.py

Frequently Asked Questions

What library does the Hiring Agent use to open PDF files?

The Hiring Agent uses PyMuPDF (pymupdf.open) to open and process PDF documents. This library provides the low-level document manipulation capabilities required for extracting text, fonts, and layout information.

How does the Hiring Agent detect headings in PDF documents?

The system uses the IdentifyHeaders class to analyze font sizes throughout the document. By comparing text metrics and identifying size hierarchies, it maps PDF text styles to appropriate Markdown heading levels (H1, H2, H3).

Can the Hiring Agent extract tables and images from PDFs?

Yes. The to_markdown function in pymupdf_rag.py extracts and renders tables, images, and vector graphics during the conversion process. It preserves these elements within the Markdown output while maintaining their document context.

Where does the converted Markdown go after PDF processing?

After conversion, the Markdown text passes to transform.py for post-processing. This module parses the Markdown into structured JSON resume objects, enabling section-wise extraction of candidate information such as work history, education, and contact details.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →