PDF to Markdown Extraction Process Using PyMuPDF in the Hiring Agent
The Hiring Agent converts résumé PDFs into structured Markdown by opening documents with PyMuPDF, detecting headers via font-size analysis, extracting text while preserving layout columns, and processing embedded images and tables before assembling the final output.
The interviewstreet/hiring-agent repository implements a sophisticated document parsing pipeline that transforms unstructured PDF résumés into clean Markdown suitable for large language model (LLM) consumption. At the core of this system lies a PyMuPDF-based extraction engine that preserves semantic structure while filtering visual noise. This article explains the complete PDF to Markdown extraction process using PyMuPDF, referencing the actual implementation in pymupdf_rag.py and pdf.py.
High-Level Architecture
The extraction pipeline is orchestrated by PDFHandler in pdf.py, which delegates the heavy lifting to the to_markdown function in pymupdf_rag.py. This separation of concerns allows the handler to manage document lifecycle and LLM interactions while the conversion engine focuses on content extraction.
Document Initialization
When PDFHandler.extract_text_from_pdf receives a file path, it initializes a PyMuPDF Document object using pymupdf.open(pdf_path) (lines 52-56 of pdf.py). The handler immediately establishes a page range iterator to support partial document processing.
with pymupdf.open(pdf_path) as doc: # pdf.py → line 52-56
pages = range(doc.page_count)
The document and page range are then passed to to_markdown (lines 54-57).
resume_text = to_markdown(doc, pages=pages) # pdf.py → line 54-57
The Conversion Pipeline
The to_markdown function executes an eight-stage pipeline to transform raw PDF content into semantic Markdown.
Header Detection via Font Analysis
Before extracting text, the system analyzes typographic hierarchy using the IdentifyHeaders class (lines 84-90 and 153-164 of pymupdf_rag.py). This utility scans every page to count characters per font size, designates the most frequent size as body text, and maps larger sizes to Markdown header prefixes (# through ######). For documents with existing table of contents structures, the alternative TocHeaders class provides mapping based on document metadata rather than font metrics. Header mapping is applied later via get_header_id.
Layout Normalization
For re-flowable PDFs, the engine forces content into a single tall page layout using doc.layout (around lines 88-107). This stage normalizes margins, clears page rotation metadata, and optionally detects background colors to distinguish content regions from decorative elements.
Visual Element Extraction
The system identifies non-textual content through three specialized collectors:
- Images: Gathered via
page.get_image_info()and filtered by dimensional thresholds. Images are either persisted to disk usingwrite_imagesor embedded as base64 data URIs viaembed_images(lines 131-165). - Vector Graphics: Collected through
page.get_drawings(), spatially clustered usingcluster_drawings, and filtered for significance viais_significant(lines 182-210). - Tables: Detected using
page.find_tables()and rendered to Markdown syntax throughTable.to_markdown(lines 230-255).
Text Extraction and Formatting
Text processing begins with page.get_textpage for high-performance access to raw content. The column_boxes function computes reading-order column boundaries while excluding rectangles occupied by images, graphics, and tables. The write_text helper (starting at line 298) iterates line-by-line through text rectangles, applying inline formatting (bold, italic, code, strike-through) and inserting appropriate header prefixes via get_header_id. Hyperlinks within text spans are resolved to Markdown link syntax using resolve_links.
Final Assembly
In the concluding stages (lines 984-1000 of pymupdf_rag.py), the engine interleaves table and image markdown with processed text while respecting page order and separator directives. The system cleans duplicate whitespace and removes stray control characters before returning the finalized Markdown string to PDFHandler.
Implementation Examples
You can extract Markdown from a résumé PDF using the high-level PDFHandler wrapper:
from pdf import PDFHandler
handler = PDFHandler()
# Convert PDF → Markdown
markdown = handler.extract_text_from_pdf("resume.pdf")
print("--- Markdown preview ---")
print(markdown[:500]) # first 500 characters
# Convert Markdown → JSON resume (uses LLM behind the scenes)
json_resume = handler.extract_json_from_pdf("resume.pdf")
print("\n--- Structured resume ---")
print(json_resume) # pydantic model representation
For direct access to the conversion engine without the LLM pipeline:
import pymupdf
from pymupdf_rag import to_markdown
doc = pymupdf.open("resume.pdf")
md = to_markdown(doc) # full document
print(md)
Key Components and Responsibilities
The extraction system spans five primary files:
pdf.py: ContainsPDFHandler, the high-level wrapper that opens PDFs, invokesto_markdown, and orchestrates LLM section extraction.pymupdf_rag.py: Houses the core PDF-to-Markdown engine, includingto_markdown,IdentifyHeaders, and handlers for images, tables, and vector graphics.transform.py: Normalizes LLM JSON output into the internalJSONResumemodel after Markdown extraction completes.models.py: Defines Pydantic schemas forJSONResume,Basics,Work, and other structured data types.prompts/template_manager.py: Manages system-message and user-prompt templates delivered to the LLM after Markdown extraction.
Summary
- The PDF to Markdown extraction process using PyMuPDF begins with
PDFHandleropening documents viapymupdf.open()and delegating toto_markdown. - Header detection relies on
IdentifyHeadersto map font sizes to Markdown header levels (#through######). - Visual elements including images, vector graphics, and tables are extracted via
page.get_image_info(),page.get_drawings(), andpage.find_tables()respectively. - Text extraction uses
page.get_textpageandcolumn_boxesto respect multi-column layouts while avoiding overlaid graphics. - The final output is assembled in
pymupdf_rag.py(lines 984-1000) and returned toPDFHandlerfor downstream LLM processing.
Frequently Asked Questions
How does the system determine which text should be formatted as headers?
The IdentifyHeaders class in pymupdf_rag.py analyzes the entire document to count character frequencies per font size, establishing the most common size as body text. Any text rendered in larger font sizes is mapped to Markdown header prefixes (# through ######) via the get_header_id method (lines 153-164). Alternatively, documents with existing table of contents metadata can use the TocHeaders class for structure detection.
What happens to images embedded in the PDF?
Images are extracted using page.get_image_info() and filtered by size thresholds to exclude icons or decorative elements. Depending on configuration, the system either saves images to disk via write_images or embeds them directly as base64 data URIs using embed_images (lines 131-165 of pymupdf_rag.py).
How does the extraction handle complex table layouts?
Tables are detected using PyMuPDF's page.find_tables() method and converted to Markdown syntax through Table.to_markdown (lines 230-255). During text extraction, table regions are excluded from column boundary calculations to prevent text interference, ensuring clean separation between tabular data and flowing prose.
Can the extraction process handle rotated or multi-column PDFs?
Yes. The pipeline explicitly clears rotation metadata during layout normalization (around lines 88-107). For multi-column documents, the column_boxes function computes reading-order boundaries while avoiding rectangles occupied by images, graphics, and tables, preserving the logical text flow regardless of physical layout complexity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →