How the Hiring Agent Converts PDF to Markdown: A Technical Deep Dive
The Hiring Agent converts PDF résumés to clean Markdown through a two-step pipeline that uses PyMuPDF for document loading and a custom to_markdown function for semantic extraction and formatting.
The interviewstreet/hiring-agent repository implements a sophisticated document processing pipeline that transforms static PDF résumés into structured Markdown text. This conversion enables downstream LLM parsers to extract candidate information section-by-section. Understanding how this system converts PDF to Markdown reveals the intricate logic behind modern document parsing architectures.
The Two-Step PDF to Markdown Pipeline
The conversion process orchestrates two distinct phases: document ingestion and semantic Markdown generation.
Step 1: PDF Loading and Page Selection with PDFHandler
The PDFHandler class in pdf.py initiates the process by opening the PDF using PyMuPDF (pymupdf.open). According to the source code at PDFHandler.extract_text_from_pdf, this method builds a list of page numbers to process and validates the document structure before extraction.
Step 2: Markdown Conversion via pymupdf_rag.py
After loading, the selected document and page range pass to the to_markdown function defined in pymupdf_rag.py. This implementation (to_markdown) serves as the primary entry point for converting PDF content to Markdown format.
Technical Implementation Details
The conversion logic extends beyond simple text extraction to preserve document semantics.
Header Hierarchy Detection via Font Size Analysis
The IdentifyHeaders class analyzes font sizes across the document to detect hierarchical structure. By comparing text metrics, the system distinguishes between H1, H2, and H3 equivalents in the original PDF, mapping them to appropriate Markdown heading syntax.
Table, Image, and Vector Graphics Extraction
The to_markdown function extracts tables, images, and vector graphics during the page walk. It renders these elements into Markdown-compatible formats while maintaining their positional relationships within the document flow.
Text Formatting and Hyperlink Resolution
As the parser walks the page column-by-column, it recognizes formatting patterns including bullet lists, code blocks, bold, italic, and strike-through styling. The write_text utility resolves hyperlinks into proper Markdown syntax ([text](url)) and assembles extracted pieces into a clean Markdown string, stripping artifacts and normalizing whitespace.
Code Examples
Hiring Agent exposes both high-level and low-level APIs for PDF conversion.
Using PDFHandler for High-Level Conversion
# Example: Convert a PDF résumé to Markdown using the Hiring Agent
from pdf import PDFHandler
pdf_path = "candidate_resume.pdf"
handler = PDFHandler()
# Step 1 – extract raw Markdown from the PDF
markdown_resume = handler.extract_text_from_pdf(pdf_path)
print(markdown_resume[:500]) # show the first 500 characters
Direct Low-Level Conversion with to_markdown
# Direct use of the low-level conversion utility
import pymupdf
from pymupdf_rag import to_markdown
doc = pymupdf.open("candidate_resume.pdf")
# Convert the whole document (pages are zero-based)
md = to_markdown(doc, pages=range(doc.page_count))
print(md) # full Markdown representation of the PDF
Summary
- The Hiring Agent converts PDF to Markdown using a two-step pipeline involving
PDFHandlerandto_markdown - PyMuPDF (
pymupdf.open) handles document loading and page selection inpdf.py - The
to_markdownfunction inpymupdf_rag.pyhandles header detection, table extraction, and formatting preservation - Supporting classes like
IdentifyHeadersandwrite_textmanage semantic analysis and text assembly - The resulting Markdown feeds into LLM parsers for structured resume extraction in
transform.py
Frequently Asked Questions
What library does the Hiring Agent use to open PDF files?
The Hiring Agent uses PyMuPDF (pymupdf.open) to open and process PDF documents. This library provides the low-level document manipulation capabilities required for extracting text, fonts, and layout information.
How does the Hiring Agent detect headings in PDF documents?
The system uses the IdentifyHeaders class to analyze font sizes throughout the document. By comparing text metrics and identifying size hierarchies, it maps PDF text styles to appropriate Markdown heading levels (H1, H2, H3).
Can the Hiring Agent extract tables and images from PDFs?
Yes. The to_markdown function in pymupdf_rag.py extracts and renders tables, images, and vector graphics during the conversion process. It preserves these elements within the Markdown output while maintaining their document context.
Where does the converted Markdown go after PDF processing?
After conversion, the Markdown text passes to transform.py for post-processing. This module parses the Markdown into structured JSON resume objects, enabling section-wise extraction of candidate information such as work history, education, and contact details.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →