What Libraries Are Used for PDF to Markdown Conversion in Hiring Agent?
Hiring Agent uses PyMuPDF (pymupdf) version 1.26.3 for low-level PDF parsing and pymupdf4llm version 0.0.27 for high-level Markdown generation, working together in pdf.py and pymupdf_rag.py to convert resume PDFs into clean, structured Markdown.
The interviewstreet/hiring-agent repository processes candidate resumes by converting PDF documents into Markdown format suitable for LLM prompting. This conversion relies on a two-stage pipeline that separates raw document parsing from intelligent content formatting. Understanding these libraries and their interaction is essential for anyone modifying the resume ingestion workflow or troubleshooting extraction quality.
The Two-Library Architecture for PDF Processing
Hiring Agent employs a layered approach to PDF conversion, delegating specific responsibilities to specialized libraries.
PyMuPDF (pymupdf) — Low-Level PDF Parsing
PyMuPDF serves as the foundational engine for document access. In requirements.txt, the project pins this dependency to version 1.26.3. This library handles the critical first step of opening PDF files and extracting raw page content, including text, images, vector graphics, and metadata. The PDFHandler.extract_text_from_pdf method in pdf.py initiates the process by calling pymupdf.open() to load the document into memory, creating a document object that subsequent processing layers consume.
pymupdf4llm — High-Level Markdown Generation
pymupdf4llm (version 0.0.27 as specified in requirements.txt) provides the intelligent formatting layer. Rather than returning unstructured text, this library analyzes the visual layout using PyMuPDF's underlying data structures to recognize structural elements. The core routine to_markdown, defined in pymupdf_rag.py, orchestrates this transformation. It employs helpers such as get_raw_lines for text extraction, column_boxes for layout detection, and IdentifyHeaders to compute header levels from font size variations.
How the Conversion Pipeline Works in Code
The actual execution flow spans two primary source files, demonstrating a clear separation between interface and implementation.
Entry Point: PDFHandler.extract_text_from_pdf
The PDFHandler class in pdf.py provides the high-level interface that the rest of the application calls. When processing a resume, extract_text_from_pdf opens the target document using PyMuPDF, selects the relevant pages (supporting partial document processing), and passes the document object along with page ranges to the conversion layer. This method abstracts the complexity of the underlying libraries, offering a simple API for PDF ingestion.
Core Logic: to_markdown in pymupdf_rag.py
The to_markdown function in pymupdf_rag.py executes the heavy lifting of format conversion. This function receives the PyMuPDF document object and iterates through specified pages, applying several analytical steps:
- Layout Analysis: Uses
column_boxesto detect text columns and distinguish body text from tables or graphics. - Header Detection: Leverages
IdentifyHeadersto calculate hierarchical levels based on font size deltas across the document. - Content Rendering: Calls
write_textto emit Markdown syntax, with specialized handlers foroutput_tablesand image processing.
The function returns a complete Markdown string that preserves the document's structural hierarchy, making resumes machine-readable for downstream LLM processing.
Practical Implementation Example
The following code demonstrates the exact pattern used within Hiring Agent's PDF processing workflow:
from pymupdf import open as open_pdf # PyMuPDF v1.26.3
from pymupdf_rag import to_markdown # pymupdf4llm v0.0.27
pdf_path = "candidate_resume.pdf"
# Stage 1: Open PDF with PyMuPDF (equivalent to PDFHandler.extract_text_from_pdf)
doc = open_pdf(pdf_path)
# Stage 2: Convert to Markdown (as implemented in pymupdf_rag.py)
markdown_output = to_markdown(
doc,
pages=range(doc.page_count), # Process all pages; accepts list for subsets
hdr_info=None, # Auto-detect headers from font metrics
write_images=False, # Optional: save images to filesystem
embed_images=False, # Optional: base64-encode images inline
ignore_images=False,
ignore_graphics=False,
detect_bg_color=True,
)
print(markdown_output[:1000]) # Preview first 1000 characters
This implementation mirrors the production code in pdf.py, showing how the application combines PyMuPDF's document handling with pymupdf4llm's formatting intelligence.
Key Files in the Conversion Pipeline
Three critical files define Hiring Agent's PDF processing capabilities:
pdf.py: Contains thePDFHandlerclass withextract_text_from_pdf, serving as the primary entry point for resume ingestion.pymupdf_rag.py: Houses theto_markdownfunction and all supporting logic for Markdown generation, including header detection and table formatting.requirements.txt: Pins the exact library versions (pymupdf==1.26.3andpymupdf4llm==0.0.27) ensuring reproducible PDF processing behavior across deployments.
Summary
- Hiring Agent uses two complementary libraries: PyMuPDF for raw PDF parsing and pymupdf4llm for intelligent Markdown generation.
- Version pinning matters: The project specifically requires PyMuPDF
1.26.3and pymupdf4llm0.0.27as defined inrequirements.txt. - Processing occurs across two files:
pdf.pyhandles high-level document management whilepymupdf_rag.pycontains the coreto_markdownlogic. - Pipeline preserves document structure: The conversion detects headers, tables, and columns using font metrics and layout analysis rather than simple text extraction.
Frequently Asked Questions
What version of PyMuPDF does Hiring Agent use?
Hiring Agent pins PyMuPDF to version 1.26.3 in its requirements.txt file. This specific version ensures stable behavior for the low-level PDF parsing operations that feed into the Markdown conversion pipeline.
How does Hiring Agent handle tables and images during PDF conversion?
The to_markdown function in pymupdf_rag.py includes specialized logic for both elements. It uses output_tables to render tabular data as Markdown tables and provides parameters like write_images and embed_images to control whether images are saved as external files, embedded as base64 strings, or ignored entirely based on the configuration passed from PDFHandler.
What is the difference between PyMuPDF and pymupdf4llm?
PyMuPDF is a general-purpose PDF library that provides access to document content at the structural level. pymupdf4llm is a specialized wrapper that builds on PyMuPDF to specifically convert PDF content into clean Markdown format, adding capabilities like header detection (IdentifyHeaders) and column analysis (column_boxes) that raw PyMuPDF does not provide.
Can the PDF to Markdown conversion be configured to ignore specific elements?
Yes. The to_markdown function accepts boolean flags including ignore_images and ignore_graphics that allow the caller to skip specific content types. The PDFHandler in pdf.py can pass these parameters based on application requirements, enabling flexible processing for different resume formats or downstream LLM contexts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →