What Libraries Are Used for PDF to Markdown Conversion in Hiring Agent?

Hiring Agent uses PyMuPDF (pymupdf) version 1.26.3 for low-level PDF parsing and pymupdf4llm version 0.0.27 for high-level Markdown generation, working together in pdf.py and pymupdf_rag.py to convert resume PDFs into clean, structured Markdown.

The interviewstreet/hiring-agent repository processes candidate resumes by converting PDF documents into Markdown format suitable for LLM prompting. This conversion relies on a two-stage pipeline that separates raw document parsing from intelligent content formatting. Understanding these libraries and their interaction is essential for anyone modifying the resume ingestion workflow or troubleshooting extraction quality.

The Two-Library Architecture for PDF Processing

Hiring Agent employs a layered approach to PDF conversion, delegating specific responsibilities to specialized libraries.

PyMuPDF (pymupdf) — Low-Level PDF Parsing

PyMuPDF serves as the foundational engine for document access. In requirements.txt, the project pins this dependency to version 1.26.3. This library handles the critical first step of opening PDF files and extracting raw page content, including text, images, vector graphics, and metadata. The PDFHandler.extract_text_from_pdf method in pdf.py initiates the process by calling pymupdf.open() to load the document into memory, creating a document object that subsequent processing layers consume.

pymupdf4llm — High-Level Markdown Generation

pymupdf4llm (version 0.0.27 as specified in requirements.txt) provides the intelligent formatting layer. Rather than returning unstructured text, this library analyzes the visual layout using PyMuPDF's underlying data structures to recognize structural elements. The core routine to_markdown, defined in pymupdf_rag.py, orchestrates this transformation. It employs helpers such as get_raw_lines for text extraction, column_boxes for layout detection, and IdentifyHeaders to compute header levels from font size variations.

How the Conversion Pipeline Works in Code

The actual execution flow spans two primary source files, demonstrating a clear separation between interface and implementation.

Entry Point: PDFHandler.extract_text_from_pdf

The PDFHandler class in pdf.py provides the high-level interface that the rest of the application calls. When processing a resume, extract_text_from_pdf opens the target document using PyMuPDF, selects the relevant pages (supporting partial document processing), and passes the document object along with page ranges to the conversion layer. This method abstracts the complexity of the underlying libraries, offering a simple API for PDF ingestion.

Core Logic: to_markdown in pymupdf_rag.py

The to_markdown function in pymupdf_rag.py executes the heavy lifting of format conversion. This function receives the PyMuPDF document object and iterates through specified pages, applying several analytical steps:

  • Layout Analysis: Uses column_boxes to detect text columns and distinguish body text from tables or graphics.
  • Header Detection: Leverages IdentifyHeaders to calculate hierarchical levels based on font size deltas across the document.
  • Content Rendering: Calls write_text to emit Markdown syntax, with specialized handlers for output_tables and image processing.

The function returns a complete Markdown string that preserves the document's structural hierarchy, making resumes machine-readable for downstream LLM processing.

Practical Implementation Example

The following code demonstrates the exact pattern used within Hiring Agent's PDF processing workflow:

from pymupdf import open as open_pdf          # PyMuPDF v1.26.3

from pymupdf_rag import to_markdown            # pymupdf4llm v0.0.27

pdf_path = "candidate_resume.pdf"

# Stage 1: Open PDF with PyMuPDF (equivalent to PDFHandler.extract_text_from_pdf)

doc = open_pdf(pdf_path)

# Stage 2: Convert to Markdown (as implemented in pymupdf_rag.py)

markdown_output = to_markdown(
    doc,
    pages=range(doc.page_count),   # Process all pages; accepts list for subsets

    hdr_info=None,                 # Auto-detect headers from font metrics

    write_images=False,            # Optional: save images to filesystem

    embed_images=False,            # Optional: base64-encode images inline

    ignore_images=False,
    ignore_graphics=False,
    detect_bg_color=True,
)

print(markdown_output[:1000])      # Preview first 1000 characters

This implementation mirrors the production code in pdf.py, showing how the application combines PyMuPDF's document handling with pymupdf4llm's formatting intelligence.

Key Files in the Conversion Pipeline

Three critical files define Hiring Agent's PDF processing capabilities:

  • pdf.py: Contains the PDFHandler class with extract_text_from_pdf, serving as the primary entry point for resume ingestion.
  • pymupdf_rag.py: Houses the to_markdown function and all supporting logic for Markdown generation, including header detection and table formatting.
  • requirements.txt: Pins the exact library versions (pymupdf==1.26.3 and pymupdf4llm==0.0.27) ensuring reproducible PDF processing behavior across deployments.

Summary

  • Hiring Agent uses two complementary libraries: PyMuPDF for raw PDF parsing and pymupdf4llm for intelligent Markdown generation.
  • Version pinning matters: The project specifically requires PyMuPDF 1.26.3 and pymupdf4llm 0.0.27 as defined in requirements.txt.
  • Processing occurs across two files: pdf.py handles high-level document management while pymupdf_rag.py contains the core to_markdown logic.
  • Pipeline preserves document structure: The conversion detects headers, tables, and columns using font metrics and layout analysis rather than simple text extraction.

Frequently Asked Questions

What version of PyMuPDF does Hiring Agent use?

Hiring Agent pins PyMuPDF to version 1.26.3 in its requirements.txt file. This specific version ensures stable behavior for the low-level PDF parsing operations that feed into the Markdown conversion pipeline.

How does Hiring Agent handle tables and images during PDF conversion?

The to_markdown function in pymupdf_rag.py includes specialized logic for both elements. It uses output_tables to render tabular data as Markdown tables and provides parameters like write_images and embed_images to control whether images are saved as external files, embedded as base64 strings, or ignored entirely based on the configuration passed from PDFHandler.

What is the difference between PyMuPDF and pymupdf4llm?

PyMuPDF is a general-purpose PDF library that provides access to document content at the structural level. pymupdf4llm is a specialized wrapper that builds on PyMuPDF to specifically convert PDF content into clean Markdown format, adding capabilities like header detection (IdentifyHeaders) and column analysis (column_boxes) that raw PyMuPDF does not provide.

Can the PDF to Markdown conversion be configured to ignore specific elements?

Yes. The to_markdown function accepts boolean flags including ignore_images and ignore_graphics that allow the caller to skip specific content types. The PDFHandler in pdf.py can pass these parameters based on application requirements, enabling flexible processing for different resume formats or downstream LLM contexts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →