What is PyMuPDF Used for in the Hiring Agent? PDF Parsing and Résumé Extraction Explained

PyMuPDF serves as the core PDF parsing engine in the Hiring Agent, converting raw résumé PDFs into structured Markdown that feeds into LLM-based extraction pipelines.

The interviewstreet/hiring-agent repository is a résumé-parsing tool that transforms PDF documents into structured JSON objects. At the heart of this system lies PyMuPDF (imported as pymupdf), which handles all low-level PDF processing operations according to the source code. The library bridges the gap between binary PDF files and the LLM-based extraction pipeline by converting document content into clean, normalized Markdown format.

Core PDF Processing Operations

The Hiring Agent relies on PyMuPDF for fundamental document operations that power the entire extraction workflow, from opening files to extracting text in reading order.

Opening PDF Documents

In pdf.py (lines 52-57), PyMuPDF initializes document objects through pymupdf.open(pdf_path). This creates a Document object representing the entire PDF file, which serves as the entry point for all subsequent parsing operations.


# From pdf.py:52-57 - Document initialization

doc = pymupdf.open(pdf_path)

Extracting Text in Reading Order

The to_markdown function in pymupdf_rag.py (lines 33-41) orchestrates text extraction by walking through each page, collecting text blocks, and identifying structural elements like headings, tables, and code sections. This helper returns a normalized Markdown string that preserves document semantics while enabling downstream LLM processing.


# Conceptual flow from pymupdf_rag.py:33-41

markdown_content = to_markdown(doc)

Advanced Document Analysis Features

Beyond basic text extraction, PyMuPDF enables sophisticated document analysis capabilities critical for parsing complex résumé layouts.

Table Detection and Conversion

PyMuPDF identifies tabular data through page.find_tables() as implemented in pymupdf_rag.py (lines 191-196). The Hiring Agent converts detected tables into Markdown table format, preserving structured data found in candidate résumés such as skills matrices and employment timelines.

Image and Graphics Handling

The system detects embedded media using page.get_image_info() and page.get_links() (found in pymupdf_rag.py, lines 111-118). The save_image helper optionally writes images to disk or embeds them as Base-64 data URIs, ensuring visual content from portfolios or certifications remains accessible to the extraction pipeline.

Header Detection Strategies

Two distinct header identification strategies operate within pymupdf_rag.py (lines 70-84):

  • IdentifyHeaders: Analyzes font sizes to detect hierarchical headings
  • TocHeaders: Leverages the PDF's Table of Contents to generate Markdown header markers (#)

OCR and Special Case Handling

PyMuPDF handles scanned documents through OCR detection logic in pymupdf_rag.py (lines 1000-1009). The page_is_ocr function checks for "ignore-text" glyphs generated by OCR engines, determining whether to treat a page as text-only or image-based content that requires alternative processing.

Integration with the LLM Pipeline

The PDFHandler.extract_text_from_pdf method in pdf.py (lines 47-58) serves as the integration point between PyMuPDF and the LLM. This method:

  1. Opens the document using pymupdf.open()
  2. Invokes to_markdown to convert content to Markdown
  3. Returns raw Markdown that feeds LLM-based section extractors defined in pdf.py
from pdf import PDFHandler

# Initialize the handler

handler = PDFHandler()

# Convert PDF to markdown (PyMuPDF does the heavy lifting)

markdown = handler.extract_text_from_pdf("resume.pdf")

# Full pipeline: PDF → markdown → LLM → JSONResume

json_resume = handler.extract_json_from_pdf("resume.pdf")

The dependency is declared in requirements.txt as pymupdf, cementing PyMuPDF as the backbone of the Hiring Agent's PDF handling capabilities.

Summary

  • PyMuPDF provides the foundational PDF parsing layer for the interviewstreet/hiring-agent repository according to the source code
  • The pymupdf.open() function in pdf.py:52-57 initializes document processing
  • Text extraction occurs through to_markdown in pymupdf_rag.py:33-41, which handles reading order, headers, and code blocks
  • Table detection uses page.find_tables() in pymupdf_rag.py:191-196 to preserve structured data
  • Image handling via page.get_image_info() and save_image in pymupdf_rag.py:111-118 processes visual content
  • OCR detection through page_is_ocr in pymupdf_rag.py:1000-1009 handles scanned documents
  • The PDFHandler class in pdf.py:47-58 bridges PyMuPDF output with LLM-based JSON extraction

Frequently Asked Questions

How does PyMuPDF handle scanned PDFs in the Hiring Agent?

PyMuPDF detects OCR-generated content through the page_is_ocr function in pymupdf_rag.py (lines 1000-1009). This helper checks for "ignore-text" glyphs that indicate OCR processing, allowing the system to determine whether to extract text directly or process the page as an image-based document requiring special handling.

What is the difference between pdf.py and pymupdf_rag.py in the Hiring Agent?

pdf.py contains the high-level PDFHandler class that orchestrates the extraction pipeline and integrates with the LLM. pymupdf_rag.py implements the core PyMuPDF logic, including the to_markdown function, header detection strategies (IdentifyHeaders and TocHeaders), table extraction, and image handling. The former manages workflow while the latter performs the actual PDF parsing operations.

Can PyMuPDF extract tables from résumés in the Hiring Agent?

Yes, PyMuPDF identifies tables through page.find_tables() as implemented in pymupdf_rag.py (lines 191-196). The Hiring Agent converts these tables into Markdown format, preserving the structured layout of skills matrices, employment histories, and other tabular data found in candidate documents.

What format does PyMuPDF output in the Hiring Agent?

PyMuPDF outputs extracted content as Markdown through the to_markdown helper function in pymupdf_rag.py. This normalized format preserves document structure—including headers, tables, and code blocks—making it suitable for downstream LLM processing that ultimately generates structured JSON résumé objects.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →