How PDF Text Extraction Using PyMuPDF Works in the Hiring-Agent Repository
The hiring-agent repository implements PDF text extraction using PyMuPDF by combining a high-level PDFHandler class with a specialized to_markdown conversion function that transforms document pages into structured Markdown while automatically detecting headers, tables, and graphics.
The interviewstreet/hiring-agent project processes resume documents through a dedicated pipeline for PDF text extraction using PyMuPDF (the Python binding for MuPDF). This system converts binary PDF files into clean Markdown format, enabling downstream LLM-based parsing of candidate information. The architecture separates file handling concerns from document layout analysis, utilizing PyMuPDF's fitz interface for low-level page access.
Core Architecture Components
The extraction system relies on two primary components that work sequentially to process documents.
PDFHandler High-Level API
The PDFHandler class in pdf.py provides the main entry point for document processing. Its extract_text_from_pdf method (lines 47-58) orchestrates the opening of PDF files and delegates conversion to the markdown generator. This wrapper handles file I/O context management and page range selection before invoking the core transformation logic.
to_markdown Conversion Engine
The to_markdown function in pymupdf_rag.py (lines 302-336) performs the heavy lifting of content extraction. It creates a Parameters dataclass to store page-level metadata, iterates through specified pages using page.get_textpage(), and applies layout detection algorithms including IdentifyHeaders for heading hierarchy and page.find_tables() for tabular data recognition.
The PDF Text Extraction Pipeline
The process follows a four-stage pipeline from file ingestion to structured text output.
1. Document Opening and Validation
PyMuPDF opens the file stream through its pymupdf.open() context manager. As implemented in pdf.py (lines 52-53), this creates a Document object that provides access to page count, metadata, and content streams while ensuring proper resource cleanup.
with pymupdf.open(pdf_path) as doc:
pages = range(doc.page_count)
2. Page Selection and Preparation
The handler extracts all pages by default using range(doc.page_count), though the to_markdown signature accepts a customizable pages parameter for selective extraction (lines 53-56 in pdf.py). This allows processing of specific page ranges without loading unnecessary content into memory.
3. Content Processing and Layout Detection
The to_markdown function processes each page through several specialized subsystems:
- Text Block Extraction: Uses
get_raw_linesandcolumn_boxeshelpers to extract raw text while preserving reading order - Header Detection: The
IdentifyHeadersclass analyzes font size hierarchies to automatically assign Markdown heading levels (H1-H6) - Table Recognition: Calls
page.find_tables()to detect tabular structures and render them as Markdown tables - Image Handling: Processes graphics either by saving to disk (
write_images=True), embedding as Base64 (embed_images=True), or skipping entirely (ignore_images=True)
4. Markdown Assembly and Return
The function aggregates processed blocks into a single Markdown string, which extract_text_from_pdf returns to callers for ingestion by the LLM resume parser defined in subsequent pipeline stages.
Implementation Examples
Using the PDFHandler Wrapper
For standard use cases, instantiate PDFHandler to leverage automatic resource management and logging.
from pdf import PDFHandler
handler = PDFHandler()
markdown_text = handler.extract_text_from_pdf("samples/resume.pdf")
if markdown_text:
print("Extracted Markdown (first 200 chars):")
print(markdown_text[:200])
This approach handles the pymupdf.open() context, page enumeration, and error handling automatically.
Direct PyMuPDF Integration
For fine-grained control over page ranges or conversion parameters, invoke to_markdown directly.
import pymupdf
from pymupdf_rag import to_markdown
doc = pymupdf.open("samples/resume.pdf")
# Extract only pages 0-4 (first five pages)
markdown = to_markdown(doc, pages=range(5))
print(markdown)
Optimizing for Text-Only Extraction
When processing speed takes priority over image content, disable graphic extraction to reduce I/O overhead.
markdown = to_markdown(
doc,
pages=None, # all pages
write_images=False,
embed_images=False,
ignore_images=True, # skip image extraction
force_text=True, # still try OCR on image-only pages
)
Setting ignore_images=True bypasses the write_images and embed_images processing paths, significantly accelerating throughput for text-heavy documents.
Key Source Files and Responsibilities
| File | Primary Concern |
|---|---|
pdf.py |
PDFHandler class and high-level orchestration of LLM extraction |
pymupdf_rag.py |
Core conversion logic including to_markdown, header detection, and table processing |
models.py |
Pydantic models receiving parsed JSON after text extraction |
transform.py |
Helper utilities converting LLM JSON output to JSONResume objects |
llm_utils.py |
Downstream LLM interaction utilities consuming the extracted Markdown |
These modules collectively form the complete processing chain: PDF → PyMuPDF → Markdown → LLM → Structured Resume JSON.
Summary
- The
PDFHandler.extract_text_from_pdfmethod inpdf.pyserves as the primary interface for PDF text extraction using PyMuPDF, handling file opening and page selection. - The
to_markdownfunction inpymupdf_rag.pyperforms content conversion usingpage.get_textpage(),IdentifyHeaders, andpage.find_tables(). - Layout detection automatically distinguishes between headers, paragraphs, tables, and graphics without manual template configuration.
- ** Performance optimization** is achievable through parameters like
ignore_images=Trueand selective page ranging. - The extracted Markdown feeds directly into the repository's LLM pipeline defined in
models.pyandllm_utils.py.
Frequently Asked Questions
How does the system handle PDF files with complex multi-column layouts?
The to_markdown function utilizes column_boxes and get_raw_lines helpers to analyze page geometry and preserve proper reading order across columns. PyMuPDF's TextPage extraction captures positional data that the conversion logic uses to reconstruct linear text flow, ensuring that content reads correctly even when source documents use sophisticated print-style layouts.
What is the difference between the PDFHandler class and the to_markdown function?
PDFHandler (in pdf.py) provides a high-level abstraction that manages file contexts, page enumeration, and error handling, making it ideal for standard batch processing. to_markdown (in pymupdf_rag.py) is the low-level conversion engine that accepts an already-opened PyMuPDF Document object and exposes granular parameters for image handling, page selection, and layout detection, suited for custom integration scenarios.
Can I extract text from specific pages rather than the entire document?
Yes. While PDFHandler defaults to range(doc.page_count), you can pass a custom pages parameter directly to to_markdown. Supply any iterable of page indices (e.g., range(5) for the first five pages or [0, 2, 4] for non-sequential pages) to process only specific sections of large PDF files.
How are images and graphics processed during text extraction?
The system supports three image handling modes controlled via boolean flags: write_images=True saves extracted images to disk with file references in the Markdown; embed_images=True encodes graphics as Base64 data URIs within the document; and ignore_images=True skips image extraction entirely for faster text-only processing. When force_text=True is set alongside image flags, the system attempts OCR on image-only pages using PyMuPDF's text extraction capabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →