How the Hiring-Agent Repository Handles Different PDF Formats and Layouts During Extraction

The Hiring-Agent repository processes diverse PDF formats and layouts through a two-stage pipeline where PDFHandler orchestrates PyMuPDF-based extraction via the to_markdown function, converting everything from multi-column résumés to scanned documents into layout-preserving Markdown before LLM structuring.

The interviewstreet/hiring-agent repository extracts structured résumé data from PDFs ranging from clean digital documents to complex scanned portfolios. Its extraction engine gracefully handles varying PDF formats and layouts through a specialized Markdown conversion layer in pymupdf_rag.py that preserves visual hierarchy before semantic parsing.

Core Extraction Architecture

PDFHandler Orchestration

Located in pdf.py (lines 47-64), the PDFHandler class serves as the primary entry point. The extract_text_from_pdf method opens documents using pymupdf.open, delegates to to_markdown for layout-aware conversion, and feeds the resulting Markdown to extract_json_from_pdf for LLM-driven JSON structuring. This separation ensures that visual layout information reaches the language model while remaining robust against malformed inputs.

Layout-Aware Markdown Conversion

The to_markdown function in pymupdf_rag.py (lines 31-38) performs the heavy lifting of analyzing font sizes, spatial relationships, content types, and reading order. It walks through PyMuPDF documents page-by-page to generate Markdown strings that faithfully reflect the original visual hierarchy, including headers, columns, tables, and graphics.

Handling Specific PDF Characteristics

Native Text and Header Detection

For vector-based text PDFs, the system extracts raw text spans using page.get_text("dict"). The IdentifyHeaders class analyzes font sizes to infer Markdown header levels (ranging from # to ######). When available, TocHeaders reads the document outline via doc.get_toc() to map entries to header tags, providing more accurate detection than font-size heuristics alone (lines 67-84).

Multi-Column Layouts

To maintain Western left-to-right reading order in multi-column documents, the pipeline employs column_boxes from pymupdf4llm.helpers.multi_column (lines 59-62). This utility splits pages into distinct columnar text blocks before extraction, ensuring that text flows correctly across complex columnar arrangements rather than extracting column-by-column in sequence.

Tables and Structured Data

The system detects tables using page.find_tables with table_strategy="lines_strict" by default (lines 91-98). Small tables containing fewer than two rows or columns are filtered out to avoid noise. Valid tables are converted to Markdown pipe syntax (|...|) and inserted at their original positions relative to surrounding text (lines 124-146).

Images and Vector Graphics

Image extraction relies on page.get_image_info with configurable image_size_limit filtering to skip tiny artifacts (lines 131-140). Large images can be saved to disk via write_images=True or embedded as Base64 via embed_images=True. For vector graphics, page.cluster_drawings groups drawing commands, and the is_significant helper (lines 56-74) distinguishes meaningful diagrams from decorative lines. Background colors are detected via get_bg_color sampling page corners (lines 162-173), allowing the system to ignore decorative fills that might otherwise clutter extraction.

Scanned and OCR PDFs

When page_is_ocr detects that all text objects are ignore-text types (lines 100-108), the pipeline recognizes the page as image-only. It relies on PyMuPDF's built-in OCR capabilities via page.get_text when force_text is enabled, ensuring textual extraction from scanned documents that lack embedded text layers.

Text spans intersecting with PDF link rectangles are transformed into Markdown hyperlink syntax ([text](url)). The _resolve_multiple_links helper intelligently splits spans containing multiple overlapping links (lines 32-66), preserving navigation integrity in the Markdown output.

Error Resilience and Logging

Throughout to_markdown, failures such as malformed tables or missing images trigger logger.debug or logger.error warnings while allowing processing to continue (lines 590-597). This ensures the pipeline returns the most complete Markdown possible rather than failing entirely when encountering individual page errors or corrupted elements.

Practical Usage Examples

Extracting Structured JSON from a PDF Résumé

from pdf import PDFHandler

handler = PDFHandler()
json_resume = handler.extract_json_from_pdf("samples/resume.pdf")

if json_resume:
    # json_resume is a Pydantic model; serialize to JSON:

    print(json_resume.json(indent=2))
else:
    print("Extraction failed.")

Key flow: PDFHandler.extract_json_from_pdf → extract_text_from_pdf → to_markdown → LLM structuring.
Source: pdf.py (lines 47-64)

Debugging with Raw Markdown Output

from pdf import PDFHandler

handler = PDFHandler()
markdown = handler.extract_text_from_pdf("samples/resume.pdf")
print(markdown[:1000])  # Preview first 1000 characters

Key method: PDFHandler.extract_text_from_pdf wraps to_markdown.
Source: pdf.py (lines 47-61)

Customizing Extraction Parameters

from pymupdf_rag import to_markdown
import pymupdf

doc = pymupdf.open("samples/complex.pdf")
md = to_markdown(
    doc,
    pages=range(doc.page_count),
    embed_images=True,      # Embed images as base64

    ignore_graphics=True,   # Skip vector graphics

    table_strategy="lines_strict"
)
print(md)

Key parameters: embed_images, ignore_graphics, table_strategy.
Source: pymupdf_rag.py (lines 31-38)

Summary

  • The Hiring-Agent repository handles different PDF formats and layouts through a two-stage pipeline: layout-aware Markdown extraction followed by LLM structuring.
  • PDFHandler in pdf.py orchestrates the high-level flow from PDF ingestion to structured JSON output.
  • to_markdown in pymupdf_rag.py manages multi-column layouts via column_boxes, tables via page.find_tables, and graphics via clustering and significance filtering.
  • OCR detection via page_is_ocr enables processing of scanned documents without embedded text layers.
  • Error resilience throughout the extraction pipeline ensures partial extractions succeed even when individual elements encounter failures.

Frequently Asked Questions

How does the system handle multi-column PDF résumés?

The to_markdown function uses column_boxes from pymupdf4llm.helpers.multi_column (lines 59-62) to identify distinct columnar regions before text extraction begins. This ensures the Markdown output follows natural Western reading order (left-to-right, top-to-bottom) across the entire page rather than extracting text column-by-column in isolation.

Can it extract text from scanned PDFs?

Yes. The page_is_ocr check (lines 100-108) detects when pages contain only image-based content by verifying whether all text objects are ignore-text types. When detected, the system invokes PyMuPDF's OCR capabilities via page.get_text with force_text enabled to extract textual content from scanned documents.

What happens to images and graphics embedded in the PDF?

Images are extracted using page.get_image_info and filtered by image_size_limit to remove tiny artifacts (lines 131-140). Large images can be saved to disk or embedded as Base64 data URIs. Vector graphics are clustered via page.cluster_drawings and assessed by the is_significant helper to determine if they should be rendered as image placeholders or filtered out as decorative elements (lines 56-74).

How are tables converted during the extraction process?

Tables are detected using page.find_tables with table_strategy="lines_strict" (lines 91-98). The system ignores small tables with fewer than two rows or columns. Valid tables are converted to Markdown pipe syntax and positioned within the text flow according to their original spatial relationships to surrounding content (lines 124-146).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →