# How the Hiring-Agent Repository Handles Different PDF Formats and Layouts During Extraction

> Learn how the Hiring-Agent repository handles diverse PDF formats and layouts. Discover its two-stage pipeline for extracting information with PyMuPDF and LLM structuring.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-05

---

**The Hiring-Agent repository processes diverse PDF formats and layouts through a two-stage pipeline where `PDFHandler` orchestrates PyMuPDF-based extraction via the `to_markdown` function, converting everything from multi-column résumés to scanned documents into layout-preserving Markdown before LLM structuring.**

The `interviewstreet/hiring-agent` repository extracts structured résumé data from PDFs ranging from clean digital documents to complex scanned portfolios. Its extraction engine gracefully handles varying PDF formats and layouts through a specialized Markdown conversion layer in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) that preserves visual hierarchy before semantic parsing.

## Core Extraction Architecture

### PDFHandler Orchestration

Located in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 47-64), the `PDFHandler` class serves as the primary entry point. The `extract_text_from_pdf` method opens documents using `pymupdf.open`, delegates to `to_markdown` for layout-aware conversion, and feeds the resulting Markdown to `extract_json_from_pdf` for LLM-driven JSON structuring. This separation ensures that visual layout information reaches the language model while remaining robust against malformed inputs.

### Layout-Aware Markdown Conversion

The `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 31-38) performs the heavy lifting of analyzing font sizes, spatial relationships, content types, and reading order. It walks through PyMuPDF documents page-by-page to generate Markdown strings that faithfully reflect the original visual hierarchy, including headers, columns, tables, and graphics.

## Handling Specific PDF Characteristics

### Native Text and Header Detection

For vector-based text PDFs, the system extracts raw text spans using `page.get_text("dict")`. The `IdentifyHeaders` class analyzes font sizes to infer Markdown header levels (ranging from `#` to `######`). When available, `TocHeaders` reads the document outline via `doc.get_toc()` to map entries to header tags, providing more accurate detection than font-size heuristics alone (lines 67-84).

### Multi-Column Layouts

To maintain Western left-to-right reading order in multi-column documents, the pipeline employs `column_boxes` from `pymupdf4llm.helpers.multi_column` (lines 59-62). This utility splits pages into distinct columnar text blocks before extraction, ensuring that text flows correctly across complex columnar arrangements rather than extracting column-by-column in sequence.

### Tables and Structured Data

The system detects tables using `page.find_tables` with `table_strategy="lines_strict"` by default (lines 91-98). Small tables containing fewer than two rows or columns are filtered out to avoid noise. Valid tables are converted to Markdown pipe syntax (`|...|`) and inserted at their original positions relative to surrounding text (lines 124-146).

### Images and Vector Graphics

Image extraction relies on `page.get_image_info` with configurable `image_size_limit` filtering to skip tiny artifacts (lines 131-140). Large images can be saved to disk via `write_images=True` or embedded as Base64 via `embed_images=True`. For vector graphics, `page.cluster_drawings` groups drawing commands, and the `is_significant` helper (lines 56-74) distinguishes meaningful diagrams from decorative lines. Background colors are detected via `get_bg_color` sampling page corners (lines 162-173), allowing the system to ignore decorative fills that might otherwise clutter extraction.

### Scanned and OCR PDFs

When `page_is_ocr` detects that all text objects are ignore-text types (lines 100-108), the pipeline recognizes the page as image-only. It relies on PyMuPDF's built-in OCR capabilities via `page.get_text` when `force_text` is enabled, ensuring textual extraction from scanned documents that lack embedded text layers.

### Hyperlinks and Navigation

Text spans intersecting with PDF link rectangles are transformed into Markdown hyperlink syntax (`[text](url)`). The `_resolve_multiple_links` helper intelligently splits spans containing multiple overlapping links (lines 32-66), preserving navigation integrity in the Markdown output.

## Error Resilience and Logging

Throughout `to_markdown`, failures such as malformed tables or missing images trigger `logger.debug` or `logger.error` warnings while allowing processing to continue (lines 590-597). This ensures the pipeline returns the most complete Markdown possible rather than failing entirely when encountering individual page errors or corrupted elements.

## Practical Usage Examples

### Extracting Structured JSON from a PDF Résumé

```python
from pdf import PDFHandler

handler = PDFHandler()
json_resume = handler.extract_json_from_pdf("samples/resume.pdf")

if json_resume:
    # json_resume is a Pydantic model; serialize to JSON:

    print(json_resume.json(indent=2))
else:
    print("Extraction failed.")

```

*Key flow*: `PDFHandler.extract_json_from_pdf` → `extract_text_from_pdf` → `to_markdown` → LLM structuring.  
Source: [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 47-64)

### Debugging with Raw Markdown Output

```python
from pdf import PDFHandler

handler = PDFHandler()
markdown = handler.extract_text_from_pdf("samples/resume.pdf")
print(markdown[:1000])  # Preview first 1000 characters

```

*Key method*: `PDFHandler.extract_text_from_pdf` wraps `to_markdown`.  
Source: [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 47-61)

### Customizing Extraction Parameters

```python
from pymupdf_rag import to_markdown
import pymupdf

doc = pymupdf.open("samples/complex.pdf")
md = to_markdown(
    doc,
    pages=range(doc.page_count),
    embed_images=True,      # Embed images as base64

    ignore_graphics=True,   # Skip vector graphics

    table_strategy="lines_strict"
)
print(md)

```

*Key parameters*: `embed_images`, `ignore_graphics`, `table_strategy`.  
Source: [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 31-38)

## Summary

- The Hiring-Agent repository handles different PDF formats and layouts through a two-stage pipeline: layout-aware Markdown extraction followed by LLM structuring.
- `PDFHandler` in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) orchestrates the high-level flow from PDF ingestion to structured JSON output.
- `to_markdown` in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) manages multi-column layouts via `column_boxes`, tables via `page.find_tables`, and graphics via clustering and significance filtering.
- OCR detection via `page_is_ocr` enables processing of scanned documents without embedded text layers.
- Error resilience throughout the extraction pipeline ensures partial extractions succeed even when individual elements encounter failures.

## Frequently Asked Questions

### How does the system handle multi-column PDF résumés?

The `to_markdown` function uses `column_boxes` from `pymupdf4llm.helpers.multi_column` (lines 59-62) to identify distinct columnar regions before text extraction begins. This ensures the Markdown output follows natural Western reading order (left-to-right, top-to-bottom) across the entire page rather than extracting text column-by-column in isolation.

### Can it extract text from scanned PDFs?

Yes. The `page_is_ocr` check (lines 100-108) detects when pages contain only image-based content by verifying whether all text objects are ignore-text types. When detected, the system invokes PyMuPDF's OCR capabilities via `page.get_text` with `force_text` enabled to extract textual content from scanned documents.

### What happens to images and graphics embedded in the PDF?

Images are extracted using `page.get_image_info` and filtered by `image_size_limit` to remove tiny artifacts (lines 131-140). Large images can be saved to disk or embedded as Base64 data URIs. Vector graphics are clustered via `page.cluster_drawings` and assessed by the `is_significant` helper to determine if they should be rendered as image placeholders or filtered out as decorative elements (lines 56-74).

### How are tables converted during the extraction process?

Tables are detected using `page.find_tables` with `table_strategy="lines_strict"` (lines 91-98). The system ignores small tables with fewer than two rows or columns. Valid tables are converted to Markdown pipe syntax and positioned within the text flow according to their original spatial relationships to surrounding content (lines 124-146).