# How the Hiring Agent Extracts Data from PDFs: A Technical Deep Dive

> Discover how the Hiring Agent extracts data from PDFs. Learn about the technical pipeline using PyMuPDF, LLMs, and Pydantic to convert résumés to structured JSON.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-06-28

---

**The Hiring Agent converts PDF résumés into structured JSON using a pipeline that combines PyMuPDF for text extraction, LLM-powered section parsing, and Pydantic model validation.**

The **interviewstreet/hiring-agent** repository provides a robust solution to **extract data from PDFs** and transform unstructured résumés into machine-readable formats. The system implements a multi-stage pipeline in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) that processes documents through specialized handlers and AI-driven parsing to produce type-safe JSON output.

## The PDF-to-JSON Processing Pipeline

### Step 1: PDF Loading and Text Extraction

The `PDFHandler.extract_text_from_pdf` method opens files using **PyMuPDF** and iterates over all pages. According to the source code in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 52-58), each page converts to Markdown via `pymupdf_rag.to_markdown`, producing clean plain text for downstream processing.

### Step 2: LLM-Driven Section Parsing

Once raw text is available, `_extract_all_sections_separately` (lines 71-73) orchestrates extraction of six major résumé sections: **basics**, **work**, **education**, **skills**, **projects**, and **awards**. For each section:

- **TemplateManager** renders a system-specific prompt (lines 79-82)
- The LLM provider processes the prompt via `self.provider.chat` (lines 88-94)
- JSON extraction occurs through `extract_json_from_response`, followed by field normalization via `transform_parsed_data` (lines 110-119)

### Step 3: Model-Based Data Conversion

After section collection, the handler instantiates Pydantic models defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) (lines 101-112). Successful conversions assemble into a single `JSONResume` instance, ensuring type-safe, validated output before returning to the caller.

## High-Level API and Usage

The public method `PDFHandler.extract_json_from_pdf` serves as the primary entry point (lines 199-212). It orchestrates text extraction, section parsing, and model validation into a single call.

```python
from pdf import PDFHandler

# Initialize the handler with LLM provider and templates

handler = PDFHandler()

# Extract structured data from a PDF résumé

resume_json = handler.extract_json_from_pdf("samples/jane_doe_resume.pdf")

if resume_json:
    print("Name:", resume_json.basics.name)
    print("Experience entries:", len(resume_json.work))
    print("Skills:", resume_json.skills)

```

The [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) script demonstrates production usage (lines 251-254), creating a `PDFHandler` instance and processing files through the same pipeline.

## Key Components and File Structure

- **[`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)**: Contains the `PDFHandler` class implementing the core extraction logic
- **[`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)**: Defines Pydantic schemas (`JSONResume`, `Basics`, `Work`) for data validation
- **[`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py)**: Manages system-message templates for each résumé section
- **[`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)**: Helper module converting PyMuPDF pages to Markdown
- **[`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py)**: Reference implementation showing end-to-end workflow

## Summary

- **PyMuPDF integration**: The `extract_text_from_pdf` method uses PyMuPDF and `pymupdf_rag.to_markdown` to convert PDF pages to Markdown text.
- **LLM-powered parsing**: The `_extract_all_sections_separately` method extracts six specific résumé sections using templated prompts and LLM inference.
- **Type-safe output**: Extracted data validates against Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), producing a structured `JSONResume` object.
- **Single-entry API**: The `extract_json_from_pdf` method combines all steps for easy integration, as demonstrated in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).

## Frequently Asked Questions

### How does the Hiring Agent handle different PDF formats?

The system uses **PyMuPDF** (fitz) to open documents and converts each page to Markdown using `pymupdf_rag.to_markdown`. This approach standardizes formatting variations into consistent plain text before LLM processing, handling diverse PDF structures including multi-column layouts and varied fonts.

### What résumé sections can the Hiring Agent extract?

The `_extract_all_sections_separately` method targets six specific sections: **basics** (contact info), **work** (experience), **education**, **skills**, **projects**, and **awards**. Each section uses dedicated prompts via `TemplateManager` to optimize extraction accuracy for that specific data type.

### How is the extracted data validated?

After LLM extraction, the system attempts to instantiate Pydantic models defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) (such as `Basics`, `Work`, and `JSONResume`). This model-based conversion ensures type safety, validates field formats, and guarantees the output conforms to the expected JSON schema before returning the final object.

### Can I use the PDF extraction without the scoring functionality?

Yes. The `PDFHandler` class in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) operates independently. You can import it directly, instantiate it, and call `extract_json_from_pdf` to receive a structured `JSONResume` object without invoking any scoring logic from [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).