# How Hiring Agent Extracts Data from PDF Resumes: A Complete Technical Guide

> Discover how Hiring Agent extracts data from PDF resumes. This guide details the four-stage pipeline using PyMuPDF, Jinja2, and Pydantic for structured JSON output.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Hiring Agent parses unstructured PDF resumes into structured JSON using a four-stage pipeline that combines PyMuPDF for text extraction, Jinja2 templating for LLM prompts, and Pydantic models for type-safe validation.**

The hiring-agent repository by Interview Street provides an open-source solution for automated resume parsing. This tool transforms raw PDF documents into strongly-typed JSON objects through a sophisticated pipeline that isolates PDF parsing, prompt engineering, and data validation into distinct, testable stages.

## The Four-Stage Extraction Pipeline

The extraction logic lives primarily in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) and orchestrates the transformation through four distinct phases. Each stage handles a specific concern, making the system modular and extensible.

### Stage 1: PDF to Markdown Conversion

The process begins in `PDFHandler.extract_text_from_pdf` (`pdf.py:47-61`), which utilizes **PyMuPDF** (fitz) to open the document. The method calls `to_markdown` from [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py), walking every page and converting the visual layout into plain-text markdown. This produces a single string containing the entire resume content while preserving structural cues like headers and bullet points.

### Stage 2: Prompt Generation with Jinja Templates

For each resume section—basics, work, education, skills, projects, and awards—the handler invokes the **`TemplateManager`** (`prompts/template_manager.py:21-44`). This component renders Jinja templates stored under `prompts/templates/*.jinja`, injecting the raw markdown text (`text_content`) into structured prompts. Each template contains system instructions and user prompts that direct the LLM to output JSON matching specific schema requirements.

### Stage 3: LLM-Powered Data Extraction

The rendered prompt is sent to the configured **LLM provider** (Ollama or Gemini) via `initialize_llm_provider`. The `_call_llm_for_section` method (`pdf.py:66-114`) constructs the chat payload including `model`, `messages`, and `options` using global defaults (`DEFAULT_MODEL`, `MODEL_PARAMETERS`). When a `return_model` is specified, its JSON schema is passed via the `format` keyword argument, enabling structured output generation.

The raw LLM response is processed through `extract_json_from_response` ([`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)), which strips the content to the first `{…}` block. The extracted JSON undergoes normalization via `transform_parsed_data` before validation.

### Stage 4: Pydantic Model Construction

Parsed data for each section is instantiated into corresponding **Pydantic section models** defined in `models.py:65-104`. The handler creates `BasicsSection`, `WorkSection`, `EducationSection`, and other section objects, then assembles them into a final `JSONResume` object (`pdf.py:261-279`). This guarantees type safety and validation, ensuring downstream code receives a fully-typed representation. Any error during extraction causes immediate abortion, preventing partial or corrupted resume data.

## Implementation Example

The following code demonstrates the public API for extracting structured data from PDF resumes:

```python
from pdf import PDFHandler

# Initialize the handler (loads templates and selects the LLM provider)

handler = PDFHandler()

# Path to a candidate's resume PDF

pdf_path = "candidates/alice_smith_resume.pdf"

# Extract a typed JSONResume object

resume = handler.extract_json_from_pdf(pdf_path)

if resume:
    # Access structured fields safely

    print("Name:", resume.basics.name)
    print("Primary skills:", [s.name for s in resume.skills or []])
    print("Work history entries:", len(resume.work or []))
else:
    print("Failed to parse the resume.")

```

This entry point handles the complete pipeline internally, from PDF text extraction through final model validation.

## Key Components and File Structure

Understanding the repository structure helps when extending or debugging the extraction process:

- **[`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)**: Core orchestrator containing `PDFHandler` class. Implements PDF-to-markdown conversion, section-wise LLM prompting, JSON parsing, and Pydantic model assembly.
- **[`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)**: Pydantic schema definitions for every resume section (`BasicsSection`, `WorkSection`, etc.) and the top-level `JSONResume` class.
- **[`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py)**: Loads and renders Jinja templates for each extraction section.
- **[`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)**: Provider initialization (`initialize_llm_provider`) and response cleaning utilities (`extract_json_from_response`).
- **[`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)**: Implements `to_markdown`, converting PyMuPDF documents into structured markdown text.

## Summary

- **Hiring Agent** uses a four-stage pipeline to extract data from PDF resumes: PDF parsing, prompt generation, LLM extraction, and Pydantic validation.
- **PyMuPDF** handles the initial text extraction via `extract_text_from_pdf` in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), converting visual layouts to markdown.
- **Jinja2 templates** in [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) create structured prompts for each resume section, injecting the raw markdown content.
- **LLM providers** (Ollama or Gemini) process prompts through `initialize_llm_provider`, with JSON schema enforcement via the `format` parameter.
- **Pydantic models** in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) ensure type-safe validation, with `JSONResume` serving as the root container for all extracted data.
- The **public API** `extract_json_from_pdf` encapsulates the entire workflow, returning fully-typed objects or aborting on any extraction error.

## Frequently Asked Questions

### How does Hiring Agent handle different PDF layouts and formats?

Hiring Agent relies on **PyMuPDF's** `to_markdown` function to normalize diverse PDF layouts into consistent markdown text. This preserves structural information like headers and lists while stripping formatting variations, allowing the LLM to process standardized text regardless of the original PDF's visual design.

### Which LLM providers does Hiring Agent support for resume extraction?

The system supports **Ollama** and **Gemini** through the `initialize_llm_provider` utility in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py). The provider's `chat` method handles the actual API calls, with configuration managed via global defaults including `DEFAULT_MODEL` and `MODEL_PARAMETERS` parameters.

### What happens if the LLM returns malformed JSON or invalid data?

The pipeline includes multiple safeguards. The `extract_json_from_response` function isolates the first JSON block from the LLM output, while `transform_parsed_data` normalizes the structure. Finally, Pydantic validation in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) enforces type constraints. If any stage fails, the entire extraction aborts to prevent partial data corruption, as implemented in `pdf.py:261-279`.

### Can I extract specific sections without processing the entire resume?

Yes. The `PDFHandler` class provides individual extraction methods such as `extract_basics_section`, `extract_work_section`, and `extract_skills_section`. These methods follow the same pipeline but target specific sections, allowing modular extraction when you only need particular data points like work history or skills.