# How the Hiring Agent Extracts Structured Data from PDF Resumes: A Four-Stage Pipeline

> Learn how the Hiring Agent extracts structured data from PDF resumes through a four stage pipeline. Convert unstructured PDFs to validated JSON using PyMuPDF, Jinja, LLMs, and Pydantic.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-06-29

---

**The Hiring Agent converts unstructured PDF résumés into validated JSON objects using a four-stage pipeline that extracts markdown text with PyMuPDF, generates section-specific prompts via Jinja templates, processes them through an LLM provider (Ollama or Gemini), and validates the output against Pydantic models.**

The **interviewstreet/hiring-agent** repository implements a robust extraction system that transforms raw PDF documents into strongly-typed data structures. This pipeline leverages **PyMuPDF** for document parsing, **Jinja2** for prompt templating, and **Pydantic** for data validation to ensure reliable structured data extraction from PDF resumes.

## Stage 1: Converting PDF Visual Layout to Markdown Text

The extraction process begins in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) with the `PDFHandler.extract_text_from_pdf` method (lines 47–61). This function opens the PDF document using `pymupdf.open(pdf_path)` and converts the visual layout into plain-text markdown.

The conversion relies on the `to_markdown` function implemented in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py), which walks every page of the document and builds a markdown string preserving the document structure. The method processes all pages via `to_markdown(doc, pages=range(doc.page_count))` and returns a single `text_content` string containing the entire résumé text.

## Stage 2: Generating Section-Specific LLM Prompts

Once the markdown text is extracted, the system prepares structured prompts for each résumé section. The `TemplateManager` class in [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) (lines 21–44) renders Jinja templates stored under `prompts/templates/*.jinja`.

The handler iterates over a fixed list of sections—**basics**, **work**, **education**, **skills**, **projects**, and **awards**—calling `TemplateManager.render_template("<section>", text_content=resume_text)` for each. Each template injects the raw markdown content into a system message that instructs the LLM to output a specific JSON structure matching the target Pydantic model for that section.

## Stage 3: LLM Extraction and Response Normalization

The rendered prompts are sent to the selected LLM provider via the `initialize_llm_provider` utility (from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)). The provider's `chat` method handles the actual inference, supporting both **Ollama** and **Gemini** backends.

In [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 66–114), the `_call_llm_for_section` method constructs the chat payload using `DEFAULT_MODEL` and `MODEL_PARAMETERS` from the global configuration. If the caller supplies a `return_model`, its JSON schema is passed to the LLM via the `format` keyword argument, enabling structured output generation.

The raw LLM response is processed through `extract_json_from_response` in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py), which strips the content to the first JSON block. The extracted JSON is parsed with `json.loads` and normalized using `transform_parsed_data` to ensure consistent field mapping before model construction.

## Stage 4: Pydantic Validation and Model Construction

The final stage transforms the parsed dictionaries into type-safe objects. In [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) (lines 261–279), each section's data is wrapped in its corresponding Pydantic model—`BasicsSection`, `WorkSection`, `EducationSection`, `SkillsSection`, `ProjectsSection`, or `AwardsSection`—as defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) (lines 65–104).

These section objects are assembled into a root `JSONResume` model, creating a fully-typed, validated representation. The design ensures that any validation errors at this stage cause the entire extraction to abort, preventing partially-filled or corrupted résumé data from propagating downstream.

## Complete Extraction Workflow

The `PDFHandler` orchestrates the full pipeline through the following execution flow:

1. **File validation** – `extract_text_from_pdf` raises `FileNotFoundError` if the path is invalid.
2. **Text extraction** – `pymupdf.open()` creates a document object; `to_markdown()` processes all pages into markdown.
3. **Section iteration** – `_extract_all_sections_separately` loops through the fixed section list, delegating to `_extract_section_data` and specific `extract_*_section` methods.
4. **Prompt rendering** – Each section calls `TemplateManager.render_template()` to generate the final LLM prompt.
5. **LLM invocation** – `_call_llm_for_section` executes the chat completion with optional JSON schema constraints.
6. **Response parsing** – The raw message is stripped to the first `{…}` block, parsed, and normalized via `transform_parsed_data`.
7. **Object instantiation** – Section dictionaries are fed to Pydantic classes (`Basics(**…)`, etc.), culminating in `JSONResume(**complete_resume)`.

## Code Example: Extracting a Structured Resume

```python
from pdf import PDFHandler

# Initialise the handler (loads templates and selects the LLM provider)

handler = PDFHandler()

# Path to a candidate's résumé PDF

pdf_path = "candidates/alice_smith_resume.pdf"

# Extract a typed JSONResume object

resume = handler.extract_json_from_pdf(pdf_path)

if resume:
    # Access structured fields safely

    print("Name:", resume.basics.name)
    print("Primary skills:", [s.name for s in resume.skills or []])
    print("Work history entries:", len(resume.work or []))
else:
    print("Failed to parse the résumé.")

```

This snippet demonstrates the public API (`PDFHandler.extract_json_from_pdf`) that downstream components use to obtain fully-validated résumé data.

## Summary

- **PyMuPDF** handles the initial PDF-to-markdown conversion in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), preserving document layout while creating machine-readable text.
- **Jinja templates** in [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) generate section-specific prompts that guide the LLM to output structured JSON.
- **LLM providers** (Ollama/Gemini) process prompts through `initialize_llm_provider`, with optional JSON schema enforcement via the `format` parameter.
- **Pydantic models** in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) validate the extracted data, assembling sections into a type-safe `JSONResume` object.
- The pipeline aborts on any extraction error, ensuring data integrity and preventing partial results.

## Frequently Asked Questions

### What PDF library does the Hiring Agent use to extract text?

The system uses **PyMuPDF** (`pymupdf`) to open PDF documents and extract content. Specifically, the `to_markdown` function in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) converts the visual layout into markdown text, preserving structural information while creating a format suitable for LLM processing.

### How does the system ensure consistent JSON output from the LLM?

The `TemplateManager` renders Jinja templates containing strict system instructions that specify the required JSON schema. Additionally, when calling the LLM via `_call_llm_for_section`, the code passes the Pydantic model's JSON schema through the `format` keyword argument, which supported providers (Ollama and Gemini) can use to constrain their output to valid JSON structures.

### Can I extend the pipeline to extract custom resume sections?

Yes. The modular architecture in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) allows extension by adding new section names to the extraction list, creating corresponding Jinja templates in `prompts/templates/`, and defining Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py). The `_extract_all_sections_separately` method iterates over the section list dynamically, making it straightforward to add new fields without modifying core orchestration logic.

### What happens if the LLM returns malformed JSON or invalid data?

The extraction pipeline includes error handling at the parsing and validation stages. If `json.loads` fails to parse the LLM response, or if the Pydantic models reject the data during `JSONResume(**complete_resume)` construction, the error is logged and the entire extraction aborts. This prevents partially-filled or corrupted résumé objects from being returned to downstream consumers.