How Hiring Agent Extracts Data from PDF Resumes: A Complete Technical Guide

Hiring Agent parses unstructured PDF resumes into structured JSON using a four-stage pipeline that combines PyMuPDF for text extraction, Jinja2 templating for LLM prompts, and Pydantic models for type-safe validation.

The hiring-agent repository by Interview Street provides an open-source solution for automated resume parsing. This tool transforms raw PDF documents into strongly-typed JSON objects through a sophisticated pipeline that isolates PDF parsing, prompt engineering, and data validation into distinct, testable stages.

The Four-Stage Extraction Pipeline

The extraction logic lives primarily in pdf.py and orchestrates the transformation through four distinct phases. Each stage handles a specific concern, making the system modular and extensible.

Stage 1: PDF to Markdown Conversion

The process begins in PDFHandler.extract_text_from_pdf (pdf.py:47-61), which utilizes PyMuPDF (fitz) to open the document. The method calls to_markdown from pymupdf_rag.py, walking every page and converting the visual layout into plain-text markdown. This produces a single string containing the entire resume content while preserving structural cues like headers and bullet points.

Stage 2: Prompt Generation with Jinja Templates

For each resume section—basics, work, education, skills, projects, and awards—the handler invokes the TemplateManager (prompts/template_manager.py:21-44). This component renders Jinja templates stored under prompts/templates/*.jinja, injecting the raw markdown text (text_content) into structured prompts. Each template contains system instructions and user prompts that direct the LLM to output JSON matching specific schema requirements.

Stage 3: LLM-Powered Data Extraction

The rendered prompt is sent to the configured LLM provider (Ollama or Gemini) via initialize_llm_provider. The _call_llm_for_section method (pdf.py:66-114) constructs the chat payload including model, messages, and options using global defaults (DEFAULT_MODEL, MODEL_PARAMETERS). When a return_model is specified, its JSON schema is passed via the format keyword argument, enabling structured output generation.

The raw LLM response is processed through extract_json_from_response (llm_utils.py), which strips the content to the first {…} block. The extracted JSON undergoes normalization via transform_parsed_data before validation.

Stage 4: Pydantic Model Construction

Parsed data for each section is instantiated into corresponding Pydantic section models defined in models.py:65-104. The handler creates BasicsSection, WorkSection, EducationSection, and other section objects, then assembles them into a final JSONResume object (pdf.py:261-279). This guarantees type safety and validation, ensuring downstream code receives a fully-typed representation. Any error during extraction causes immediate abortion, preventing partial or corrupted resume data.

Implementation Example

The following code demonstrates the public API for extracting structured data from PDF resumes:

from pdf import PDFHandler

# Initialize the handler (loads templates and selects the LLM provider)

handler = PDFHandler()

# Path to a candidate's resume PDF

pdf_path = "candidates/alice_smith_resume.pdf"

# Extract a typed JSONResume object

resume = handler.extract_json_from_pdf(pdf_path)

if resume:
    # Access structured fields safely

    print("Name:", resume.basics.name)
    print("Primary skills:", [s.name for s in resume.skills or []])
    print("Work history entries:", len(resume.work or []))
else:
    print("Failed to parse the resume.")

This entry point handles the complete pipeline internally, from PDF text extraction through final model validation.

Key Components and File Structure

Understanding the repository structure helps when extending or debugging the extraction process:

  • pdf.py: Core orchestrator containing PDFHandler class. Implements PDF-to-markdown conversion, section-wise LLM prompting, JSON parsing, and Pydantic model assembly.
  • models.py: Pydantic schema definitions for every resume section (BasicsSection, WorkSection, etc.) and the top-level JSONResume class.
  • prompts/template_manager.py: Loads and renders Jinja templates for each extraction section.
  • llm_utils.py: Provider initialization (initialize_llm_provider) and response cleaning utilities (extract_json_from_response).
  • pymupdf_rag.py: Implements to_markdown, converting PyMuPDF documents into structured markdown text.

Summary

  • Hiring Agent uses a four-stage pipeline to extract data from PDF resumes: PDF parsing, prompt generation, LLM extraction, and Pydantic validation.
  • PyMuPDF handles the initial text extraction via extract_text_from_pdf in pdf.py, converting visual layouts to markdown.
  • Jinja2 templates in prompts/template_manager.py create structured prompts for each resume section, injecting the raw markdown content.
  • LLM providers (Ollama or Gemini) process prompts through initialize_llm_provider, with JSON schema enforcement via the format parameter.
  • Pydantic models in models.py ensure type-safe validation, with JSONResume serving as the root container for all extracted data.
  • The public API extract_json_from_pdf encapsulates the entire workflow, returning fully-typed objects or aborting on any extraction error.

Frequently Asked Questions

How does Hiring Agent handle different PDF layouts and formats?

Hiring Agent relies on PyMuPDF's to_markdown function to normalize diverse PDF layouts into consistent markdown text. This preserves structural information like headers and lists while stripping formatting variations, allowing the LLM to process standardized text regardless of the original PDF's visual design.

Which LLM providers does Hiring Agent support for resume extraction?

The system supports Ollama and Gemini through the initialize_llm_provider utility in llm_utils.py. The provider's chat method handles the actual API calls, with configuration managed via global defaults including DEFAULT_MODEL and MODEL_PARAMETERS parameters.

What happens if the LLM returns malformed JSON or invalid data?

The pipeline includes multiple safeguards. The extract_json_from_response function isolates the first JSON block from the LLM output, while transform_parsed_data normalizes the structure. Finally, Pydantic validation in models.py enforces type constraints. If any stage fails, the entire extraction aborts to prevent partial data corruption, as implemented in pdf.py:261-279.

Can I extract specific sections without processing the entire resume?

Yes. The PDFHandler class provides individual extraction methods such as extract_basics_section, extract_work_section, and extract_skills_section. These methods follow the same pipeline but target specific sections, allowing modular extraction when you only need particular data points like work history or skills.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →