How Hiring Agent Extracts Data from PDF Resumes: A Complete Technical Guide
Hiring Agent parses unstructured PDF resumes into structured JSON using a four-stage pipeline that combines PyMuPDF for text extraction, Jinja2 templating for LLM prompts, and Pydantic models for type-safe validation.
The hiring-agent repository by Interview Street provides an open-source solution for automated resume parsing. This tool transforms raw PDF documents into strongly-typed JSON objects through a sophisticated pipeline that isolates PDF parsing, prompt engineering, and data validation into distinct, testable stages.
The Four-Stage Extraction Pipeline
The extraction logic lives primarily in pdf.py and orchestrates the transformation through four distinct phases. Each stage handles a specific concern, making the system modular and extensible.
Stage 1: PDF to Markdown Conversion
The process begins in PDFHandler.extract_text_from_pdf (pdf.py:47-61), which utilizes PyMuPDF (fitz) to open the document. The method calls to_markdown from pymupdf_rag.py, walking every page and converting the visual layout into plain-text markdown. This produces a single string containing the entire resume content while preserving structural cues like headers and bullet points.
Stage 2: Prompt Generation with Jinja Templates
For each resume section—basics, work, education, skills, projects, and awards—the handler invokes the TemplateManager (prompts/template_manager.py:21-44). This component renders Jinja templates stored under prompts/templates/*.jinja, injecting the raw markdown text (text_content) into structured prompts. Each template contains system instructions and user prompts that direct the LLM to output JSON matching specific schema requirements.
Stage 3: LLM-Powered Data Extraction
The rendered prompt is sent to the configured LLM provider (Ollama or Gemini) via initialize_llm_provider. The _call_llm_for_section method (pdf.py:66-114) constructs the chat payload including model, messages, and options using global defaults (DEFAULT_MODEL, MODEL_PARAMETERS). When a return_model is specified, its JSON schema is passed via the format keyword argument, enabling structured output generation.
The raw LLM response is processed through extract_json_from_response (llm_utils.py), which strips the content to the first {…} block. The extracted JSON undergoes normalization via transform_parsed_data before validation.
Stage 4: Pydantic Model Construction
Parsed data for each section is instantiated into corresponding Pydantic section models defined in models.py:65-104. The handler creates BasicsSection, WorkSection, EducationSection, and other section objects, then assembles them into a final JSONResume object (pdf.py:261-279). This guarantees type safety and validation, ensuring downstream code receives a fully-typed representation. Any error during extraction causes immediate abortion, preventing partial or corrupted resume data.
Implementation Example
The following code demonstrates the public API for extracting structured data from PDF resumes:
from pdf import PDFHandler
# Initialize the handler (loads templates and selects the LLM provider)
handler = PDFHandler()
# Path to a candidate's resume PDF
pdf_path = "candidates/alice_smith_resume.pdf"
# Extract a typed JSONResume object
resume = handler.extract_json_from_pdf(pdf_path)
if resume:
# Access structured fields safely
print("Name:", resume.basics.name)
print("Primary skills:", [s.name for s in resume.skills or []])
print("Work history entries:", len(resume.work or []))
else:
print("Failed to parse the resume.")
This entry point handles the complete pipeline internally, from PDF text extraction through final model validation.
Key Components and File Structure
Understanding the repository structure helps when extending or debugging the extraction process:
pdf.py: Core orchestrator containingPDFHandlerclass. Implements PDF-to-markdown conversion, section-wise LLM prompting, JSON parsing, and Pydantic model assembly.models.py: Pydantic schema definitions for every resume section (BasicsSection,WorkSection, etc.) and the top-levelJSONResumeclass.prompts/template_manager.py: Loads and renders Jinja templates for each extraction section.llm_utils.py: Provider initialization (initialize_llm_provider) and response cleaning utilities (extract_json_from_response).pymupdf_rag.py: Implementsto_markdown, converting PyMuPDF documents into structured markdown text.
Summary
- Hiring Agent uses a four-stage pipeline to extract data from PDF resumes: PDF parsing, prompt generation, LLM extraction, and Pydantic validation.
- PyMuPDF handles the initial text extraction via
extract_text_from_pdfinpdf.py, converting visual layouts to markdown. - Jinja2 templates in
prompts/template_manager.pycreate structured prompts for each resume section, injecting the raw markdown content. - LLM providers (Ollama or Gemini) process prompts through
initialize_llm_provider, with JSON schema enforcement via theformatparameter. - Pydantic models in
models.pyensure type-safe validation, withJSONResumeserving as the root container for all extracted data. - The public API
extract_json_from_pdfencapsulates the entire workflow, returning fully-typed objects or aborting on any extraction error.
Frequently Asked Questions
How does Hiring Agent handle different PDF layouts and formats?
Hiring Agent relies on PyMuPDF's to_markdown function to normalize diverse PDF layouts into consistent markdown text. This preserves structural information like headers and lists while stripping formatting variations, allowing the LLM to process standardized text regardless of the original PDF's visual design.
Which LLM providers does Hiring Agent support for resume extraction?
The system supports Ollama and Gemini through the initialize_llm_provider utility in llm_utils.py. The provider's chat method handles the actual API calls, with configuration managed via global defaults including DEFAULT_MODEL and MODEL_PARAMETERS parameters.
What happens if the LLM returns malformed JSON or invalid data?
The pipeline includes multiple safeguards. The extract_json_from_response function isolates the first JSON block from the LLM output, while transform_parsed_data normalizes the structure. Finally, Pydantic validation in models.py enforces type constraints. If any stage fails, the entire extraction aborts to prevent partial data corruption, as implemented in pdf.py:261-279.
Can I extract specific sections without processing the entire resume?
Yes. The PDFHandler class provides individual extraction methods such as extract_basics_section, extract_work_section, and extract_skills_section. These methods follow the same pipeline but target specific sections, allowing modular extraction when you only need particular data points like work history or skills.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →