How the Hiring Agent Extracts Structured Data from PDF Resumes: A Four-Stage Pipeline
The Hiring Agent converts unstructured PDF résumés into validated JSON objects using a four-stage pipeline that extracts markdown text with PyMuPDF, generates section-specific prompts via Jinja templates, processes them through an LLM provider (Ollama or Gemini), and validates the output against Pydantic models.
The interviewstreet/hiring-agent repository implements a robust extraction system that transforms raw PDF documents into strongly-typed data structures. This pipeline leverages PyMuPDF for document parsing, Jinja2 for prompt templating, and Pydantic for data validation to ensure reliable structured data extraction from PDF resumes.
Stage 1: Converting PDF Visual Layout to Markdown Text
The extraction process begins in pdf.py with the PDFHandler.extract_text_from_pdf method (lines 47–61). This function opens the PDF document using pymupdf.open(pdf_path) and converts the visual layout into plain-text markdown.
The conversion relies on the to_markdown function implemented in pymupdf_rag.py, which walks every page of the document and builds a markdown string preserving the document structure. The method processes all pages via to_markdown(doc, pages=range(doc.page_count)) and returns a single text_content string containing the entire résumé text.
Stage 2: Generating Section-Specific LLM Prompts
Once the markdown text is extracted, the system prepares structured prompts for each résumé section. The TemplateManager class in prompts/template_manager.py (lines 21–44) renders Jinja templates stored under prompts/templates/*.jinja.
The handler iterates over a fixed list of sections—basics, work, education, skills, projects, and awards—calling TemplateManager.render_template("<section>", text_content=resume_text) for each. Each template injects the raw markdown content into a system message that instructs the LLM to output a specific JSON structure matching the target Pydantic model for that section.
Stage 3: LLM Extraction and Response Normalization
The rendered prompts are sent to the selected LLM provider via the initialize_llm_provider utility (from llm_utils.py). The provider's chat method handles the actual inference, supporting both Ollama and Gemini backends.
In pdf.py (lines 66–114), the _call_llm_for_section method constructs the chat payload using DEFAULT_MODEL and MODEL_PARAMETERS from the global configuration. If the caller supplies a return_model, its JSON schema is passed to the LLM via the format keyword argument, enabling structured output generation.
The raw LLM response is processed through extract_json_from_response in llm_utils.py, which strips the content to the first JSON block. The extracted JSON is parsed with json.loads and normalized using transform_parsed_data to ensure consistent field mapping before model construction.
Stage 4: Pydantic Validation and Model Construction
The final stage transforms the parsed dictionaries into type-safe objects. In pdf.py (lines 261–279), each section's data is wrapped in its corresponding Pydantic model—BasicsSection, WorkSection, EducationSection, SkillsSection, ProjectsSection, or AwardsSection—as defined in models.py (lines 65–104).
These section objects are assembled into a root JSONResume model, creating a fully-typed, validated representation. The design ensures that any validation errors at this stage cause the entire extraction to abort, preventing partially-filled or corrupted résumé data from propagating downstream.
Complete Extraction Workflow
The PDFHandler orchestrates the full pipeline through the following execution flow:
- File validation –
extract_text_from_pdfraisesFileNotFoundErrorif the path is invalid. - Text extraction –
pymupdf.open()creates a document object;to_markdown()processes all pages into markdown. - Section iteration –
_extract_all_sections_separatelyloops through the fixed section list, delegating to_extract_section_dataand specificextract_*_sectionmethods. - Prompt rendering – Each section calls
TemplateManager.render_template()to generate the final LLM prompt. - LLM invocation –
_call_llm_for_sectionexecutes the chat completion with optional JSON schema constraints. - Response parsing – The raw message is stripped to the first
{…}block, parsed, and normalized viatransform_parsed_data. - Object instantiation – Section dictionaries are fed to Pydantic classes (
Basics(**…), etc.), culminating inJSONResume(**complete_resume).
Code Example: Extracting a Structured Resume
from pdf import PDFHandler
# Initialise the handler (loads templates and selects the LLM provider)
handler = PDFHandler()
# Path to a candidate's résumé PDF
pdf_path = "candidates/alice_smith_resume.pdf"
# Extract a typed JSONResume object
resume = handler.extract_json_from_pdf(pdf_path)
if resume:
# Access structured fields safely
print("Name:", resume.basics.name)
print("Primary skills:", [s.name for s in resume.skills or []])
print("Work history entries:", len(resume.work or []))
else:
print("Failed to parse the résumé.")
This snippet demonstrates the public API (PDFHandler.extract_json_from_pdf) that downstream components use to obtain fully-validated résumé data.
Summary
- PyMuPDF handles the initial PDF-to-markdown conversion in
pdf.py, preserving document layout while creating machine-readable text. - Jinja templates in
prompts/template_manager.pygenerate section-specific prompts that guide the LLM to output structured JSON. - LLM providers (Ollama/Gemini) process prompts through
initialize_llm_provider, with optional JSON schema enforcement via theformatparameter. - Pydantic models in
models.pyvalidate the extracted data, assembling sections into a type-safeJSONResumeobject. - The pipeline aborts on any extraction error, ensuring data integrity and preventing partial results.
Frequently Asked Questions
What PDF library does the Hiring Agent use to extract text?
The system uses PyMuPDF (pymupdf) to open PDF documents and extract content. Specifically, the to_markdown function in pymupdf_rag.py converts the visual layout into markdown text, preserving structural information while creating a format suitable for LLM processing.
How does the system ensure consistent JSON output from the LLM?
The TemplateManager renders Jinja templates containing strict system instructions that specify the required JSON schema. Additionally, when calling the LLM via _call_llm_for_section, the code passes the Pydantic model's JSON schema through the format keyword argument, which supported providers (Ollama and Gemini) can use to constrain their output to valid JSON structures.
Can I extend the pipeline to extract custom resume sections?
Yes. The modular architecture in pdf.py allows extension by adding new section names to the extraction list, creating corresponding Jinja templates in prompts/templates/, and defining Pydantic models in models.py. The _extract_all_sections_separately method iterates over the section list dynamically, making it straightforward to add new fields without modifying core orchestration logic.
What happens if the LLM returns malformed JSON or invalid data?
The extraction pipeline includes error handling at the parsing and validation stages. If json.loads fails to parse the LLM response, or if the Pydantic models reject the data during JSONResume(**complete_resume) construction, the error is logged and the entire extraction aborts. This prevents partially-filled or corrupted résumé objects from being returned to downstream consumers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →