What Data Can Be Extracted from PDFs Using Hiring-Agent: The Complete Technical Guide
The hiring-agent extracts six structured resume sections—basics, work experience, education, skills, projects, and awards—from PDFs and converts them into validated JSON objects using PyMuPDF for text extraction and LLM-driven parsing for semantic understanding.
The interviewstreet/hiring-agent repository provides a robust pipeline for PDF data extraction, specifically designed to transform unstructured resume documents into machine-readable JSON. By combining PyMuPDF for document parsing with large language models for intelligent section segmentation, the tool converts arbitrary PDF resumes into strictly typed Pydantic models defined in models.py.
The PDF-to-JSON Extraction Pipeline
The extraction logic lives in pdf.py within the PDFHandler class, which orchestrates a three-stage pipeline to convert binary PDF documents into structured JSONResume instances.
Text Extraction with PyMuPDF
The PDFHandler.extract_text_from_pdf method handles initial document ingestion. It opens PDF files using PyMuPDF, iterates over all pages, and converts each page to Markdown format via pymupdf_rag.to_markdown (source lines 52-58). This produces clean plain text that preserves document structure while removing binary formatting artifacts.
LLM-Powered Section Parsing
Once raw text is available, PDFHandler._extract_all_sections_separately orchestrates the extraction of six major resume sections (source lines 71-73). For each section, the system:
- Renders a system prompt specific to the target section using
TemplateManagerfromprompts/template_manager.py(source lines 79-82) - Invokes the configured LLM provider via
self.provider.chatwith the rendered prompt and resume text as the user message (source lines 88-94) - Extracts JSON from the LLM response using
extract_json_from_response, then normalizes fields withtransform_parsed_data(source lines 110-119)
Pydantic Model Validation
After collecting all sections, the handler instantiates Pydantic models from models.py—including Basics, Work, Education, Skills, Projects, and Awards—and assembles them into a single JSONResume instance (source lines 101-112). This ensures type safety and schema compliance before returning the final object.
Six Structured Resume Sections Extracted
The hiring-agent specifically targets these six semantic categories when extracting data from PDFs:
- Basics: Personal information including name, contact details, and location
- Work: Employment history with company names, titles, dates, and descriptions
- Education: Academic institutions, degrees, fields of study, and graduation dates
- Skills: Technical competencies, languages, and proficiency classifications
- Projects: Personal or professional project titles, descriptions, and URLs
- Awards: Certifications, honors, achievements, and recognition details
End-to-End Implementation Example
The public method PDFHandler.extract_json_from_pdf serves as the high-level entry point, tying together text extraction, section parsing, and model validation (source lines 199-212). The score.py script demonstrates typical usage (source lines 251-254).
from pdf import PDFHandler
# 1️⃣ Initialise the handler (loads the LLM provider & templates)
handler = PDFHandler()
# 2️⃣ Pass a local PDF path – the method returns a JSONResume instance
resume_json = handler.extract_json_from_pdf("samples/jane_doe_resume.pdf")
if resume_json:
# 3️⃣ Access structured fields
print("Name:", resume_json.basics.name)
print("Work experience entries:", len(resume_json.work))
print("Skills:", resume_json.skills)
else:
print("Failed to parse the PDF")
You can also run the extraction via the command line:
python score.py path/to/resume.pdf
Summary
- The hiring-agent extracts six specific resume sections (basics, work, education, skills, projects, awards) from PDF documents using a hybrid approach.
- PyMuPDF and
pymupdf_rag.pyhandle the initial text-to-Markdown conversion inPDFHandler.extract_text_from_pdf. - LLM-driven parsing via
_extract_all_sections_separatelyextracts structured JSON from the plain text using section-specific prompts fromTemplateManager. - Pydantic validation in
models.pyensures the extracted data conforms to theJSONResumeschema before returning the final object. - The
extract_json_from_pdfmethod inpdf.pyprovides the primary public interface, as demonstrated inscore.py.
Frequently Asked Questions
What specific resume sections can hiring-agent extract from PDFs?
The hiring-agent extracts six specific sections defined in the JSONResume schema: basics (contact information), work (employment history), education (academic background), skills (competencies), projects (portfolio items), and awards (certifications and honors). Each section is processed separately by the _extract_all_sections_separately method using dedicated LLM prompts to ensure accurate categorization.
How does hiring-agent convert PDF content to structured JSON?
The conversion happens in three stages. First, PDFHandler.extract_text_from_pdf uses PyMuPDF to convert PDF pages to Markdown text. Second, PDFHandler._extract_all_sections_separately sends this text to an LLM with section-specific prompts from TemplateManager to extract structured data. Third, the extracted JSON is validated against Pydantic models in models.py and assembled into a JSONResume object by extract_json_from_pdf.
What Python libraries does hiring-agent use for PDF processing?
The repository uses PyMuPDF (fitz) for opening and reading PDF files, combined with a custom pymupdf_rag.py helper that converts PyMuPDF page objects into Markdown format. This Markdown text is then processed by the LLM provider configured in the PDFHandler class, with final data validation handled by Pydantic models defined in models.py.
Can hiring-agent extract data from any PDF or only resumes?
According to the source code in pdf.py, the hiring-agent is specifically optimized for resume extraction. The _extract_all_sections_separately method and TemplateManager prompts are hardcoded to identify resume-specific sections (work, education, skills, etc.). While the PyMuPDF text extraction would work on any PDF, the LLM parsing logic expects resume formatting and would likely produce suboptimal results for non-resume documents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →