How Hiring-Agent Handles PDF Extraction: A Two-Stage Pipeline
Hiring-Agent extracts structured data from PDF resumes by first converting pages to markdown with PyMuPDF, then prompting an LLM with Jinja templates to parse individual sections into a JSONResume Pydantic model.
The open-source interviewstreet/hiring-agent repository implements a robust PDF extraction pipeline that transforms unstructured resume documents into machine-readable JSON. Unlike simple text scrapers, this system preserves semantic structure by chaining a layout-aware markdown converter with a large language model. The entire workflow is encapsulated in the PDFHandler class within pdf.py and integrated into the evaluation flow via score.py.
Stage 1: Raw Text Extraction with PyMuPDF
The pipeline begins by converting visual PDF content into clean markdown text. This stage leverages the fitz library (PyMuPDF) to maintain document hierarchy including headings, lists, and tables.
The extract_text_from_pdf Method
Located in pdf.py at lines 47-62, this method validates the file path, opens the document, and delegates markdown conversion to a helper function:
def extract_text_from_pdf(self, pdf_path: str) -> Optional[str]:
if not os.path.exists(pdf_path):
raise FileNotFoundError(...)
with pymupdf.open(pdf_path) as doc:
pages = range(doc.page_count)
resume_text = to_markdown(doc, pages=pages) # ← pymupdf_rag helper
return resume_text
Key implementation details include:
- Path validation raises
FileNotFoundErrorbefore attempting to open the document. pymupdf.opencreates a context-managed document object.to_markdown(defined inpymupdf_rag.py) performs the heavy lifting of layout analysis, converting each page into markdown syntax.- The concatenated string is returned for downstream processing, or
Noneif errors occur.
Stage 2: Structured JSON Generation via LLM
Once the resume exists as markdown text, the system invokes a series of LLM prompts to extract structured data. This approach separates concerns between content extraction (Stage 1) and semantic understanding (Stage 2).
Orchestrating Section Extraction
The entry point extract_json_from_pdf (lines 199-213 in pdf.py) coordinates the workflow:
def extract_json_from_pdf(self, pdf_path: str) -> Optional[JSONResume]:
text_content = self.extract_text_from_pdf(pdf_path)
if not text_content:
return None
return self._extract_all_sections_separately(text_content)
This method ensures text extraction succeeds before delegating to the section parser.
Processing Individual Sections
The _extract_all_sections_separately method (lines 66-74, 89-99, and 266-298) iterates over predefined resume categories:
def _extract_all_sections_separately(self, text_content: str) -> Optional[JSONResume]:
sections = ["basics", "work", "education", "skills", "projects", "awards"]
for section_name in sections:
section_data = self._extract_section_data(text_content, section_name)
if not section_data:
return None
complete_resume.update(section_data)
return JSONResume(**complete_resume)
For each section, the system renders a Jinja template (stored in prompts/templates/{section_name}.jinja) containing the full markdown resume and specific extraction instructions. The TemplateManager injects the text content into these templates before sending them to the LLM.
LLM Integration and Prompt Rendering
The _call_llm_for_section method (lines 66-100 and 105-118) constructs the chat payload and handles provider communication:
chat_params = {
"model": DEFAULT_MODEL,
"messages": [
{"role": "system", "content": section_system_message},
{"role": "user", "content": prompt},
],
"options": {
"stream": False,
"temperature": model_params["temperature"],
"top_p": model_params["top_p"],
},
}
response = self.provider.chat(**chat_params, **kwargs)
response_text = extract_json_from_response(response["message"]["content"])
parsed_data = json.loads(response_text)
The provider instance is initialized via initialize_llm_provider. The response undergoes JSON sanitization through extract_json_from_response before json.loads parses it into Python dictionaries.
Building the Final Model
After all six sections are successfully parsed, lines 311-319 instantiate the final data structure:
JSONResume(**complete_resume)
This Pydantic model (defined in models.py) validates the assembled dictionary against schema definitions for Basics, Work, Education, and other nested types, ensuring type safety throughout the application.
Integration with the Evaluation Pipeline
In score.py (lines 51-53), the extractor integrates into the broader hiring workflow:
pdf_handler = PDFHandler()
resume_data = pdf_handler.extract_json_from_pdf(pdf_path)
The score.py module instantiates PDFHandler without arguments, calls extract_json_from_pdf with the candidate's file path, and passes the resulting JSONResume object into subsequent evaluation logic.
Complete Working Example
To use the PDF extraction capabilities standalone:
from pdf import PDFHandler
pdf_path = "examples/jane_doe_resume.pdf"
handler = PDFHandler()
json_resume = handler.extract_json_from_pdf(pdf_path)
if json_resume:
print("✅ Extraction succeeded!")
print(json_resume.model_dump_json(indent=2))
else:
print("❌ Extraction failed.")
This demonstrates the public API surface: instantiate PDFHandler, invoke extract_json_from_pdf, and receive a validated Pydantic model or None if extraction fails.
Summary
- Two-stage architecture: PyMuPDF converts PDF to markdown, then an LLM parses sections into structured JSON.
- Core implementation: The
PDFHandlerclass inpdf.pyorchestrates extraction viaextract_text_from_pdfandextract_json_from_pdf. - Section parsing: Six distinct resume sections (basics, work, education, skills, projects, awards) are processed through individual Jinja templates.
- Type safety: All extracted data is validated against Pydantic models in
models.pybefore returning aJSONResumeobject. - Entry point:
score.pylines 51-53 demonstrate production usage of the extractor.
Frequently Asked Questions
What library does hiring-agent use for PDF extraction?
The system uses PyMuPDF (fitz) to open and read PDF documents. Specifically, pdf.py calls pymupdf.open() to access page content, then utilizes a helper function to_markdown from pymupdf_rag.py to convert the binary layout into semantic markdown text while preserving headings and table structures.
How does the system handle different resume sections?
Hiring-Agent processes resumes section-by-section using a predefined list: ["basics", "work", "education", "skills", "projects", "awards"]. For each section, _extract_all_sections_separately renders a specific Jinja template (e.g., basics.jinja, work.jinja) containing the full resume markdown and tailored extraction instructions, then sends this prompt to the LLM.
What happens if the LLM fails to parse a section?
If any section returns empty or invalid data from _extract_section_data, the _extract_all_sections_separately method immediately returns None (lines 89-99 in pdf.py). This fail-fast approach ensures that incomplete extractions do not propagate partial data downstream, maintaining data integrity for the evaluation pipeline.
Where is the final structured data defined?
The final output conforms to the JSONResume Pydantic model defined in models.py. This model includes nested structures for Basics, Work, Education, and other standard resume fields. Lines 311-319 in pdf.py instantiate this model with the aggregated section data, providing runtime validation and JSON serialization capabilities via model_dump_json().
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →