# How Hiring-Agent Handles PDF Extraction: A Two-Stage Pipeline

> Discover how Hiring-Agent extracts structured data from PDFs. Learn about its two-stage pipeline: PyMuPDF conversion to markdown and LLM prompting for JSONResume Pydantic parsing.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-06-30

---

**Hiring-Agent extracts structured data from PDF resumes by first converting pages to markdown with PyMuPDF, then prompting an LLM with Jinja templates to parse individual sections into a JSONResume Pydantic model.**

The open-source **interviewstreet/hiring-agent** repository implements a robust **PDF extraction** pipeline that transforms unstructured resume documents into machine-readable JSON. Unlike simple text scrapers, this system preserves semantic structure by chaining a layout-aware markdown converter with a large language model. The entire workflow is encapsulated in the `PDFHandler` class within [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) and integrated into the evaluation flow via [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).

## Stage 1: Raw Text Extraction with PyMuPDF

The pipeline begins by converting visual PDF content into clean markdown text. This stage leverages the `fitz` library (PyMuPDF) to maintain document hierarchy including headings, lists, and tables.

### The extract_text_from_pdf Method

Located in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) at lines **47-62**, this method validates the file path, opens the document, and delegates markdown conversion to a helper function:

```python
def extract_text_from_pdf(self, pdf_path: str) -> Optional[str]:
    if not os.path.exists(pdf_path):
        raise FileNotFoundError(...)
    with pymupdf.open(pdf_path) as doc:
        pages = range(doc.page_count)
        resume_text = to_markdown(doc, pages=pages)   # ← pymupdf_rag helper

    return resume_text

```

Key implementation details include:

- **Path validation** raises `FileNotFoundError` before attempting to open the document.
- **`pymupdf.open`** creates a context-managed document object.
- **`to_markdown`** (defined in [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py)) performs the heavy lifting of layout analysis, converting each page into markdown syntax.
- The concatenated string is returned for downstream processing, or `None` if errors occur.

## Stage 2: Structured JSON Generation via LLM

Once the resume exists as markdown text, the system invokes a series of LLM prompts to extract structured data. This approach separates concerns between content extraction (Stage 1) and semantic understanding (Stage 2).

### Orchestrating Section Extraction

The entry point `extract_json_from_pdf` (lines **199-213** in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)) coordinates the workflow:

```python
def extract_json_from_pdf(self, pdf_path: str) -> Optional[JSONResume]:
    text_content = self.extract_text_from_pdf(pdf_path)
    if not text_content:
        return None
    return self._extract_all_sections_separately(text_content)

```

This method ensures text extraction succeeds before delegating to the section parser.

### Processing Individual Sections

The `_extract_all_sections_separately` method (lines **66-74**, **89-99**, and **266-298**) iterates over predefined resume categories:

```python
def _extract_all_sections_separately(self, text_content: str) -> Optional[JSONResume]:
    sections = ["basics", "work", "education", "skills", "projects", "awards"]
    for section_name in sections:
        section_data = self._extract_section_data(text_content, section_name)
        if not section_data:
            return None
        complete_resume.update(section_data)
    return JSONResume(**complete_resume)

```

For each section, the system renders a Jinja template (stored in `prompts/templates/{section_name}.jinja`) containing the full markdown resume and specific extraction instructions. The `TemplateManager` injects the text content into these templates before sending them to the LLM.

### LLM Integration and Prompt Rendering

The `_call_llm_for_section` method (lines **66-100** and **105-118**) constructs the chat payload and handles provider communication:

```python
chat_params = {
    "model": DEFAULT_MODEL,
    "messages": [
        {"role": "system", "content": section_system_message},
        {"role": "user",   "content": prompt},
    ],
    "options": {
        "stream": False,
        "temperature": model_params["temperature"],
        "top_p": model_params["top_p"],
    },
}
response = self.provider.chat(**chat_params, **kwargs)
response_text = extract_json_from_response(response["message"]["content"])
parsed_data = json.loads(response_text)

```

The provider instance is initialized via `initialize_llm_provider`. The response undergoes JSON sanitization through `extract_json_from_response` before `json.loads` parses it into Python dictionaries.

### Building the Final Model

After all six sections are successfully parsed, lines **311-319** instantiate the final data structure:

```python
JSONResume(**complete_resume)

```

This Pydantic model (defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)) validates the assembled dictionary against schema definitions for `Basics`, `Work`, `Education`, and other nested types, ensuring type safety throughout the application.

## Integration with the Evaluation Pipeline

In [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines **51-53**), the extractor integrates into the broader hiring workflow:

```python
pdf_handler = PDFHandler()
resume_data = pdf_handler.extract_json_from_pdf(pdf_path)

```

The [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) module instantiates `PDFHandler` without arguments, calls `extract_json_from_pdf` with the candidate's file path, and passes the resulting `JSONResume` object into subsequent evaluation logic.

## Complete Working Example

To use the PDF extraction capabilities standalone:

```python
from pdf import PDFHandler

pdf_path = "examples/jane_doe_resume.pdf"

handler = PDFHandler()
json_resume = handler.extract_json_from_pdf(pdf_path)

if json_resume:
    print("✅ Extraction succeeded!")
    print(json_resume.model_dump_json(indent=2))
else:
    print("❌ Extraction failed.")

```

This demonstrates the public API surface: instantiate `PDFHandler`, invoke `extract_json_from_pdf`, and receive a validated Pydantic model or `None` if extraction fails.

## Summary

- **Two-stage architecture**: PyMuPDF converts PDF to markdown, then an LLM parses sections into structured JSON.
- **Core implementation**: The `PDFHandler` class in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) orchestrates extraction via `extract_text_from_pdf` and `extract_json_from_pdf`.
- **Section parsing**: Six distinct resume sections (basics, work, education, skills, projects, awards) are processed through individual Jinja templates.
- **Type safety**: All extracted data is validated against Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) before returning a `JSONResume` object.
- **Entry point**: [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) lines **51-53** demonstrate production usage of the extractor.

## Frequently Asked Questions

### What library does hiring-agent use for PDF extraction?

The system uses **PyMuPDF** (`fitz`) to open and read PDF documents. Specifically, [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) calls `pymupdf.open()` to access page content, then utilizes a helper function `to_markdown` from [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) to convert the binary layout into semantic markdown text while preserving headings and table structures.

### How does the system handle different resume sections?

Hiring-Agent processes resumes section-by-section using a predefined list: `["basics", "work", "education", "skills", "projects", "awards"]`. For each section, `_extract_all_sections_separately` renders a specific Jinja template (e.g., `basics.jinja`, `work.jinja`) containing the full resume markdown and tailored extraction instructions, then sends this prompt to the LLM.

### What happens if the LLM fails to parse a section?

If any section returns empty or invalid data from `_extract_section_data`, the `_extract_all_sections_separately` method immediately returns `None` (lines **89-99** in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)). This fail-fast approach ensures that incomplete extractions do not propagate partial data downstream, maintaining data integrity for the evaluation pipeline.

### Where is the final structured data defined?

The final output conforms to the `JSONResume` Pydantic model defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py). This model includes nested structures for `Basics`, `Work`, `Education`, and other standard resume fields. Lines **311-319** in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) instantiate this model with the aggregated section data, providing runtime validation and JSON serialization capabilities via `model_dump_json()`.