# What Data Can Be Extracted from PDFs Using Hiring-Agent: The Complete Technical Guide

> Learn what data hiring-agent extracts from PDFs. This technical guide covers structured resume sections like work experience education skills and more into validated JSON objects.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-06-30

---

**The hiring-agent extracts six structured resume sections—basics, work experience, education, skills, projects, and awards—from PDFs and converts them into validated JSON objects using PyMuPDF for text extraction and LLM-driven parsing for semantic understanding.**

The interviewstreet/hiring-agent repository provides a robust pipeline for PDF data extraction, specifically designed to transform unstructured resume documents into machine-readable JSON. By combining PyMuPDF for document parsing with large language models for intelligent section segmentation, the tool converts arbitrary PDF resumes into strictly typed Pydantic models defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py).

## The PDF-to-JSON Extraction Pipeline

The extraction logic lives in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) within the `PDFHandler` class, which orchestrates a three-stage pipeline to convert binary PDF documents into structured `JSONResume` instances.

### Text Extraction with PyMuPDF

The `PDFHandler.extract_text_from_pdf` method handles initial document ingestion. It opens PDF files using **PyMuPDF**, iterates over all pages, and converts each page to Markdown format via `pymupdf_rag.to_markdown` ([source lines 52-58](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L52-L58)). This produces clean plain text that preserves document structure while removing binary formatting artifacts.

### LLM-Powered Section Parsing

Once raw text is available, `PDFHandler._extract_all_sections_separately` orchestrates the extraction of six major resume sections ([source lines 71-73](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L71-L73)). For each section, the system:

- Renders a **system prompt** specific to the target section using `TemplateManager` from [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) ([source lines 79-82](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L79-L82))
- Invokes the configured LLM provider via `self.provider.chat` with the rendered prompt and resume text as the user message ([source lines 88-94](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L88-L94))
- Extracts JSON from the LLM response using `extract_json_from_response`, then normalizes fields with `transform_parsed_data` ([source lines 110-119](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L110-L119))

### Pydantic Model Validation

After collecting all sections, the handler instantiates Pydantic models from [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)—including `Basics`, `Work`, `Education`, `Skills`, `Projects`, and `Awards`—and assembles them into a single `JSONResume` instance ([source lines 101-112](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L101-L112)). This ensures type safety and schema compliance before returning the final object.

## Six Structured Resume Sections Extracted

The hiring-agent specifically targets these six semantic categories when extracting data from PDFs:

- **Basics**: Personal information including name, contact details, and location
- **Work**: Employment history with company names, titles, dates, and descriptions
- **Education**: Academic institutions, degrees, fields of study, and graduation dates
- **Skills**: Technical competencies, languages, and proficiency classifications
- **Projects**: Personal or professional project titles, descriptions, and URLs
- **Awards**: Certifications, honors, achievements, and recognition details

## End-to-End Implementation Example

The public method `PDFHandler.extract_json_from_pdf` serves as the high-level entry point, tying together text extraction, section parsing, and model validation ([source lines 199-212](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py#L199-L212)). The [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) script demonstrates typical usage ([source lines 251-254](https://github.com/interviewstreet/hiring-agent/blob/main/score.py#L251-L254)).

```python
from pdf import PDFHandler

# 1️⃣ Initialise the handler (loads the LLM provider & templates)

handler = PDFHandler()

# 2️⃣ Pass a local PDF path – the method returns a JSONResume instance

resume_json = handler.extract_json_from_pdf("samples/jane_doe_resume.pdf")

if resume_json:
    # 3️⃣ Access structured fields

    print("Name:", resume_json.basics.name)
    print("Work experience entries:", len(resume_json.work))
    print("Skills:", resume_json.skills)
else:
    print("Failed to parse the PDF")

```

You can also run the extraction via the command line:

```bash
python score.py path/to/resume.pdf

```

## Summary

- The hiring-agent extracts **six specific resume sections** (basics, work, education, skills, projects, awards) from PDF documents using a hybrid approach.
- **PyMuPDF** and [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) handle the initial text-to-Markdown conversion in `PDFHandler.extract_text_from_pdf`.
- **LLM-driven parsing** via `_extract_all_sections_separately` extracts structured JSON from the plain text using section-specific prompts from `TemplateManager`.
- **Pydantic validation** in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) ensures the extracted data conforms to the `JSONResume` schema before returning the final object.
- The `extract_json_from_pdf` method in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) provides the primary public interface, as demonstrated in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).

## Frequently Asked Questions

### What specific resume sections can hiring-agent extract from PDFs?

The hiring-agent extracts six specific sections defined in the `JSONResume` schema: **basics** (contact information), **work** (employment history), **education** (academic background), **skills** (competencies), **projects** (portfolio items), and **awards** (certifications and honors). Each section is processed separately by the `_extract_all_sections_separately` method using dedicated LLM prompts to ensure accurate categorization.

### How does hiring-agent convert PDF content to structured JSON?

The conversion happens in three stages. First, `PDFHandler.extract_text_from_pdf` uses **PyMuPDF** to convert PDF pages to Markdown text. Second, `PDFHandler._extract_all_sections_separately` sends this text to an LLM with section-specific prompts from `TemplateManager` to extract structured data. Third, the extracted JSON is validated against Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) and assembled into a `JSONResume` object by `extract_json_from_pdf`.

### What Python libraries does hiring-agent use for PDF processing?

The repository uses **PyMuPDF** (fitz) for opening and reading PDF files, combined with a custom [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) helper that converts PyMuPDF page objects into Markdown format. This Markdown text is then processed by the LLM provider configured in the `PDFHandler` class, with final data validation handled by **Pydantic** models defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py).

### Can hiring-agent extract data from any PDF or only resumes?

According to the source code in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), the hiring-agent is specifically optimized for resume extraction. The `_extract_all_sections_separately` method and `TemplateManager` prompts are hardcoded to identify resume-specific sections (work, education, skills, etc.). While the PyMuPDF text extraction would work on any PDF, the LLM parsing logic expects resume formatting and would likely produce suboptimal results for non-resume documents.