Complete Data Flow from PDF Input to Evaluation Output in Hiring Agent

Hiring Agent transforms a raw résumé PDF into a structured evaluation score through eight deterministic stages: PDF extraction, markdown conversion, section-wise LLM parsing, JSONResume assembly, optional GitHub enrichment, text transformation, LLM-based evaluation, and final scoring presentation.

The interviewstreet/hiring-agent repository implements a fully automated pipeline that converts unstructured résumé documents into quantified hiring assessments. Understanding the complete data flow from PDF input to evaluation output reveals how modular Python components and Large Language Model (LLM) integrations work together to produce standardized candidate scores.

Step 1: PDF Ingestion and Markdown Extraction

CLI Entry Point in score.py

The pipeline begins in score.py where the main(pdf_path) function serves as the command-line interface. When invoked via python score.py résumé.pdf, the function creates a cache identifier, checks for existing cached JSON when DEVELOPMENT_MODE=True, and instantiates a PDFHandler to begin extraction.

Low-Level Text Extraction with PyMuPDF

Inside pdf.py, the PDFHandler.extract_text_from_pdf method opens the PDF using PyMuPDF and delegates markdown conversion to to_markdown defined in pymupdf_rag.py. This function iterates through each page, detects structural elements including headings, tables, and images, and returns a Markdown-formatted string that preserves the document's semantic layout.

Step 2: Section-Wise Structured Parsing

Template-Driven LLM Calls in pdf.py

The extracted markdown feeds into a sophisticated parsing system within pdf.py. The handler renders Jinja templates stored under prompts/templates/—including basics.jinja, work.jinja, education.jinja, skills.jinja, projects.jinja, and awards.jinja—to create section-specific prompts. Each request combines the rendered template with a system message from system_message.jinja and sends it to the active LLM provider via self.provider.chat.

Data Transformation and JSONResume Assembly

The LLM response undergoes cleaning and JSON decoding before transform_parsed_data from transform.py post-processes the output to match the JSONResume Pydantic schema. The _extract_all_sections_separately method orchestrates this process across all six core sections, assembling a complete JSONResume instance that structures the candidate's information into typed fields.

Step 3: GitHub Profile Enrichment

After constructing the base resume object, score.py examines resume_data.basics.profiles for entries where the network field equals "Github". When found, fetch_and_display_github_info from github.py retrieves the public user profile and repository list. The system then applies the github_project_selection.jinja template to prompt the LLM to identify up to seven representative projects, enriching the candidate profile with verifiable open-source contributions.

Step 4: Text Conversion for Evaluation

Before final scoring, structured data converts back to plain text using helper functions in transform.py. The convert_json_resume_to_text, convert_github_data_to_text, and convert_blog_data_to_text functions flatten the JSONResume object and GitHub data into concatenated text blocks. These strings form the contextual foundation for the evaluation prompts.

Step 5: Final Evaluation and Scoring

LLM-Based Assessment in evaluator.py

The ResumeEvaluator.evaluate_resume method in evaluator.py loads two critical templates: resume_evaluation_criteria.jinja (containing the detailed scoring rubric) and resume_evaluation_system_message.jinja (providing system-level instructions). The method sends a structured output request to the LLM provider with format=EvaluationData.model_json_schema(), ensuring the response conforms to the expected schema.

Score Computation and Output

The returned JSON payload parses into an EvaluationData Pydantic model. Back in score.py, print_evaluation_results calculates total scores, category breakdowns, bonuses, and deductions, presenting a human-readable report. When running in DEVELOPMENT_MODE, the system appends results to resume_evaluations.csv for batch analysis and tracking.

Complete Code Examples

Run the full pipeline from the command line:

python score.py path/to/resume.pdf

Use the library programmatically to access the evaluation data:

from score import main

evaluation = main("resume.pdf")

# evaluation is an instance of models.EvaluationData

print(evaluation.scores.open_source.score)

Extract the intermediate JSON resume without invoking the evaluator:

from pdf import PDFHandler

handler = PDFHandler()
json_resume = handler.extract_json_from_pdf("resume.pdf")
print(json_resume.model_dump())

Fetch GitHub enrichment data independently:

from github import fetch_and_display_github_info

github_info = fetch_and_display_github_info("https://github.com/username")
print(github_info["profile"]["login"])

Summary

  • Entry Point: The main(pdf_path) function in score.py orchestrates the entire pipeline with optional caching via DEVELOPMENT_MODE.
  • PDF Processing: pymupdf_rag.py converts PDF bytes to Markdown, preserving document structure through PyMuPDF.
  • Structured Extraction: pdf.py uses section-specific Jinja templates to prompt LLMs, with transform.py ensuring schema compliance via transform_parsed_data.
  • External Enrichment: github.py conditionally augments profiles using fetch_and_display_github_info and project selection prompts.
  • Evaluation: evaluator.py generates standardized scores using structured output requests to EvaluationData models.
  • Output: Final results display through print_evaluation_results in score.py, with optional CSV export to resume_evaluations.csv.

Frequently Asked Questions

How does Hiring Agent handle different PDF formats?

Hiring Agent uses PyMuPDF in pymupdf_rag.py to normalize diverse PDF layouts into Markdown. The to_markdown function detects headings, tables, and images across pages, creating a consistent text representation that subsequent LLM prompts can process regardless of the original document's formatting.

What LLM providers does Hiring Agent support?

The codebase abstracts provider selection through self.provider.chat interfaces in pdf.py and evaluator.py. The system supports Ollama and Gemini configurations, with provider mapping and API keys managed centrally in prompt.py and config.py.

Can I run the pipeline without the GitHub enrichment step?

Yes. The GitHub enrichment in score.py only executes when the candidate's basics.profiles contains a GitHub URL. If no such profile exists or if the network field does not match "Github", the pipeline proceeds directly to evaluation using only the parsed resume data.

Where is the evaluation data cached?

When DEVELOPMENT_MODE=True, score.py caches intermediate JSON results and writes final evaluations to resume_evaluations.csv. The cache name derives from the input PDF path, allowing rapid re-running of the pipeline without repeating LLM calls during development and testing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →