How to Test the Evaluation Pipeline with Sample Resumes in Hiring-Agent

You can test the hiring-agent evaluation pipeline by running python score.py <resume.pdf> against any sample PDF, which executes the full flow from PDF extraction through GitHub enrichment to final scoring while caching intermediate results for inspection.

The evaluation pipeline in the interviewstreet/hiring-agent repository is a fully automated end-to-end system that transforms raw PDF resumes into structured, scored evaluations. Understanding how to test this pipeline with sample resumes allows you to validate each processing stage—from initial text extraction to LLM-based scoring—without relying on production data.

Understanding the Pipeline Architecture

Before testing, you should understand the seven-stage flow implemented in the source code:

  1. PDF Ingestion - pymupdf_rag.py uses PyMuPDF to convert each page into a Markdown-ish string.
  2. Structured Parsing - pdf.py sends the Markdown to an LLM using Jinja templates stored in prompts/templates/*.jinja, returning JSON-Resume style objects.
  3. GitHub Enrichment - github.py fetches profile and repository data when a GitHub URL is detected, classifying projects and selecting the top 7.
  4. Evaluation Scoring - evaluator.py loads the resume evaluation criteria template, builds a chat prompt, and parses the structured EvaluationData response using extract_json_from_response.
  5. Orchestration - score.py wires all stages together, handling caching, CSV export (when DEVELOPMENT_MODE=True), and pretty printing.
  6. Model abstraction - models.py defines Pydantic schemas (JSONResume, EvaluationData) and provider wrappers.
  7. Data transformation - transform.py converts raw JSON into printable text and CSV rows.

Prerequisites and Environment Setup

First, clone the repository and install dependencies:

git clone https://github.com/interviewstreet/hiring-agent
cd hiring-agent
python -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate

pip install -r requirements.txt

Configure your environment by copying the example file:

cp .env.example .env

Edit .env to select your LLM provider. For local testing without external APIs, use Ollama:


# .env

LLM_PROVIDER=ollama
DEFAULT_MODEL=gemma3:4b
DEVELOPMENT_MODE=True

Set DEVELOPMENT_MODE=True to enable caching and CSV export during testing.

Creating a Sample Resume PDF

You need a PDF file to test the pipeline. Create a minimal test resume using Python's ReportLab:


# create_test_resume.py

from reportlab.lib.pagesizes import LETTER
from reportlab.pdfgen import canvas

c = canvas.Canvas("test_resume.pdf", pagesize=LETTER)
c.setFont("Helvetica", 12)
c.drawString(72, 720, "Jane Smith")
c.drawString(72, 700, "Senior Software Engineer")
c.drawString(72, 680, "Email: jane.smith@example.com")
c.drawString(72, 660, "GitHub: https://github.com/janesmith")
c.drawString(72, 640, "Skills: Python, Rust, Kubernetes")
c.drawString(72, 620, "Experience:")
c.drawString(90, 600, "TechCorp – Staff Engineer (2019-2024)")
c.showPage()
c.save()

Run the script to generate your test file:

python create_test_resume.py

This creates test_resume.pdf, which contains structured data including a GitHub URL to trigger the enrichment stage in github.py.

Running the Full Evaluation Pipeline

Execute the pipeline using the CLI entry point in score.py:

python score.py test_resume.pdf

The orchestration flow in score.py performs the following actions:

Inspecting Intermediate Artifacts

Because the pipeline caches intermediate results, you can inspect the JSON structures at each stage without re-running the LLM calls.

View the parsed resume structure:

cat cache/resumecache_test_resume.json | jq .

View the fetched GitHub data:

cat cache/githubcache_test_resume.json | jq .

To re-run the full pipeline fresh, delete the cache files:

rm cache/resumecache_test_resume.json cache/githubcache_test_resume.json

Subsequent runs without deletion will use cached data, making them nearly instant.

Configuring Alternative LLM Providers

While Ollama works offline, you can test with Google's Gemini for comparison.

Update your .env:

LLM_PROVIDER=gemini
DEFAULT_MODEL=gemini-2.5-pro
GEMINI_API_KEY=YOUR_KEY_HERE

Run the pipeline again:

python score.py test_resume.pdf

The provider logic in prompt.py automatically routes requests to the appropriate backend based on LLM_PROVIDER and DEFAULT_MODEL environment variables.

Summary

Testing the hiring-agent evaluation pipeline requires only a sample PDF and the score.py CLI command:

  • Environment: Set DEVELOPMENT_MODE=True to enable caching and CSV output in config.py.
  • Sample Data: Create a PDF with ReportLab or use any existing resume containing a GitHub URL to trigger full enrichment.
  • Execution: Run python score.py <file.pdf> to trigger the chain through pymupdf_rag.py, pdf.py, github.py, and evaluator.py.
  • Validation: Inspect cache/resumecache_*.json and cache/githubcache_*.json to verify extraction and enrichment stages.
  • Iteration: Delete cache files to force re-processing, or rely on caching for rapid iterative testing.

Frequently Asked Questions

How do I verify that the GitHub enrichment stage is working?

Check for the existence of cache/githubcache_<filename>.json after running score.py. If this file contains repository data and profile information, the github.py module successfully extracted the URL from your PDF and queried the GitHub API. If the file is missing or empty, verify that your sample resume includes a valid https://github.com/<username> URL format.

Can I test the pipeline without an internet connection?

Yes. Set LLM_PROVIDER=ollama in your .env file and ensure you have the Ollama server running locally with your chosen model (e.g., gemma3:4b). According to the models.py implementation, the Ollama provider routes requests to localhost:11434, requiring no external API calls. However, GitHub enrichment requires internet access unless you manually create a cached githubcache_*.json file before running.

What file handles the final scoring output format?

The evaluator.py file builds the evaluation prompt and parses the LLM response into an EvaluationData Pydantic model. The transform.py module then converts this structured data into both console-friendly text and CSV rows. When DEVELOPMENT_MODE=True, score.py appends results to resume_evaluations.csv using the helpers in transform.py.

How do I debug failures in the resume parsing stage?

Inspect the intermediate output in cache/resumecache_<filename>.json to see the raw JSON produced by the LLM in pdf.py. If this file is malformed or missing sections, check the Jinja templates in prompts/templates/ to ensure they match the expected schema defined in models.py for JSONResume. You can also enable verbose logging or modify pdf.py to print the raw Markdown output from pymupdf_rag.py before it reaches the LLM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →