How to Write Unit and Integration Tests for New Pipeline Stages in the Hiring-Agent Repository

Mock external dependencies like LLM providers and GitHub APIs to test pipeline stages in isolation, then verify end-to-end orchestration with synthetic fixtures.

The interviewstreet/hiring-agent repository implements a modular resume evaluation pipeline where each stage is encapsulated in separate modules. When you need to write unit and integration tests for new pipeline stages, you can leverage the existing loosely-coupled architecture to isolate components and verify their behavior independently.

Understanding the Hiring-Agent Pipeline Architecture

The pipeline consists of five distinct stages, each exposing a well-defined interface that makes testing straightforward:

  • PDF extraction: Implemented in pdf.py via the PDFHandler class, which uses PyMuPDF to convert resume PDFs to markdown-like text.
  • Section parsing: Handled by PDFHandler._call_llm_for_section in the same file, sending extracted text to LLM providers (Ollama or Gemini) using Jinja templates.
  • GitHub enrichment: Located in github.py, fetching profile and repository data via HTTP requests.
  • Evaluation: The ResumeEvaluator class in evaluator.py generates structured assessments using system-message templates.
  • Orchestration: score.py coordinates all stages and outputs CSV results when DEVELOPMENT_MODE=True.

Because each stage is a public class with clear inputs and outputs, you can write unit tests that target them in isolation, while integration tests verify the complete data flow.

Writing Unit Tests for Individual Pipeline Stages

Unit tests should reside in tests/unit/ and use pytest (already included in requirements.txt). Focus on mocking external I/O to test logic in isolation.

Testing PDF Extraction in pdf.py

For PDFHandler.extract_text_from_pdf, patch pymupdf.open to return a deterministic fake document rather than reading actual files.


# tests/unit/test_pdf_handler.py

import pytest
from unittest.mock import MagicMock, patch
from pdf import PDFHandler

@pytest.fixture
def pdf_handler():
    return PDFHandler()

def test_extract_text_from_pdf_success(pdf_handler):
    fake_doc = MagicMock()
    fake_doc.page_count = 1
    fake_doc.__iter__.return_value = []
    with patch("pymupdf.open", return_value=fake_doc):
        text = pdf_handler.extract_text_from_pdf("dummy.pdf")
    assert isinstance(text, str)

Testing LLM Integration Across Stages

Both PDFHandler._call_llm_for_section and ResumeEvaluator.evaluate_resume depend on initialize_llm_provider from llm_utils. Replace the provider with a stub that returns predefined JSON payloads.


# tests/unit/test_evaluator.py

import json
from unittest.mock import MagicMock, patch
import pytest
from evaluator import ResumeEvaluator

@pytest.fixture
def evaluator():
    return ResumeEvaluator()

def test_evaluate_resume_parses_json(evaluator):
    fake_provider = MagicMock()
    fake_provider.chat.return_value = {
        "message": {"content": json.dumps({"overall_score": 85, "category_scores": {}})}
    }
    with patch("llm_utils.initialize_llm_provider", return_value=fake_provider):
        data = evaluator.evaluate_resume("sample resume text")
    assert data.overall_score == 85
    assert isinstance(data.category_scores, dict)

Testing GitHub API Enrichment

In github.py, patch requests.get (or the httpx client) to return fixture data instead of hitting the real GitHub API.


# tests/unit/test_github.py

from unittest.mock import MagicMock, patch
import pytest
from github import fetch_github_projects

def test_github_enrichment_handles_empty_repos():
    mock_response = {"login": "alice", "public_repos": []}
    with patch("github.requests.get", return_value=MagicMock(json=lambda: mock_response)):
        result = fetch_github_projects("alice")
    assert result == []

Validating Pydantic Data Models

The models.py file defines JSONResume and EvaluationData schemas. Test that invalid data raises pydantic.ValidationError.


# tests/unit/test_models.py

import pytest
from pydantic import ValidationError
from models import JSONResume

def test_json_resume_rejects_malformed_data():
    bad_dict = {"basics": {"name": None}}  # invalid type

    with pytest.raises(ValidationError):
        JSONResume(**bad_dict)

Creating Integration Tests for End-to-End Validation

Integration tests in tests/integration/ verify that all stages cooperate correctly using a synthetic fixture.

Setting Up Test Fixtures

Place a minimal PDF (e.g., 1 KB) under tests/integration/fixtures/resume.pdf. This provides a deterministic input without bundling large binary files.

Mocking External Services in Integration Tests

Even in integration tests, avoid network calls by maintaining the same mocks used in unit tests. Use monkeypatch to override initialize_llm_provider and requests.get globally for the test session.


# tests/integration/test_full_pipeline.py

import json
from pathlib import Path
from unittest.mock import MagicMock
import pytest

@pytest.fixture(autouse=True)
def mock_external_services(monkeypatch):
    fake_provider = MagicMock()
    fake_provider.chat.return_value = {
        "message": {"content": json.dumps({"basics": {"name": "Bob"}})}
    }
    monkeypatch.setattr(
        "llm_utils.initialize_llm_provider", lambda _: fake_provider
    )
    monkeypatch.setattr(
        "github.requests.get",
        lambda *_: type("Resp", (), {"json": lambda: {"login": "bob", "public_repos": []}})()
    )

Verifying the Complete Pipeline Flow

Import the main function from score.py and assert that it returns valid objects and writes CSV output.

def test_end_to_end(tmp_path):
    from score import main as score_main

    pdf_path = Path("tests/integration/fixtures/resume.pdf")
    result = score_main(pdf_path)

    assert result.basics.name == "Bob"
    
    csv_path = Path("resume_evaluations.csv")
    assert csv_path.is_file()
    last_row = csv_path.read_text().splitlines()[-1]
    assert "Bob" in last_row

Organize tests to mirror the source structure, keeping reusable mocks in conftest.py:


tests/
├── unit/
│   ├── test_pdf_handler.py
│   ├── test_github.py
│   ├── test_evaluator.py
│   ├── test_models.py
│   └── conftest.py          # shared fixtures and mock factories

└── integration/
    ├── test_full_pipeline.py
    └── fixtures/
        └── resume.pdf       # tiny synthetic PDF

Store reusable mock objects in conftest.py so each unit test can import them via from conftest import fake_llm_provider.

Automating Tests in CI/CD

Add a GitHub Actions workflow that leverages the existing requirements.txt to install dependencies and run the test suite.

name: Tests
on: [push, pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"
      - name: Install dependencies
        run: pip install -r requirements.txt
      - name: Run tests
        run: pytest -q --maxfail=5 --disable-warnings

When adding new pipeline stages, create a corresponding unit test file in tests/unit/ and extend the integration test to cover the new code path.

Summary

  • Mock external dependencies: Patch pymupdf.open, initialize_llm_provider, and requests.get to test logic without network calls or file I/O.
  • Isolate unit tests: Target individual classes like PDFHandler and ResumeEvaluator in tests/unit/ with deterministic inputs.
  • Verify integration: Use a synthetic PDF fixture in tests/integration/ to confirm all stages cooperate correctly in score.py.
  • Validate data models: Ensure JSONResume and EvaluationData schemas reject malformed inputs using Pydantic validation tests.
  • Organize consistently: Maintain separate unit/ and integration/ directories with shared fixtures in conftest.py.

Frequently Asked Questions

How do I mock the LLM provider when testing pipeline stages?

Patch llm_utils.initialize_llm_provider to return a MagicMock object with a chat method that returns a predefined dictionary structure. For example, fake_provider.chat.return_value = {"message": {"content": '{"basics": {"name": "Alice"}}'}} allows you to test section parsing in PDFHandler._call_llm_for_section without calling Ollama or Gemini.

Where should I place test fixtures like sample PDFs?

Place minimal synthetic files under tests/integration/fixtures/. Keep these files small (under 10 KB) to avoid bloating the repository while still providing real binary data for end-to-end testing. The score.py orchestrator can process these just like production PDFs.

How do I test the GitHub enrichment stage without hitting rate limits?

Mock the HTTP client used in github.py by patching requests.get (or httpx.get if applicable) to return a fixture object with a .json() method. This isolates your tests from external API dependencies and eliminates flakiness caused by network issues or GitHub rate limiting.

Should I test the TemplateManager class separately?

The TemplateManager in prompts/template_manager.py only reads Jinja files from disk and renders them, so it typically does not require mocking unless you want to verify specific template names were requested. You can use the real implementation in most tests, as template rendering is lightweight and deterministic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →