# How to Write Unit and Integration Tests for New Pipeline Stages in the Hiring-Agent Repository

> Learn to write effective unit and integration tests for hiring-agent pipeline stages. Mock external dependencies and test end-to-end orchestration with synthetic fixtures.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-05

---

**Mock external dependencies like LLM providers and GitHub APIs to test pipeline stages in isolation, then verify end-to-end orchestration with synthetic fixtures.**

The `interviewstreet/hiring-agent` repository implements a modular resume evaluation pipeline where each stage is encapsulated in separate modules. When you need to write unit and integration tests for new pipeline stages, you can leverage the existing loosely-coupled architecture to isolate components and verify their behavior independently.

## Understanding the Hiring-Agent Pipeline Architecture

The pipeline consists of five distinct stages, each exposing a well-defined interface that makes testing straightforward:

- **PDF extraction**: Implemented in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py) via the `PDFHandler` class, which uses **PyMuPDF** to convert resume PDFs to markdown-like text.
- **Section parsing**: Handled by `PDFHandler._call_llm_for_section` in the same file, sending extracted text to LLM providers (Ollama or Gemini) using Jinja templates.
- **GitHub enrichment**: Located in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), fetching profile and repository data via HTTP requests.
- **Evaluation**: The `ResumeEvaluator` class in [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) generates structured assessments using system-message templates.
- **Orchestration**: [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) coordinates all stages and outputs CSV results when `DEVELOPMENT_MODE=True`.

Because each stage is a public class with clear inputs and outputs, you can write unit tests that target them in isolation, while integration tests verify the complete data flow.

## Writing Unit Tests for Individual Pipeline Stages

Unit tests should reside in `tests/unit/` and use **pytest** (already included in [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt)). Focus on mocking external I/O to test logic in isolation.

### Testing PDF Extraction in [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)

For `PDFHandler.extract_text_from_pdf`, patch `pymupdf.open` to return a deterministic fake document rather than reading actual files.

```python

# tests/unit/test_pdf_handler.py

import pytest
from unittest.mock import MagicMock, patch
from pdf import PDFHandler

@pytest.fixture
def pdf_handler():
    return PDFHandler()

def test_extract_text_from_pdf_success(pdf_handler):
    fake_doc = MagicMock()
    fake_doc.page_count = 1
    fake_doc.__iter__.return_value = []
    with patch("pymupdf.open", return_value=fake_doc):
        text = pdf_handler.extract_text_from_pdf("dummy.pdf")
    assert isinstance(text, str)

```

### Testing LLM Integration Across Stages

Both `PDFHandler._call_llm_for_section` and `ResumeEvaluator.evaluate_resume` depend on `initialize_llm_provider` from `llm_utils`. Replace the provider with a stub that returns predefined JSON payloads.

```python

# tests/unit/test_evaluator.py

import json
from unittest.mock import MagicMock, patch
import pytest
from evaluator import ResumeEvaluator

@pytest.fixture
def evaluator():
    return ResumeEvaluator()

def test_evaluate_resume_parses_json(evaluator):
    fake_provider = MagicMock()
    fake_provider.chat.return_value = {
        "message": {"content": json.dumps({"overall_score": 85, "category_scores": {}})}
    }
    with patch("llm_utils.initialize_llm_provider", return_value=fake_provider):
        data = evaluator.evaluate_resume("sample resume text")
    assert data.overall_score == 85
    assert isinstance(data.category_scores, dict)

```

### Testing GitHub API Enrichment

In [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), patch `requests.get` (or the `httpx` client) to return fixture data instead of hitting the real GitHub API.

```python

# tests/unit/test_github.py

from unittest.mock import MagicMock, patch
import pytest
from github import fetch_github_projects

def test_github_enrichment_handles_empty_repos():
    mock_response = {"login": "alice", "public_repos": []}
    with patch("github.requests.get", return_value=MagicMock(json=lambda: mock_response)):
        result = fetch_github_projects("alice")
    assert result == []

```

### Validating Pydantic Data Models

The [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) file defines `JSONResume` and `EvaluationData` schemas. Test that invalid data raises `pydantic.ValidationError`.

```python

# tests/unit/test_models.py

import pytest
from pydantic import ValidationError
from models import JSONResume

def test_json_resume_rejects_malformed_data():
    bad_dict = {"basics": {"name": None}}  # invalid type

    with pytest.raises(ValidationError):
        JSONResume(**bad_dict)

```

## Creating Integration Tests for End-to-End Validation

Integration tests in `tests/integration/` verify that all stages cooperate correctly using a synthetic fixture.

### Setting Up Test Fixtures

Place a minimal PDF (e.g., 1 KB) under `tests/integration/fixtures/resume.pdf`. This provides a deterministic input without bundling large binary files.

### Mocking External Services in Integration Tests

Even in integration tests, avoid network calls by maintaining the same mocks used in unit tests. Use `monkeypatch` to override `initialize_llm_provider` and `requests.get` globally for the test session.

```python

# tests/integration/test_full_pipeline.py

import json
from pathlib import Path
from unittest.mock import MagicMock
import pytest

@pytest.fixture(autouse=True)
def mock_external_services(monkeypatch):
    fake_provider = MagicMock()
    fake_provider.chat.return_value = {
        "message": {"content": json.dumps({"basics": {"name": "Bob"}})}
    }
    monkeypatch.setattr(
        "llm_utils.initialize_llm_provider", lambda _: fake_provider
    )
    monkeypatch.setattr(
        "github.requests.get",
        lambda *_: type("Resp", (), {"json": lambda: {"login": "bob", "public_repos": []}})()
    )

```

### Verifying the Complete Pipeline Flow

Import the `main` function from [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) and assert that it returns valid objects and writes CSV output.

```python
def test_end_to_end(tmp_path):
    from score import main as score_main

    pdf_path = Path("tests/integration/fixtures/resume.pdf")
    result = score_main(pdf_path)

    assert result.basics.name == "Bob"
    
    csv_path = Path("resume_evaluations.csv")
    assert csv_path.is_file()
    last_row = csv_path.read_text().splitlines()[-1]
    assert "Bob" in last_row

```

## Recommended Test Directory Structure

Organize tests to mirror the source structure, keeping reusable mocks in [`conftest.py`](https://github.com/interviewstreet/hiring-agent/blob/main/conftest.py):

```

tests/
├── unit/
│   ├── test_pdf_handler.py
│   ├── test_github.py
│   ├── test_evaluator.py
│   ├── test_models.py
│   └── conftest.py          # shared fixtures and mock factories

└── integration/
    ├── test_full_pipeline.py
    └── fixtures/
        └── resume.pdf       # tiny synthetic PDF

```

Store reusable mock objects in [`conftest.py`](https://github.com/interviewstreet/hiring-agent/blob/main/conftest.py) so each unit test can import them via `from conftest import fake_llm_provider`.

## Automating Tests in CI/CD

Add a GitHub Actions workflow that leverages the existing [`requirements.txt`](https://github.com/interviewstreet/hiring-agent/blob/main/requirements.txt) to install dependencies and run the test suite.

```yaml
name: Tests
on: [push, pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"
      - name: Install dependencies
        run: pip install -r requirements.txt
      - name: Run tests
        run: pytest -q --maxfail=5 --disable-warnings

```

When adding new pipeline stages, create a corresponding unit test file in `tests/unit/` and extend the integration test to cover the new code path.

## Summary

- **Mock external dependencies**: Patch `pymupdf.open`, `initialize_llm_provider`, and `requests.get` to test logic without network calls or file I/O.
- **Isolate unit tests**: Target individual classes like `PDFHandler` and `ResumeEvaluator` in `tests/unit/` with deterministic inputs.
- **Verify integration**: Use a synthetic PDF fixture in `tests/integration/` to confirm all stages cooperate correctly in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).
- **Validate data models**: Ensure `JSONResume` and `EvaluationData` schemas reject malformed inputs using Pydantic validation tests.
- **Organize consistently**: Maintain separate `unit/` and `integration/` directories with shared fixtures in [`conftest.py`](https://github.com/interviewstreet/hiring-agent/blob/main/conftest.py).

## Frequently Asked Questions

### How do I mock the LLM provider when testing pipeline stages?

Patch `llm_utils.initialize_llm_provider` to return a `MagicMock` object with a `chat` method that returns a predefined dictionary structure. For example, `fake_provider.chat.return_value = {"message": {"content": '{"basics": {"name": "Alice"}}'}}` allows you to test section parsing in `PDFHandler._call_llm_for_section` without calling Ollama or Gemini.

### Where should I place test fixtures like sample PDFs?

Place minimal synthetic files under `tests/integration/fixtures/`. Keep these files small (under 10 KB) to avoid bloating the repository while still providing real binary data for end-to-end testing. The [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) orchestrator can process these just like production PDFs.

### How do I test the GitHub enrichment stage without hitting rate limits?

Mock the HTTP client used in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) by patching `requests.get` (or `httpx.get` if applicable) to return a fixture object with a `.json()` method. This isolates your tests from external API dependencies and eliminates flakiness caused by network issues or GitHub rate limiting.

### Should I test the TemplateManager class separately?

The `TemplateManager` in [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) only reads Jinja files from disk and renders them, so it typically does not require mocking unless you want to verify specific template names were requested. You can use the real implementation in most tests, as template rendering is lightweight and deterministic.