Where to Find Documentation for the Hiring-Agent: Complete Guide

The primary documentation for the hiring-agent is located in the repository's README.md file, which covers architecture, installation, configuration, and CLI usage, supplemented by inline docstrings in source files like score.py, pdf.py, and evaluator.py.

The interviewstreet/hiring-agent repository provides an open-source LLM-powered pipeline for extracting structured data from resume PDFs and evaluating candidates with fairness constraints. Understanding where to find documentation for the hiring-agent is essential for developers integrating the pipeline into recruitment workflows or extending the evaluation logic.

Primary Documentation Location

The central documentation hub lives at [README.md](https://github.com/interviewstreet/hiring-agent/blob/main/README.md) in the repository root. This file contains the architectural overview, installation prerequisites, environment variable configuration, and command-line usage instructions. For component-specific details, the source files contain extensive inline docstrings and comments that expand on the high-level architecture described in the README.

Architecture Overview

The hiring-agent follows a modular pipeline architecture. Each stage has dedicated source files with self-documenting code and specific responsibilities:

PDF Extraction Pipeline

The system processes PDF resumes through two complementary modules:

  • pymupdf_rag.py – Handles low-level PDF page extraction using PyMuPDF, converting pages to Markdown-like text.
  • pdf.py – Orchestrates section parsing and sends each resume section to the LLM using Jinja templates.

LLM Integration Layer

Provider abstractions and utility functions live in dedicated modules:

  • models.py – Contains Pydantic schemas and unified provider wrappers for both Ollama and Google Gemini APIs.
  • llm_utils.py – Provides helper utilities for initializing providers, handling requests, and cleaning LLM responses.

Prompt Templates

Structured extraction instructions are defined in the prompts/templates/ directory. These Jinja templates (such as basics.jinja and work.jinja) enforce strict formatting rules for each resume section parsed by the LLM.

GitHub Profile Enrichment

The github.py module extracts GitHub usernames from resume content, fetches profile and repository data, classifies projects by relevance, and uses the LLM to select the top seven contributions for evaluation.

Evaluation Engine

evaluator.py implements the fairness-aware scoring routine. It produces category scores, applies bonuses and deductions, and generates explanatory evidence for each scoring decision.

Pipeline Orchestration

score.py serves as the CLI entry point and orchestration layer. It wires all pipeline stages together, prints human-readable evaluation reports, and writes CSV output rows when DEVELOPMENT_MODE=True.

Configuration Management

config.py holds global configuration flags, primarily DEVELOPMENT_MODE. The README documents the required environment variables: LLM_PROVIDER, DEFAULT_MODEL, GEMINI_API_KEY, and GITHUB_TOKEN.

Quick Start Guide

To run the hiring-agent from the command line:


# Clone the repository and set up the environment

git clone https://github.com/interviewstreet/hiring-agent
cd hiring-agent
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

# Optional: Pull a local Ollama model

ollama pull gemma3:4b

# Run the evaluation pipeline on a resume

python score.py path/to/resume.pdf

Programmatic API Documentation

Beyond CLI usage, you can import modules directly for custom workflows.

Processing PDF Resumes

Extract structured data from PDFs using the PDFHandler class:

from pdf import PDFHandler
from models import JSONResume

# Initialize the handler (uses environment variables for LLM configuration)

handler = PDFHandler()

# Convert PDF to structured JSONResume object

resume: JSONResume = handler.process("path/to/resume.pdf")
print(resume.dict())

GitHub Data Enrichment

Enrich candidate profiles with GitHub metadata:

from github import GitHubEnricher

enricher = GitHubEnricher()

# Extract top 7 projects from the user's GitHub profile

github_data = enricher.enrich(resume.github_username)
print(github_data.top_projects)

Running Fairness-Aware Evaluations

Execute the scoring logic independently:

from evaluator import Evaluator

evaluator = Evaluator()
score_report = evaluator.evaluate(resume, github_data)
print(score_report.summary())

Configuration Reference

The hiring-agent requires specific environment variables to function:

  • LLM_PROVIDER – Set to ollama or gemini to select the backend.
  • DEFAULT_MODEL – Specifies the model name (e.g., gemma3:4b for Ollama or gemini-1.5-pro for Google).
  • GEMINI_API_KEY – Required when using Google Gemini as the provider.
  • GITHUB_TOKEN – Personal access token for GitHub API rate limits and private repo access.
  • DEVELOPMENT_MODE – Boolean flag in config.py that enables CSV output and JSON caching when set to True.

Summary

Frequently Asked Questions

Where is the main documentation for the hiring-agent?

The main documentation is located in the repository's [README.md](https://github.com/interviewstreet/hiring-agent/blob/main/README.md) file. It provides the architectural overview, installation steps, and configuration guide. For implementation details, refer to the inline docstrings within specific source files like pdf.py, github.py, and evaluator.py.

How do I configure the LLM provider for the hiring-agent?

Set the LLM_PROVIDER environment variable to either ollama or gemini. For Ollama, ensure the model is pulled locally (e.g., ollama pull gemma3:4b) and specify it in DEFAULT_MODEL. For Gemini, provide your GEMINI_API_KEY and set DEFAULT_MODEL to a valid Gemini model identifier like gemini-1.5-pro.

Can I use the hiring-agent programmatically instead of via CLI?

Yes. Import the relevant modules directly: use PDFHandler from pdf.py for resume extraction, GitHubEnricher from github.py for profile enrichment, and Evaluator from evaluator.py for scoring. These classes expose Python APIs that allow integration into custom applications beyond the score.py CLI entry point.

What are the key environment variables required to run the hiring-agent?

The essential environment variables are LLM_PROVIDER (selects the backend), DEFAULT_MODEL (specifies the model name), and GITHUB_TOKEN (enables GitHub API access). If using Google Gemini, you must also set GEMINI_API_KEY. The DEVELOPMENT_MODE flag in config.py controls whether the system outputs CSV files and caches intermediate JSON results.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →