Where to Find Documentation for the Hiring-Agent: Complete Guide
The primary documentation for the hiring-agent is located in the repository's README.md file, which covers architecture, installation, configuration, and CLI usage, supplemented by inline docstrings in source files like score.py, pdf.py, and evaluator.py.
The interviewstreet/hiring-agent repository provides an open-source LLM-powered pipeline for extracting structured data from resume PDFs and evaluating candidates with fairness constraints. Understanding where to find documentation for the hiring-agent is essential for developers integrating the pipeline into recruitment workflows or extending the evaluation logic.
Primary Documentation Location
The central documentation hub lives at [README.md](https://github.com/interviewstreet/hiring-agent/blob/main/README.md) in the repository root. This file contains the architectural overview, installation prerequisites, environment variable configuration, and command-line usage instructions. For component-specific details, the source files contain extensive inline docstrings and comments that expand on the high-level architecture described in the README.
Architecture Overview
The hiring-agent follows a modular pipeline architecture. Each stage has dedicated source files with self-documenting code and specific responsibilities:
PDF Extraction Pipeline
The system processes PDF resumes through two complementary modules:
pymupdf_rag.py– Handles low-level PDF page extraction using PyMuPDF, converting pages to Markdown-like text.pdf.py– Orchestrates section parsing and sends each resume section to the LLM using Jinja templates.
LLM Integration Layer
Provider abstractions and utility functions live in dedicated modules:
models.py– Contains Pydantic schemas and unified provider wrappers for both Ollama and Google Gemini APIs.llm_utils.py– Provides helper utilities for initializing providers, handling requests, and cleaning LLM responses.
Prompt Templates
Structured extraction instructions are defined in the prompts/templates/ directory. These Jinja templates (such as basics.jinja and work.jinja) enforce strict formatting rules for each resume section parsed by the LLM.
GitHub Profile Enrichment
The github.py module extracts GitHub usernames from resume content, fetches profile and repository data, classifies projects by relevance, and uses the LLM to select the top seven contributions for evaluation.
Evaluation Engine
evaluator.py implements the fairness-aware scoring routine. It produces category scores, applies bonuses and deductions, and generates explanatory evidence for each scoring decision.
Pipeline Orchestration
score.py serves as the CLI entry point and orchestration layer. It wires all pipeline stages together, prints human-readable evaluation reports, and writes CSV output rows when DEVELOPMENT_MODE=True.
Configuration Management
config.py holds global configuration flags, primarily DEVELOPMENT_MODE. The README documents the required environment variables: LLM_PROVIDER, DEFAULT_MODEL, GEMINI_API_KEY, and GITHUB_TOKEN.
Quick Start Guide
To run the hiring-agent from the command line:
# Clone the repository and set up the environment
git clone https://github.com/interviewstreet/hiring-agent
cd hiring-agent
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# Optional: Pull a local Ollama model
ollama pull gemma3:4b
# Run the evaluation pipeline on a resume
python score.py path/to/resume.pdf
Programmatic API Documentation
Beyond CLI usage, you can import modules directly for custom workflows.
Processing PDF Resumes
Extract structured data from PDFs using the PDFHandler class:
from pdf import PDFHandler
from models import JSONResume
# Initialize the handler (uses environment variables for LLM configuration)
handler = PDFHandler()
# Convert PDF to structured JSONResume object
resume: JSONResume = handler.process("path/to/resume.pdf")
print(resume.dict())
GitHub Data Enrichment
Enrich candidate profiles with GitHub metadata:
from github import GitHubEnricher
enricher = GitHubEnricher()
# Extract top 7 projects from the user's GitHub profile
github_data = enricher.enrich(resume.github_username)
print(github_data.top_projects)
Running Fairness-Aware Evaluations
Execute the scoring logic independently:
from evaluator import Evaluator
evaluator = Evaluator()
score_report = evaluator.evaluate(resume, github_data)
print(score_report.summary())
Configuration Reference
The hiring-agent requires specific environment variables to function:
LLM_PROVIDER– Set toollamaorgeminito select the backend.DEFAULT_MODEL– Specifies the model name (e.g.,gemma3:4bfor Ollama orgemini-1.5-profor Google).GEMINI_API_KEY– Required when using Google Gemini as the provider.GITHUB_TOKEN– Personal access token for GitHub API rate limits and private repo access.DEVELOPMENT_MODE– Boolean flag inconfig.pythat enables CSV output and JSON caching when set toTrue.
Summary
- The primary documentation for the hiring-agent resides in [
README.md](https://github.com/interviewstreet/hiring-agent/blob/main/README.md), covering installation, architecture, and CLI usage. - Source code documentation is embedded in module docstrings within
score.py,pdf.py,evaluator.py, andgithub.py. - The pipeline consists of distinct stages: PDF extraction (
pymupdf_rag.py,pdf.py), LLM interaction (models.py,llm_utils.py), GitHub enrichment (github.py), and fairness scoring (evaluator.py). - Configuration is managed through environment variables (
LLM_PROVIDER,DEFAULT_MODEL,GEMINI_API_KEY,GITHUB_TOKEN) and theDEVELOPMENT_MODEflag inconfig.py. - The system supports both CLI execution via
score.pyand programmatic integration using the Python API.
Frequently Asked Questions
Where is the main documentation for the hiring-agent?
The main documentation is located in the repository's [README.md](https://github.com/interviewstreet/hiring-agent/blob/main/README.md) file. It provides the architectural overview, installation steps, and configuration guide. For implementation details, refer to the inline docstrings within specific source files like pdf.py, github.py, and evaluator.py.
How do I configure the LLM provider for the hiring-agent?
Set the LLM_PROVIDER environment variable to either ollama or gemini. For Ollama, ensure the model is pulled locally (e.g., ollama pull gemma3:4b) and specify it in DEFAULT_MODEL. For Gemini, provide your GEMINI_API_KEY and set DEFAULT_MODEL to a valid Gemini model identifier like gemini-1.5-pro.
Can I use the hiring-agent programmatically instead of via CLI?
Yes. Import the relevant modules directly: use PDFHandler from pdf.py for resume extraction, GitHubEnricher from github.py for profile enrichment, and Evaluator from evaluator.py for scoring. These classes expose Python APIs that allow integration into custom applications beyond the score.py CLI entry point.
What are the key environment variables required to run the hiring-agent?
The essential environment variables are LLM_PROVIDER (selects the backend), DEFAULT_MODEL (specifies the model name), and GITHUB_TOKEN (enables GitHub API access). If using Google Gemini, you must also set GEMINI_API_KEY. The DEVELOPMENT_MODE flag in config.py controls whether the system outputs CSV files and caches intermediate JSON results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →