How to Retrieve Campaign History in the Hiring Agent Repository
The Hiring Agent repository does not implement a native campaign history feature, but you can reconstruct execution records using development-mode cache files and the resume_evaluations.csv log.
The interviewstreet/hiring-agent repository is an open-source pipeline designed to parse resumes, enrich them with GitHub signals, and generate structured candidate evaluations. If you are looking to retrieve campaign history—to track previous evaluation runs or audit pipeline executions—you should know that the codebase does not define "campaigns" as a domain concept. Instead, the project provides development-mode artifacts in config.py and score.py that serve as the closest equivalent to historical execution logs.
Understanding the Hiring Agent Architecture
The repository processes resumes through a sequential pipeline: PDF extraction, LLM-driven section parsing, GitHub profile enrichment, and final scoring. According to the source code in score.py, the orchestration logic handles end-to-end workflow execution, while evaluator.py applies fairness constraints and scoring rubrics. The GitHub enrichment module (github.py) fetches repository data, and the PDF conversion modules (pdf.py and pymupdf_rag.py) handle document ingestion. The models.py file defines Pydantic schemas for data validation, but contains no persistence logic for historical queries.
Why Campaign History Isn't Built-In
Campaign history implies a persistence layer that records metadata for each pipeline execution, typically including timestamps, input parameters, and output summaries. The Hiring Agent codebase contains no database models, REST endpoints, or CLI commands for storing or querying such records. The models.py file defines only runtime data structures and LLM provider abstractions, not historical storage entities. Consequently, there is no function or API to retrieve campaign history in the traditional sense.
Reconstructing Execution History via Development Caches
While native campaign tracking does not exist, the repository offers a development-mode caching mechanism that captures intermediate and final results. When enabled, this feature creates durable artifacts that function as a chronological audit trail.
Enabling Development Mode in config.py
The caching behavior is controlled by the DEVELOPMENT_MODE flag located in config.py. When set to True, the pipeline writes JSON-serialized intermediate results and appends final scores to a CSV file.
# config.py
DEVELOPMENT_MODE = True # Enables cache/ directory and CSV logging
Locating Cached JSON Artifacts
During execution, the pipeline writes two types of cache files to the cache/ directory:
cache/resumecache_<basename>.json– Stores the PDF-to-markdown extraction results frompdf.pycache/githubcache_<basename>.json– Stores the GitHub enrichment data fetched bygithub.py
These files provide granular insight into each pipeline stage and can be used to reconstruct the state of any given run.
Reading the resume_evaluations.csv Log
The final scores and metadata are appended to resume_evaluations.csv in the project root. This file serves as the closest analogue to a campaign history dashboard, containing a row for each completed evaluation executed through score.py.
Practical Code Examples for Retrieving Campaign Data
To generate and inspect campaign-style history, execute the pipeline and parse the resulting artifacts.
Running the pipeline:
import subprocess
pdf_path = "samples/resume.pdf"
# Execute CLI; development mode is typically on by default
subprocess.run(["python", "score.py", pdf_path], check=True)
Loading the CSV history:
import pandas as pd
df = pd.read_csv("resume_evaluations.csv")
print(df.tail()) # Display most recent evaluation runs
Inspecting cached intermediate results:
import json
import glob
# List all cached resume extractions
cache_files = glob.glob("cache/resumecache_*.json")
for file in sorted(cache_files):
with open(file, "r") as f:
data = json.load(f)
print(f"Processing history for: {file}")
# Access extraction timestamp or content as needed
Extending the Repository for True Campaign Tracking
If the development-mode cache is insufficient for production requirements, you must extend the architecture. Implementation paths include:
- Database Persistence: Integrate SQL or NoSQL storage in
score.pyto log run metadata, input checksums, and final scores - API Endpoints: Add FastAPI or Flask routes to query historical executions from the persistence layer
- Campaign CLI Commands: Extend the command-line interface in
score.pyto supportcampaign listorcampaign show <id>operations
Summary
- The interviewstreet/hiring-agent repository does not implement dedicated campaign history functionality
- Retrieve campaign history approximations are available through
cache/resumecache_*.json,cache/githubcache_*.json, andresume_evaluations.csvwhenDEVELOPMENT_MODE=Trueinconfig.py - Key source files involved include
score.py,config.py,evaluator.py, andgithub.py - For production-grade campaign tracking, you must implement a persistence layer and query interface on top of the existing pipeline
Frequently Asked Questions
Does the Hiring Agent repository have a database for campaign history?
No. The codebase contains no database models or storage layers for historical execution data. The models.py file only defines Pydantic schemas for runtime data validation and LLM provider configurations. Campaign history must be reconstructed from development cache files or implemented as a custom extension.
Where are campaign execution logs stored?
When DEVELOPMENT_MODE is enabled in config.py, execution artifacts are stored in the cache/ directory as JSON files (resumecache_*.json and githubcache_*.json), and final results are appended to resume_evaluations.csv in the project root. These locations provide the only durable record of past runs.
Can I retrieve campaign history via API or CLI?
Currently, there are no API endpoints or CLI commands to query campaign history. The score.py file provides only the main execution entry point. To support historical querying, you would need to build REST endpoints or additional CLI subcommands that interface with a persistent data store.
What is the difference between cache files and campaign history?
Cache files store intermediate processing results (PDF extraction and GitHub data) for debugging individual runs, while campaign history implies a structured record of all executions across time. The Hiring Agent's cache files serve a dual purpose as impromptu history logs, but they lack the relational structure and query capabilities of true campaign management systems.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →