How to Manage Candidate Applications Within a Campaign Using Hiring Agent
You can manage candidate applications within a campaign by running each resume PDF through the score.py CLI, which automatically extracts structured data, calculates evaluation scores, and appends results to a CSV file for ranking and filtering.
Hiring Agent is an open-source pipeline from Interview Street that transforms candidate resumes into structured, explainable evaluation scores. By chaining the CLI workflow across multiple PDFs, you create a campaign—a repeatable process that aggregates applicant data into a single CSV for comparison. This approach ensures every candidate is evaluated against the same deterministic criteria, making it ideal for high-volume hiring rounds.
Campaign Architecture Overview
The repository implements a four-stage pipeline that standardizes how you manage candidate applications within a campaign. Each stage is modular, allowing you to re-run specific components without reprocessing unchanged data.
PDF Extraction and Structured Parsing
The pipeline begins in pdf.py, where the PDFHandler class converts each page of a résumé PDF into markdown-like text using PyMuPDF. It then feeds this content to an LLM (Ollama or Gemini) with Jinja prompts to extract JSON-Resume sections—basics, work experience, education, skills, projects, and awards. This structured extraction ensures that downstream evaluations receive consistent data regardless of PDF formatting variations.
GitHub Profile Enrichment
If the parser detects a GitHub username in the résumé, github.py automatically pulls the candidate's profile and repository data. The enrichment module selects the seven most relevant projects based on language usage and activity, then appends this metadata to the candidate's structured record. This step is crucial for technical campaigns where open-source contributions significantly impact scoring.
LLM-Based Evaluation Scoring
The evaluator.py file contains the ResumeEvaluator class, which combines the JSON-Resume data and GitHub enrichment into a single evaluation context. It makes a second LLM call that returns a strict-scored EvaluationData object containing overall scores, per-category breakdowns, bonus points, deductions, and qualitative strengths. This deterministic scoring ensures that rerunning the same resume yields identical results, enabling fair comparison across your candidate pool.
CSV Export and Aggregation
When DEVELOPMENT_MODE=True in config.py, the score.py orchestrator appends each evaluation to resume_evaluations.csv. This CSV serves as your campaign dashboard, capturing candidate names, overall scores, category-specific metrics, and links to source PDFs. Because the export appends rather than overwrites, you can safely process candidates incrementally without losing previous results.
Running a Campaign: Step-by-Step Workflow
Managing multiple applications requires treating each résumé as a unit of work and aggregating the outputs. The workflow below assumes you have activated your virtual environment and installed dependencies as documented in the repository README.
1. Prepare Your Candidate Directory
Place all candidate PDFs in a dedicated folder. The pipeline processes files individually, so organization is purely for your convenience.
mkdir -p ./candidates
cp ~/downloads/*.pdf ./candidates/
2. Process Individual Applications
Run the CLI against a single PDF to verify your configuration and view the evaluation output. This command executes the full pipeline: PDF extraction, GitHub enrichment (if applicable), LLM evaluation, and CSV logging.
python score.py candidates/alice_resume.pdf
The terminal displays a human-readable report, while resume_evaluations.csv receives a new row containing structured scores.
3. Execute Batch Processing
To manage candidate applications within a campaign at scale, wrap the CLI in a shell loop that iterates over your entire candidate directory.
#!/usr/bin/env bash
# run_campaign.sh - Batch evaluate allPDFs in ./candidates
set -euo pipefail
CAMPAIGN_DIR="./candidates"
CSV_OUT="resume_evaluations.csv"
# Optional: Start with fresh CSV
> "$CSV_OUT"
for pdf in "$CAMPAIGN_DIR"/*.pdf; do
echo "Evaluating $(basename "$pdf")..."
python score.py "$pdf"
done
echo "Campaign complete. Results in $CSV_OUT"
Execute the script with bash run_campaign.sh to generate a complete evaluation dataset.
4. Analyze Campaign Results
Load the aggregated CSV into pandas or any analytics tool to rank candidates, filter by minimum scores, or visualize category performance.
import pandas as pd
df = pd.read_csv("resume_evaluations.csv")
# Display top 5 candidates by overall score
top_candidates = df.nlargest(5, "overall_score")
print(top_candidates[["candidate_name", "overall_score",
"open_source_score", "technical_skills_score"]])
This workflow transforms raw PDFs into actionable hiring intelligence within minutes.
Key Configuration for Campaign Management
The config.py file controls campaign behavior through environment variables. Set DEVELOPMENT_MODE=True to enable CSV export functionality—without this flag, the pipeline runs evaluations but does not persist results to resume_evaluations.csv. You can also toggle between LLM providers (Ollama for local inference or Gemini for cloud-based processing) to balance speed against evaluation quality.
Summary
- Hiring Agent converts resume PDFs into structured JSON-Resume data using
pdf.pyand thePDFHandlerclass. - GitHub enrichment via
github.pyautomatically augments technical candidates' profiles with repository metrics. - Deterministic scoring through
evaluator.ResumeEvaluatorinevaluator.pyensures fair, reproducible candidate comparisons. - CSV aggregation occurs automatically when
DEVELOPMENT_MODE=True, creating a campaign-level dataset inresume_evaluations.csv. - Batch processing via shell scripts or task runners allows you to evaluate hundreds of applications while maintaining consistent scoring criteria.
Frequently Asked Questions
How does the pipeline handle duplicate candidate evaluations?
The score.py orchestrator appends each evaluation to resume_evaluations.csv without deduplication checks. If you run the same PDF twice, you will create duplicate rows. To prevent this, implement a preprocessing step that tracks processed filenames or clear the CSV before rerunning your campaign batch script.
What specific data appears in the CSV export?
The resume_evaluations.csv file contains the candidate name, overall score, per-category scores (open source, self projects, production experience, technical skills), maximum possible scores per category, total bonus points, total deductions, and a reference to the source PDF filename. This schema allows you to calculate percentages and weighted rankings using standard spreadsheet formulas.
Can I customize scoring criteria for different campaigns?
Yes, by modifying the evaluation prompts in evaluator.py or adjusting the EvaluationData model parameters. The pipeline is deterministic—using the same model and prompts produces identical scores—so you can version your prompt files and switch between them for different campaign types (e.g., junior vs. senior roles). Rerun your batch script with the new configuration to regenerate the CSV with updated criteria.
How do I enable CSV export for campaign tracking?
Set DEVELOPMENT_MODE=True in your config.py or environment variables. This flag activates the CSV logging logic within score.py. Without this configuration, the CLI outputs results to the terminal only and does not persist data to resume_evaluations.csv, making campaign-level analysis impossible.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →