Hiring Agent Scoring Formulas: Calculating Open Source, Self Projects, Production, and Technical Skills
The hiring agent evaluates resumes using a weighted system where Open Source contributes 35 points, Self Projects 30 points, Production Experience 25 points, and Technical Skills 10 points, with raw LLM scores capped to these weights and combined with optional bonuses for a maximum total of 120 points.
The interviewstreet/hiring-agent repository automates technical resume screening using a large language model (LLM) that assigns raw scores across four distinct categories. These raw scores are then normalized against fixed category weights and aggregated into a final evaluation. Understanding the exact formulas in score.py is essential for accurately interpreting candidate rankings and debugging evaluation outputs.
Understanding the Four Scoring Categories
Fixed Category Weights
The scoring system establishes hard ceilings for each category in score.py through the category_maxes dictionary (lines 78-85):
- Open Source: 35 points maximum
- Self Projects: 30 points maximum
- Production Experience: 25 points maximum
- Technical Skills: 10 points maximum
These weights represent the contribution of each category to the base 100-point evaluation scale.
How Raw LLM Scores Are Processed
The CategoryScore Data Model
Raw evaluation data originates from the LLM as structured objects defined in models.py (lines 18-22). The CategoryScore class captures the LLM's assessment:
class CategoryScore(BaseModel):
score: float # raw score from the LLM
max: int # maximum the LLM could assign (usually > category weight)
evidence: str # textual justification
The score field contains the LLM's numeric assessment, which often exceeds the category's weight allowance.
Capping Logic in score.py
To normalize LLM output against the fixed weights, score.py implements a capping mechanism (lines 88-94). Each raw score is constrained to its category maximum using min():
capped_score = min(os_score.score, category_maxes["open_source"])
This pattern applies identically to all four categories—Self Projects, Production Experience, and Technical Skills—ensuring no single category exceeds its allocated weight regardless of the LLM's scoring generosity.
Calculating the Overall Resume Score
The Aggregation Formula
The overall score computation follows this formula implemented in score.py:
total_score = Σ capped_category_score # Open-Source + Self-Projects + Production + Technical-Skills
total_score += bonus_points.total # optional bonus (max 20)
total_score -= deductions.total # optional deductions (non-negative)
The maximum possible overall score is calculated as:
max_possible_score = Σ category_weight + 20 # 35+30+25+10 = 100 → 120 with bonus
If total_score exceeds this ceiling, it is clamped to max_possible_score (see lines 65-69 in score.py).
Bonus Points and Deductions
The system allows for up to 20 bonus points based on exceptional criteria, while deductions subtract from the total. Both values are non-negative and applied after category capping.
Code Implementation Examples
Printing Evaluation Results
To display formatted results with capped scores, use the print_evaluation_results function from score.py:
from score import print_evaluation_results
from evaluator import ResumeEvaluator
# Assume `evaluation` is an EvaluationData object returned by the LLM
print_evaluation_results(evaluation, candidate_name="Jane Doe")
This outputs the categorized breakdown:
🌐 Open Source: 32/35
🚀 Self Projects: 28/30
🏢 Production Experience: 22/25
💻 Technical Skills: 9/10
⭐ BONUS POINTS: 15
⚠️ DEDUCTIONS: -3
Computing Scores Programmatically
For custom analytics or pipeline integration, replicate the capping logic:
def compute_overall(evaluation):
cat_weights = {"open_source": 35, "self_projects": 30,
"production": 25, "technical_skills": 10}
total = 0
for cat, weight in cat_weights.items():
raw = getattr(evaluation.scores, cat).score
total += min(raw, weight) # cap to weight
total += evaluation.bonus_points.total
total -= evaluation.deductions.total
max_possible = sum(cat_weights.values()) + 20
return min(total, max_possible)
overall = compute_overall(evaluation)
print(f"Overall score: {overall:.1f}/120")
Exporting Raw Scores to CSV
When generating CSV reports, the system preserves raw LLM scores (not capped values) in transform.py (lines 76-86):
from transform import transform_evaluation_response
row = transform_evaluation_response(
file_name="resume.pdf",
evaluation=evaluation,
resume_data=resume,
github_data=github_info,
)
# Access raw scores via:
# row["open_source_score"], row["open_source_max"], etc.
This distinction is critical for audit trails—CSV exports contain the LLM's original assessment, while the overall calculation uses the normalized, capped values.
Summary
- Category weights are fixed at 35/30/25/10 points respectively and defined in
score.py→category_maxes(lines 78-85) - Raw LLM scores are capped to category weights using
min(score, weight)before aggregation (lines 88-94) - Overall scoring sums capped categories, adds bonuses (max 20), subtracts deductions, and clamps to 120 points maximum (lines 65-69)
- CSV exports in
transform.pypreserve raw scores for auditing while calculations use normalized values
Frequently Asked Questions
What is the maximum possible score in Hiring Agent?
The absolute maximum is 120 points, comprising 100 points from the four weighted categories (35+30+25+10) plus up to 20 bonus points. The final score is capped at this value in score.py (lines 65-69) to prevent overflow from excessive LLM scores or bonus calculations.
Where are the scoring category weights defined?
The weights are hardcoded in the category_maxes dictionary within score.py at lines 78-85. This dictionary maps category names to their integer maximums: Open Source (35), Self Projects (30), Production Experience (25), and Technical Skills (10).
Why does the system cap raw LLM scores instead of using them directly?
Capping ensures consistent weighting across evaluations. Since different LLM prompts or model versions might return scores on varying scales, the min() operation normalizes all inputs to the fixed 100-point category framework, preventing any single category from dominating the final score due to LLM scoring inflation.
Does the CSV export contain capped scores or raw LLM scores?
The CSV export contains raw LLM scores as implemented in transform.py (lines 76-86). The row dictionaries include fields like open_source_score and open_source_max reflecting the LLM's original output, while the capped values used for final ranking are computed separately during the evaluation display phase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →