# How the Hiring Agent Selects the Top 7 GitHub Projects for Evaluation

> Discover how the Hiring Agent selects top GitHub projects. Learn about author commit count, LLM-driven ranking for contribution volume, popularity, and complexity to find 7 unique projects.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-07-02

---

**The Hiring Agent filters repositories by a minimum author commit count of 4, then applies an LLM-driven ranking based on contribution volume, project popularity, and technical complexity to identify exactly 7 unique projects for evaluation.**

The `interviewstreet/hiring-agent` repository automates technical candidate screening by extracting and analyzing GitHub portfolios. Understanding how the system selects the top 7 GitHub projects for evaluation reveals a multi-stage pipeline that combines deterministic thresholds with AI-powered qualitative assessment.

## Data Collection and Metric Extraction

The selection process begins in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), where the `fetch_all_github_repos` function gathers every public repository associated with the extracted username. For each repository, the system computes **author_commit_count**—the definitive metric representing commits authored by the candidate—alongside auxiliary data including stars, forks, primary language, and topics.

The `generate_projects_json` function (lines 34-57) transforms this raw data into a structured JSON list. This normalization ensures the downstream ranking logic receives consistent inputs containing both quantitative activity metrics and qualitative project metadata.

## Hard Contribution Thresholds

Before any ranking occurs, the system enforces a strict eligibility rule: **only projects with `author_commit_count >= 4` qualify for consideration**. This threshold eliminates forks, template repositories, and superficial contributions.

Enforcement occurs at two independent layers for redundancy:

- **Template layer**: The Jinja prompt explicitly filters out repositories below this threshold (lines 44-48 in `prompts/templates/github_project_selection.jinja`)
- **LLM system prompt**: The system message reiterates the constraint (lines 55-57), instructing the model to disregard low-participation projects even if present in the input data

## The Prioritization Hierarchy

Once filtered, repositories are evaluated against a strictly ordered list of criteria defined in the template under "Selection Criteria (in order of importance)" (lines 14-22):

1. **Highest author commit count (≥ 15)** — Indicates substantial, sustained involvement
2. **Moderate author commit count (5-14)** — Signifies meaningful contribution without requiring project ownership
3. **Contributions to popular open-source projects (≥ 1,000 stars)** — Recognizes impact on widely-used libraries
4. **Technical complexity, real-world impact, code quality, community engagement, modern tech stack, originality** — Captures qualitative engineering excellence

This hierarchy ensures the LLM prioritizes demonstrated expertise and ownership over superficial popularity metrics.

## LLM-Driven Ranking and Enforcement

The prepared `projects_data` JSON is injected into the `github_project_selection` template. The LLM invocation (lines 78-86 in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)) includes a system message explicitly demanding **exactly 7 unique projects**, preventing variable list sizes or duplicate entries.

The template constrains the LLM to apply the prioritization hierarchy while respecting the hard commit threshold, combining deterministic rules with the model's ability to assess qualitative factors like architectural sophistication.

## Post-Processing Safeguards

After receiving the LLM response, the code at lines 96-115 in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) implements defensive validation:

- **Deduplication**: Removes duplicate project IDs from the returned array
- **Completeness check**: Verifies at least 7 unique projects are present
- **Fallback mechanism**: If the LLM returns fewer than 7 valid projects, the system automatically substitutes the top entries from a list sorted by `author_commit_count`

This ensures the final output always contains exactly 7 projects even if the LLM output is incomplete.

## Implementation Example

To retrieve the curated list of 7 projects for a candidate:

```python
from hiring_agent.github import fetch_and_display_github_info

# 1️⃣ Provide a GitHub profile URL (e.g., from a resume)

profile_url = "https://github.com/exampleUser"

# 2️⃣ Run the full enrichment pipeline

result = fetch_and_display_github_info(profile_url)

# 3️⃣ Access the exactly-7 selected projects

top_projects = result["projects"]  # List of dicts

for proj in top_projects:
    print(f"{proj['name']} – {proj['author_commit_count']} commits")

```

This end-to-end helper handles username extraction (`extract_github_username`), repository fetching (`fetch_all_github_repos`), JSON generation (`generate_projects_json`), and the LLM ranking pipeline automatically.

## Summary

- **Minimum threshold**: Projects require at least 4 author commits to qualify for evaluation
- **Ranking priority**: Commit count (highest first), then project popularity (≥1,000 stars), then technical quality metrics
- **Dual enforcement**: Thresholds are validated both in template logic and LLM system prompts
- **Guaranteed output**: Post-processing deduplicates and falls back to commit-count sorting to ensure exactly 7 projects are always returned
- **Key files**: [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) handles the pipeline, `github_project_selection.jinja` encodes the ranking rules

## Frequently Asked Questions

### What is the minimum number of commits required for a project to be considered?

A project must have an **author_commit_count of at least 4** to be eligible for selection. This threshold is enforced in both the Jinja template (`github_project_selection.jinja`, lines 44-48) and the LLM system prompt (lines 55-57) to ensure low-effort contributions are filtered out before ranking occurs.

### How does the system ensure exactly 7 projects are always returned?

The LLM is explicitly instructed to return exactly 7 unique projects in its system message ([`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), lines 78-86). Additionally, post-processing logic (lines 96-115) deduplicates results and implements a fallback mechanism that selects the top 7 projects by author commit count if the LLM returns fewer than 7 valid entries.

### What factors besides commit count influence project selection?

While commit count is the primary signal (prioritizing projects with ≥15 commits, then 5-14), the LLM also considers contributions to popular open-source projects (≥1,000 stars) and qualitative factors including technical complexity, real-world impact, code quality, community engagement, modern tech stack, and originality.

### Where are the selection criteria defined in the codebase?

The prioritization hierarchy and hard thresholds are defined in `prompts/templates/github_project_selection.jinja` (lines 14-22 for criteria ordering, lines 44-58 for threshold enforcement). The orchestration logic that invokes the LLM and handles post-processing resides in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py).