How the GitHub Project Selection Algorithm Chooses the Top 7 Projects in Hiring-Agent
The GitHub project selection algorithm uses a three-phase pipeline that fetches all public repositories, sorts them by star count, and prompts an LLM to select exactly 7 unique projects where the candidate has authored at least 4 commits, with automatic fallback mechanisms if the LLM returns fewer results.
The interviewstreet/hiring-agent repository implements a deterministic hybrid approach to surface the most impressive repositories from a candidate's GitHub profile. This GitHub project selection algorithm combines automated data preprocessing with constrained LLM reasoning to ensure consistent output quality while handling edge cases gracefully.
Overview of the Three-Phase Selection Process
The algorithm operates through distinct phases that progress from data collection to final curation, as implemented in the core github.py module.
Phase 1: Data Aggregation and Star-Based Sorting
The process begins with fetch_all_github_repos in github.py:L18-L84, which retrieves all public repositories for a candidate via the GitHub API. Each repository object is enriched with contribution statistics including author_commit_count (the candidate's personal commits), total_commit_count, and star count.
After fetching, the system sorts the entire collection descending by github_details.stars at lines 80-85:
# Sorting logic from github.py
projects.sort(key=lambda x: x.get('github_details', {}).get('stars', 0), reverse=True)
This star-based ranking ensures that the most popular repositories are considered first during downstream selection.
Phase 2: LLM Evaluation Against Strict Criteria
The sorted repository list is converted to JSON via generate_projects_json and injected into the prompts/templates/github_project_selection.jinja template. This template encodes hard-coded selection rules including:
- Author commit threshold: Minimum 4 commits by the candidate
- Exact count requirement: Precisely 7 unique projects
- Priority weighting: Preference for high-commit-count and popular open-source contributions
The system message in github.py:L82-L84 reinforces these constraints:
system_msg = {
"role": "system",
"content": (
"You are an expert technical recruiter analyzing GitHub repositories to identify "
"the most impressive projects. CRITICAL: You must select exactly 7 UNIQUE projects "
"- no duplicates allowed. Each project must be different from the others."
),
}
The LLM receives the rendered template at lines 61-89 and returns a JSON array of selected projects.
Phase 3: Deduplication and Fallback Handling
The post-processing logic in github.py:L94-L131 and github.py:L140-L166 handles the LLM response through several defensive steps:
- Parsing: Extracts the JSON array from
response["message"]["content"] - Deduplication: Tracks
seen_namesto ensure uniqueness, keeping only the first occurrence of each project name - Fill-up logic: If fewer than 7 unique projects remain, the algorithm pulls additional entries from the pre-sorted
projects_datalist until the quota is met - Safe fallback: If parsing fails entirely,
github.py:L190-L216returns the first 7 entries from the star-sorted list
Deep Dive into the Selection Criteria
The algorithm enforces specific technical thresholds that balance contribution depth with project popularity.
The Author Commit Threshold
The Jinja template (lines 44-48 of github_project_selection.jinja) requires candidates to have substantive involvement, defined as author_commit_count >= 4. This filter eliminates forked repositories or minor contributions, ensuring selected projects represent meaningful development work.
The Jinja Template Rules
The template file located at prompts/templates/github_project_selection.jinja serves as the authoritative rule engine. Beyond the commit threshold, it instructs the LLM to:
- Prioritize projects with high commit counts and star ratings
- Avoid selecting multiple repositories from the same organization unless distinctly different
- Ensure the final output contains exactly 7 unique project names
Implementation Details and Code Examples
Developers can interact with the selection algorithm through the main entry point:
from github import fetch_and_display_github_info
# Fetch and select top 7 projects for a candidate
result = fetch_and_display_github_info("https://github.com/candidate-username")
projects = result["projects"] # JSON array of 7 selected repositories
Constructing the LLM Prompt
The prompt construction pipeline renders repository data with specific instructions:
# Excerpt from github.py showing prompt preparation
projects_json = generate_projects_json(projects_data)
prompt = template_manager.render(
"github_project_selection.jinja",
projects=projects_json,
candidate_name=candidate_name
)
Handling the Response and Fallback Logic
The robust fallback mechanism ensures 7 projects are always returned:
# From github.py:L140-L166 - deduplication and fill-up
unique_projects = []
seen_names = set()
for project in llm_selected_projects:
name = project.get("name")
if name not in seen_names:
unique_projects.append(project)
seen_names.add(name)
# Fill up to 7 if needed
if len(unique_projects) < 7:
for project in projects_data:
if project.get("name") not in seen_names:
unique_projects.append(project)
if len(unique_projects) == 7:
break
Summary
- The GitHub project selection algorithm in interviewstreet/hiring-agent combines deterministic preprocessing with LLM-based curation to identify the top 7 repositories.
- Phase 1 fetches and sorts all public repositories by star count in
github.py:L18-L84. - Phase 2 uses the
github_project_selection.jinjatemplate to enforce strict criteria including a minimum 4 author commits and exactly 7 unique selections. - Phase 3 implements defensive post-processing with deduplication, fill-up logic, and a safe fallback to the top 7 starred projects if the LLM response fails.
Frequently Asked Questions
What happens if a candidate has fewer than 7 qualifying repositories?
The algorithm first attempts to secure 7 unique projects from the LLM response. If the candidate lacks sufficient repositories meeting the author-commit threshold (≥4 commits), the fill-up logic in github.py:L140-L166 pulls additional projects from the star-sorted list regardless of commit count, ensuring the output always contains up to 7 entries.
Why does the algorithm require exactly 4 author commits?
The 4-commit threshold defined in github_project_selection.jinja serves as a heuristic for meaningful contribution. According to the template logic, this minimum ensures selected projects represent substantial development work rather than minor documentation fixes or automated forks, filtering out superficial contributions while maintaining sufficient portfolio breadth.
How does the system handle duplicate project names from the LLM?
The post-processing logic in github.py:L94-L131 maintains a seen_names set to track processed projects. When parsing the LLM's JSON response, the code skips any project whose name already exists in the set, preserving only the first occurrence. This guarantees uniqueness regardless of whether the LLM accidentally includes duplicates.
What is the fallback mechanism if the LLM fails to return valid JSON?
If the LLM response cannot be parsed as JSON, the code at github.py:L190-L216 executes a safe fallback that immediately returns the first 7 entries from the projects_data list (already sorted by stars). This deterministic path ensures the hiring pipeline continues functioning even during LLM service interruptions or malformed responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →