How the Hiring-Agent Project Selection Algorithm Chooses Top GitHub Repositories
The Hiring-Agent project selection algorithm fetches all public repositories for a candidate, enriches them with contribution statistics and popularity metrics, sorts them by star count, and uses a large language model (LLM) to identify the seven most impressive and unique projects.
The interviewstreet/hiring-agent repository implements an intelligent pipeline that automates the evaluation of a software engineer's GitHub portfolio. Understanding how this project selection algorithm works helps developers optimize their public profiles and assists recruiters in interpreting automated candidate assessments. The system combines deterministic data aggregation with probabilistic AI reasoning to surface meaningful engineering work.
Fetching and Enriching Repository Data
The pipeline begins in github.py with the fetch_all_github_repos function (lines 18-84), which orchestrates the initial data harvest.
Extracting Candidate Metadata
The function first extracts the candidate's username from the supplied GitHub profile URL. It then queries the GitHub API's repos endpoint, specifically requesting results sorted by the updated timestamp to ensure recent activity.
Filtering Trivial Forks
Not every repository warrants consideration. The algorithm discards trivial forks using a hard filter: fork && forks_count < 5. This eliminates repositories that the candidate merely forked without substantial development effort, while retaining significant forks that indicate community engagement or maintenance contributions.
Computing Contribution Metrics
For each retained repository, the system calls fetch_repo_contributors and fetch_contributions_count to gather detailed statistics. It calculates the ratio of author commits versus total commits, classifying repositories as either open_source (more than one contributor) or self_project (solo work). The function stores a rich payload including stars, forks, primary language, and topics in the github_details dictionary.
Ranking by Popularity
After enrichment, the algorithm applies a deterministic sort to prioritize visibility. The code explicitly orders repositories by star count in descending order:
projects.sort(key=lambda x: x["github_details"]["stars"], reverse=True)
This sorting occurs at lines 80-81 in github.py, ensuring that the most popular and potentially influential projects appear first in the candidate's portfolio. This ranking serves as the foundation for the subsequent AI selection phase.
LLM-Powered Project Selection
The generate_projects_json function (lines 60-95 and 136-147) handles the intelligence layer of the project selection algorithm.
Preparing the Prompt
The function receives the sorted projects list and serializes it to JSON using json.dumps(projects_data, indent=2). It then renders the GitHub-project-selection Jinja template via TemplateManager.render_template, located at prompts/templates/github_project_selection.jinja. The system initializes the LLM provider through initialize_llm_provider and submits the rendered prompt with a strict system instruction demanding exactly 7 unique projects (lines 82-84).
Enforcing Uniqueness and Fallbacks
The algorithm implements robust error handling to guarantee result integrity. First, it deduplicates the LLM response by name to create unique_projects. If fewer than 7 distinct projects emerge, the system falls back to the next unused items from the original star-sorted list until the quota is met. If the LLM returns malformed JSON or unparseable output, the code triggers a hard fallback that selects the first 7 entries from the originally sorted list.
Code Implementation
You can interact with the project selection algorithm directly through the exposed functions:
# Fetch and select top projects for a candidate
from github import fetch_all_github_repos, generate_projects_json
candidate_url = "https://github.com/example-candidate"
repos = fetch_all_github_repos(candidate_url)
top_projects = generate_projects_json(repos)
# Each project contains selection metadata and GitHub details
for proj in top_projects:
print(f"{proj['name']}: {proj['github_details']['stars']} stars")
To parse LLM responses or initialize providers used in the selection process:
from llm_utils import initialize_llm_provider, extract_json_from_response
provider = initialize_llm_provider()
# Use provider.chat() and extract_json_from_response() for custom LLM workflows
Summary
- The project selection algorithm in
github.pyfilters trivial forks with fewer than 5 forks and enriches remaining repositories with contribution statistics. - Candidate repositories are sorted deterministically by star count before LLM evaluation.
- The system uses
generate_projects_jsonto prompt an LLM with a Jinja template, requiring exactly 7 unique project selections. - Fallback mechanisms ensure 7 projects are always returned, either through deduplication backfilling or defaulting to the top starred repositories.
- Supporting logic resides in
prompts/template_manager.pyfor rendering andllm_utils.pyfor LLM interaction and JSON extraction.
Frequently Asked Questions
How does the algorithm distinguish between open source contributions and personal projects?
The code classifies repositories based on contributor count. If fetch_repo_contributors returns more than one contributor, the repository is tagged as open_source; otherwise, it is classified as self_project. This distinction helps the LLM evaluate collaborative versus individual engineering capabilities.
What happens if a candidate has fewer than 7 repositories?
The algorithm attempts to fulfill the quota of exactly 7 projects through its fallback logic. If the LLM returns fewer than 7 unique selections, the system appends additional repositories from the star-sorted original list. If the LLM fails entirely, it defaults to the first 7 repositories from the sorted list, or fewer if the candidate lacks sufficient public repositories.
Why does the system filter forks with fewer than 5 forks?
The condition fork && forks_count < 5 identifies trivial forks—repositories that were likely forked for reference but never substantially modified or maintained. By filtering these out, the algorithm focuses on original work or significant forks that demonstrate active development or community maintenance, ensuring the final selection represents meaningful engineering effort.
Which files control the LLM prompt formatting and response parsing?
The Jinja template at prompts/templates/github_project_selection.jinja defines the prompt structure sent to the LLM. The TemplateManager class in prompts/template_manager.py handles template rendering. Response parsing and LLM provider initialization occur in llm_utils.py, specifically through extract_json_from_response and initialize_llm_provider.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →