# How the Hiring-Agent Project Selection Algorithm Chooses Top GitHub Repositories

> Discover how the Hiring-Agent algorithm selects top GitHub repositories. It analyzes contributions, popularity, and LLM insights to find the best projects.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-06

---

**The Hiring-Agent project selection algorithm fetches all public repositories for a candidate, enriches them with contribution statistics and popularity metrics, sorts them by star count, and uses a large language model (LLM) to identify the seven most impressive and unique projects.**

The **interviewstreet/hiring-agent** repository implements an intelligent pipeline that automates the evaluation of a software engineer's GitHub portfolio. Understanding how this project selection algorithm works helps developers optimize their public profiles and assists recruiters in interpreting automated candidate assessments. The system combines deterministic data aggregation with probabilistic AI reasoning to surface meaningful engineering work.

## Fetching and Enriching Repository Data

The pipeline begins in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) with the `fetch_all_github_repos` function (lines 18-84), which orchestrates the initial data harvest.

### Extracting Candidate Metadata

The function first extracts the candidate's username from the supplied GitHub profile URL. It then queries the GitHub API's *repos* endpoint, specifically requesting results sorted by the `updated` timestamp to ensure recent activity.

### Filtering Trivial Forks

Not every repository warrants consideration. The algorithm discards **trivial forks** using a hard filter: `fork && forks_count < 5`. This eliminates repositories that the candidate merely forked without substantial development effort, while retaining significant forks that indicate community engagement or maintenance contributions.

### Computing Contribution Metrics

For each retained repository, the system calls `fetch_repo_contributors` and `fetch_contributions_count` to gather detailed statistics. It calculates the ratio of **author commits** versus **total commits**, classifying repositories as either **open_source** (more than one contributor) or **self_project** (solo work). The function stores a rich payload including stars, forks, primary language, and topics in the `github_details` dictionary.

## Ranking by Popularity

After enrichment, the algorithm applies a deterministic sort to prioritize visibility. The code explicitly orders repositories by star count in descending order:

```python
projects.sort(key=lambda x: x["github_details"]["stars"], reverse=True)

```

This sorting occurs at lines 80-81 in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), ensuring that the most popular and potentially influential projects appear first in the candidate's portfolio. This ranking serves as the foundation for the subsequent AI selection phase.

## LLM-Powered Project Selection

The `generate_projects_json` function (lines 60-95 and 136-147) handles the intelligence layer of the project selection algorithm.

### Preparing the Prompt

The function receives the sorted `projects` list and serializes it to JSON using `json.dumps(projects_data, indent=2)`. It then renders the **GitHub-project-selection** Jinja template via `TemplateManager.render_template`, located at `prompts/templates/github_project_selection.jinja`. The system initializes the LLM provider through `initialize_llm_provider` and submits the rendered prompt with a strict system instruction demanding exactly **7 unique projects** (lines 82-84).

### Enforcing Uniqueness and Fallbacks

The algorithm implements robust error handling to guarantee result integrity. First, it deduplicates the LLM response by `name` to create `unique_projects`. If fewer than 7 distinct projects emerge, the system falls back to the next unused items from the original star-sorted list until the quota is met. If the LLM returns malformed JSON or unparseable output, the code triggers a hard fallback that selects the first 7 entries from the originally sorted list.

## Code Implementation

You can interact with the project selection algorithm directly through the exposed functions:

```python

# Fetch and select top projects for a candidate

from github import fetch_all_github_repos, generate_projects_json

candidate_url = "https://github.com/example-candidate"
repos = fetch_all_github_repos(candidate_url)
top_projects = generate_projects_json(repos)

# Each project contains selection metadata and GitHub details

for proj in top_projects:
    print(f"{proj['name']}: {proj['github_details']['stars']} stars")

```

To parse LLM responses or initialize providers used in the selection process:

```python
from llm_utils import initialize_llm_provider, extract_json_from_response

provider = initialize_llm_provider()

# Use provider.chat() and extract_json_from_response() for custom LLM workflows

```

## Summary

- The **project selection algorithm** in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) filters trivial forks with fewer than 5 forks and enriches remaining repositories with contribution statistics.
- Candidate repositories are sorted deterministically by star count before LLM evaluation.
- The system uses `generate_projects_json` to prompt an LLM with a Jinja template, requiring exactly 7 unique project selections.
- **Fallback mechanisms** ensure 7 projects are always returned, either through deduplication backfilling or defaulting to the top starred repositories.
- Supporting logic resides in [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) for rendering and [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) for LLM interaction and JSON extraction.

## Frequently Asked Questions

### How does the algorithm distinguish between open source contributions and personal projects?

The code classifies repositories based on contributor count. If `fetch_repo_contributors` returns more than one contributor, the repository is tagged as **open_source**; otherwise, it is classified as **self_project**. This distinction helps the LLM evaluate collaborative versus individual engineering capabilities.

### What happens if a candidate has fewer than 7 repositories?

The algorithm attempts to fulfill the quota of exactly 7 projects through its fallback logic. If the LLM returns fewer than 7 unique selections, the system appends additional repositories from the star-sorted original list. If the LLM fails entirely, it defaults to the first 7 repositories from the sorted list, or fewer if the candidate lacks sufficient public repositories.

### Why does the system filter forks with fewer than 5 forks?

The condition `fork && forks_count < 5` identifies trivial forks—repositories that were likely forked for reference but never substantially modified or maintained. By filtering these out, the algorithm focuses on original work or significant forks that demonstrate active development or community maintenance, ensuring the final selection represents meaningful engineering effort.

### Which files control the LLM prompt formatting and response parsing?

The **Jinja template** at `prompts/templates/github_project_selection.jinja` defines the prompt structure sent to the LLM. The `TemplateManager` class in [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) handles template rendering. Response parsing and LLM provider initialization occur in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py), specifically through `extract_json_from_response` and `initialize_llm_provider`.