# Algorithm for Selecting the Top 7 GitHub Repositories in the Hiring Agent

> Discover the algorithm for selecting top 7 GitHub repos in hiring agent. Learn how deterministic sorting and LLM curation ensure unique, high-quality results for hiring.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: algorithm-analysis
- Published: 2026-06-27

---

**The interviewstreet/hiring-agent repository implements a hybrid selection pipeline that combines deterministic star-based sorting with LLM-driven curation to return exactly seven unique GitHub repositories, using popularity-ranked fallbacks when the AI response is incomplete or invalid.**

The hiring-agent project automates technical candidate evaluation by analyzing GitHub profiles to identify the most impressive work samples. Understanding the algorithm for selecting the top 7 GitHub repositories reveals how the system balances quantitative metrics with AI-powered judgment to surface high-quality projects while handling edge cases gracefully.

## Step 1: Repository Data Collection and Contributor Analysis

### Fetching Public Repositories and Metadata

The selection process begins in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) where `fetch_all_github_repos` retrieves every public repository for a given username—defaulting to the most recently updated 100 projects. For each repository, the code invokes `fetch_repo_contributors` to calculate three critical metrics: `contributor_count` (total contributors), `author_commit_count` (commits by the repository owner), and `total_commit_count` (aggregate contributions across all contributors).

### Classifying Project Types and Filtering Forks

The algorithm distinguishes between collaborative and solo work by setting a `project_type` flag to **open_source** when `contributor_count` exceeds one, otherwise marking it as **self_project**. To eliminate low-value duplicates, the system automatically excludes forked repositories with fewer than five forks, as these rarely represent substantial independent development. All metadata is stored in a structured dictionary (lines 50-78 of [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)).

## Step 2: Deterministic Pre-Ranking by Popularity

Before AI involvement, the pipeline establishes a deterministic baseline by sorting the filtered list in descending order of GitHub stars:

```python
projects.sort(key=lambda x: x["github_details"]["stars"], reverse=True)

```

This star-based ordering creates a reliable "best-by-popularity" sequence that serves as the fallback reference throughout the selection process.

## Step 3: LLM-Driven Selection and Deduplication

### AI-Powered Project Curation

The `generate_projects_json` function serializes the star-sorted repository list into JSON and transmits it to an LLM provider instantiated via `initialize_llm_provider`. The system prompt (lines 82-84 of [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)) explicitly instructs the model to "select exactly 7 UNIQUE projects – no duplicates allowed." The response is parsed using `extract_json_from_response` (defined in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)) to extract the AI's curated selections.

### Fallback Mechanisms for Incomplete Results

The system implements robust deduplication using a `seen_names` set to remove duplicate project names from the LLM output. If the model returns fewer than seven unique repositories, the algorithm walks the original star-sorted list and appends the highest-ranked missing entries until the count reaches exactly seven (lines 15-22). When the LLM response cannot be parsed or contains invalid JSON, the code logs an error and immediately returns the first seven repositories from the popularity-sorted list (lines 35-36).

## Code Implementation Examples

```python

# Fetch and display the top 7 repositories for evaluation

from github import fetch_and_display_github_info

result = fetch_and_display_github_info("https://github.com/example_user")

# result["projects"] contains exactly 7 project objects

print(result["projects"])

```

```python

# Directly invoke the selection pipeline components

from github import generate_projects_json, fetch_all_github_repos

# Fetch raw repos (pre-sorted by stars)

raw_projects = fetch_all_github_repos("https://github.com/example_user")

# Let the LLM select the top 7 with deduplication

top7 = generate_projects_json(raw_projects)

```

## Summary

- The algorithm for selecting the top 7 GitHub repositories combines **quantitative filtering** (removing low-fork clones) with **deterministic star-sorting** and **LLM-driven curation**.
- Repository metadata—including contributor counts and commit statistics—is aggregated in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 50-78) to classify projects as open-source or self-directed.
- The LLM receives explicit instructions to choose exactly seven unique projects, with a **fallback mechanism** that supplements incomplete AI responses using the star-ranked list.
- All parsing and validation logic resides in [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py), while [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) renders the selection prompts.

## Frequently Asked Questions

### How does the hiring-agent handle candidates with fewer than 7 repositories?

If the candidate has fewer than seven public repositories meeting the criteria, the algorithm returns all available valid repositories. The fallback mechanism only activates when the LLM returns fewer than seven selections from a larger pool, not when the source data itself is limited.

### Why does the algorithm ignore forks with fewer than 5 forks?

Forked repositories with minimal fork counts are excluded because they typically represent minor contributions or untouched copies of existing projects rather than substantial original development. This filtering ensures the evaluation focuses on meaningful work.

### What happens if the LLM returns duplicate repository names?

The deduplication logic uses a `seen_names` set to automatically remove duplicates before finalizing the list. If deduplication reduces the count below seven, the system pulls additional repositories from the star-sorted fallback list to maintain the required quota.

### Where is the project classification logic (open_source vs self_project) implemented?

The classification occurs in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 50-78) where the code compares `contributor_count` against the threshold of one. Repositories with multiple contributors are flagged as **open_source**, while solo efforts are marked as **self_project**, enabling nuanced evaluation of collaborative versus independent development skills.