# How the GitHub Project Selection Algorithm Chooses the Top 7 Projects in Hiring-Agent

> Discover how the GitHub project selection algorithm picks top projects for the hiring-agent using a 3-phase pipeline, LLM prompting, and commit analysis for precisely 7 unique repositories.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-07-05

---

**The GitHub project selection algorithm uses a three-phase pipeline that fetches all public repositories, sorts them by star count, and prompts an LLM to select exactly 7 unique projects where the candidate has authored at least 4 commits, with automatic fallback mechanisms if the LLM returns fewer results.**

The interviewstreet/hiring-agent repository implements a deterministic hybrid approach to surface the most impressive repositories from a candidate's GitHub profile. This GitHub project selection algorithm combines automated data preprocessing with constrained LLM reasoning to ensure consistent output quality while handling edge cases gracefully.

## Overview of the Three-Phase Selection Process

The algorithm operates through distinct phases that progress from data collection to final curation, as implemented in the core [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) module.

### Phase 1: Data Aggregation and Star-Based Sorting

The process begins with `fetch_all_github_repos` in `github.py:L18-L84`, which retrieves all public repositories for a candidate via the GitHub API. Each repository object is enriched with contribution statistics including `author_commit_count` (the candidate's personal commits), `total_commit_count`, and star count.

After fetching, the system sorts the entire collection descending by `github_details.stars` at lines 80-85:

```python

# Sorting logic from github.py

projects.sort(key=lambda x: x.get('github_details', {}).get('stars', 0), reverse=True)

```

This star-based ranking ensures that the most popular repositories are considered first during downstream selection.

### Phase 2: LLM Evaluation Against Strict Criteria

The sorted repository list is converted to JSON via `generate_projects_json` and injected into the `prompts/templates/github_project_selection.jinja` template. This template encodes hard-coded selection rules including:

- **Author commit threshold**: Minimum 4 commits by the candidate
- **Exact count requirement**: Precisely 7 unique projects
- **Priority weighting**: Preference for high-commit-count and popular open-source contributions

The system message in `github.py:L82-L84` reinforces these constraints:

```python
system_msg = {
    "role": "system",
    "content": (
        "You are an expert technical recruiter analyzing GitHub repositories to identify "
        "the most impressive projects. CRITICAL: You must select exactly 7 UNIQUE projects "
        "- no duplicates allowed. Each project must be different from the others."
    ),
}

```

The LLM receives the rendered template at lines 61-89 and returns a JSON array of selected projects.

### Phase 3: Deduplication and Fallback Handling

The post-processing logic in `github.py:L94-L131` and `github.py:L140-L166` handles the LLM response through several defensive steps:

1. **Parsing**: Extracts the JSON array from `response["message"]["content"]`
2. **Deduplication**: Tracks `seen_names` to ensure uniqueness, keeping only the first occurrence of each project name
3. **Fill-up logic**: If fewer than 7 unique projects remain, the algorithm pulls additional entries from the pre-sorted `projects_data` list until the quota is met
4. **Safe fallback**: If parsing fails entirely, `github.py:L190-L216` returns the first 7 entries from the star-sorted list

## Deep Dive into the Selection Criteria

The algorithm enforces specific technical thresholds that balance contribution depth with project popularity.

### The Author Commit Threshold

The Jinja template (lines 44-48 of `github_project_selection.jinja`) requires candidates to have substantive involvement, defined as `author_commit_count >= 4`. This filter eliminates forked repositories or minor contributions, ensuring selected projects represent meaningful development work.

### The Jinja Template Rules

The template file located at `prompts/templates/github_project_selection.jinja` serves as the authoritative rule engine. Beyond the commit threshold, it instructs the LLM to:

- Prioritize projects with high commit counts and star ratings
- Avoid selecting multiple repositories from the same organization unless distinctly different
- Ensure the final output contains exactly 7 unique project names

## Implementation Details and Code Examples

Developers can interact with the selection algorithm through the main entry point:

```python
from github import fetch_and_display_github_info

# Fetch and select top 7 projects for a candidate

result = fetch_and_display_github_info("https://github.com/candidate-username")
projects = result["projects"]  # JSON array of 7 selected repositories

```

### Constructing the LLM Prompt

The prompt construction pipeline renders repository data with specific instructions:

```python

# Excerpt from github.py showing prompt preparation

projects_json = generate_projects_json(projects_data)
prompt = template_manager.render(
    "github_project_selection.jinja",
    projects=projects_json,
    candidate_name=candidate_name
)

```

### Handling the Response and Fallback Logic

The robust fallback mechanism ensures 7 projects are always returned:

```python

# From github.py:L140-L166 - deduplication and fill-up

unique_projects = []
seen_names = set()

for project in llm_selected_projects:
    name = project.get("name")
    if name not in seen_names:
        unique_projects.append(project)
        seen_names.add(name)

# Fill up to 7 if needed

if len(unique_projects) < 7:
    for project in projects_data:
        if project.get("name") not in seen_names:
            unique_projects.append(project)
            if len(unique_projects) == 7:
                break

```

## Summary

- The GitHub project selection algorithm in interviewstreet/hiring-agent combines deterministic preprocessing with LLM-based curation to identify the top 7 repositories.
- **Phase 1** fetches and sorts all public repositories by star count in `github.py:L18-L84`.
- **Phase 2** uses the `github_project_selection.jinja` template to enforce strict criteria including a minimum 4 author commits and exactly 7 unique selections.
- **Phase 3** implements defensive post-processing with deduplication, fill-up logic, and a safe fallback to the top 7 starred projects if the LLM response fails.

## Frequently Asked Questions

### What happens if a candidate has fewer than 7 qualifying repositories?

The algorithm first attempts to secure 7 unique projects from the LLM response. If the candidate lacks sufficient repositories meeting the author-commit threshold (≥4 commits), the fill-up logic in `github.py:L140-L166` pulls additional projects from the star-sorted list regardless of commit count, ensuring the output always contains up to 7 entries.

### Why does the algorithm require exactly 4 author commits?

The 4-commit threshold defined in `github_project_selection.jinja` serves as a heuristic for meaningful contribution. According to the template logic, this minimum ensures selected projects represent substantial development work rather than minor documentation fixes or automated forks, filtering out superficial contributions while maintaining sufficient portfolio breadth.

### How does the system handle duplicate project names from the LLM?

The post-processing logic in `github.py:L94-L131` maintains a `seen_names` set to track processed projects. When parsing the LLM's JSON response, the code skips any project whose name already exists in the set, preserving only the first occurrence. This guarantees uniqueness regardless of whether the LLM accidentally includes duplicates.

### What is the fallback mechanism if the LLM fails to return valid JSON?

If the LLM response cannot be parsed as JSON, the code at `github.py:L190-L216` executes a safe fallback that immediately returns the first 7 entries from the `projects_data` list (already sorted by stars). This deterministic path ensures the hiring pipeline continues functioning even during LLM service interruptions or malformed responses.