# How the Hiring Agent Selects Top GitHub Projects: Criteria and LLM Algorithm

> Discover what criteria the Hiring Agent uses to select top GitHub projects. It leverages an LLM algorithm with a fallback star-count ranking for impressive results.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-03

---

**The Hiring Agent selects top GitHub projects by feeding repository metadata into a Large Language Model (LLM) with a strict system prompt requesting exactly 7 unique, impressive projects, falling back to star-count ranking if the LLM response fails or returns insufficient results.**

The `interviewstreet/hiring-agent` repository automates technical candidate evaluation by analyzing GitHub profiles to identify standout projects. Understanding how this open-source tool selects top GitHub projects reveals a hybrid approach combining quantitative repository metrics with AI-driven judgment.

## Repository Metadata Collection

The selection process begins with comprehensive data gathering for each repository associated with a candidate's profile. In [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 34-62), the `fetch_all_github_repos` function collects structured metadata including:

- **Identity data**: Repository name, description, URL, live URL, and primary language (stored as `technologies`)
- **Classification**: Distinction between **open-source** projects (multiple contributors) and **self-projects** (single contributor)
- **Engagement metrics**: Star count, fork count, issue count, topics, repository size, and creation/update dates (the `github_details` block)
- **Contribution analytics**: Total contributor count, the author's specific commit count, and overall commit totals

This data structure provides the raw material for both the LLM evaluation and the fallback ranking mechanism.

## LLM-Powered Selection Criteria

Rather than relying solely on star counts, the Hiring Agent employs an LLM to interpret repository quality and significance. The core logic resides in `generate_projects_json` within [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 78-85), where collected repository data is rendered into a prompt using the `github_project_selection.jinja` template.

### The System Prompt Requirements

The LLM receives explicit instructions through a system message that constrains the selection behavior:

> "You are an expert technical recruiter analyzing GitHub repositories to identify the most impressive projects. **CRITICAL:** You must select exactly **7** UNIQUE projects – no duplicates allowed. Each project must be different from the others."

This prompt ensures the model prioritizes **project uniqueness** and **impressive qualities** over raw popularity metrics, allowing the LLM to evaluate factors like code complexity, documentation quality, and technical sophistication that star counts alone cannot capture.

### Parsing and Deduplication Logic

After the LLM returns a JSON list of selected projects, the Agent implements rigorous validation in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 90-106):

1. **JSON parsing**: The response is parsed to extract the project list
2. **Deduplication**: Any duplicate repository names are removed
3. **Quota verification**: The system ensures exactly 7 distinct projects are present
4. **Gap filling**: If the LLM returns fewer than 7 unique entries, the Agent automatically fills remaining slots with the highest-ranked repositories from the original list, sorted by star count in descending order

This hybrid approach ensures that LLM judgment takes precedence while maintaining the required output quantity through quantitative metrics when necessary.

## Error Handling and Fallback Mechanisms

The implementation includes robust error handling for LLM failures. In [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 124-136), the system catches JSON decode errors or API call failures and executes a deterministic fallback: returning the first 7 repositories from the star-sorted list as a safety net.

This dual-path architecture ensures reliability while preserving the LLM's ability to identify noteworthy projects that might have lower star counts but high technical merit.

## Source Code Implementation

The selection algorithm spans several key files in the repository:

| File | Function | Purpose |
|------|----------|---------|
| [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) | `generate_projects_json` (lines 34-106) | Orchestrates data collection, LLM prompting, and result parsing |
| [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) | Error handling block (lines 124-136) | Implements the star-count fallback for failed LLM calls |
| `prompts/templates/github_project_selection.jinja` | Template rendering | Structures repository data for LLM consumption |
| [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) | `get_template()` | Loads and renders Jinja templates |
| [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) | `GitHubProfile` | Defines the data model for user profiles accompanying project lists |

## Practical Usage Examples

To retrieve the top 7 projects for a GitHub profile using the Hiring Agent:

```python
from github import fetch_and_display_github_info

result = fetch_and_display_github_info("https://github.com/your-username")
print("Top projects:")
for proj in result["projects"]:
    print(f"- {proj['name']} ({proj['github_url']})")

```

For testing or custom implementations, you can invoke the selection logic directly:

```python
from github import fetch_all_github_repos, generate_projects_json

repos = fetch_all_github_repos("https://github.com/your-username")
top_projects = generate_projects_json(repos)   # Returns a list of up to 7 selected projects

```

## Summary

- The Hiring Agent combines **quantitative metrics** (stars, forks, contributions) with **LLM interpretation** to identify impressive projects
- The system strictly enforces **7 unique projects** through prompt engineering and programmatic deduplication
- **Fallback mechanisms** ensure reliability by defaulting to star-count rankings when LLM calls fail or return insufficient results
- Implementation centers on [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), specifically the `generate_projects_json` function and its error handling blocks
- The `github_project_selection.jinja` template structures the data sent to the LLM, while [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) provides the underlying data schema

## Frequently Asked Questions

### How does the Hiring Agent handle repositories with low star counts but high code quality?

The LLM evaluation layer allows the system to identify technically sophisticated projects regardless of popularity. While the fallback mechanism relies on star counts, the primary selection path enables the LLM to recognize factors like architectural complexity, testing coverage, and innovative implementations that pure metrics might miss.

### What happens if the LLM returns duplicate project recommendations?

The Agent implements explicit deduplication logic in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 90-106) that removes duplicate repository names before finalizing the list. If deduplication reduces the list below 7 projects, the system automatically fills the remaining slots with the next highest-ranked repositories by star count from the original candidate pool.

### Can the number of selected projects be configured instead of fixed at 7?

According to the source code in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 78-85), the requirement for exactly 7 projects is hard-coded in the system message sent to the LLM. Changing this would require modifying the system prompt template in `prompts/templates/github_project_selection.jinja` and updating the validation logic that enforces the quota during result parsing.

### Which repository metrics most heavily influence the LLM's selection decision?

While the LLM receives comprehensive metadata including stars, forks, contributor counts, and technologies, the system prompt specifically instructs it to identify "most impressive projects" based on the complete context. The model evaluates the combination of engagement metrics, contributor diversity (open-source vs. self-project classification), and repository metadata to determine technical merit.