# How GitHub Projects Are Classified by the LLM in Hiring Agent

> Discover how Hiring Agent uses an LLM to classify GitHub Projects based on contributor count differentiating between open source and solo endeavors for candidate evaluation. Learn more now.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: getting-started
- Published: 2026-06-29

---

**Hiring Agent classifies GitHub repositories by contributor count, tagging repos with multiple contributors as "open_source" and solo repos as "self_project", then passes these pre-computed tags to an LLM to select the top 7 projects for candidate evaluation.**

The `interviewstreet/hiring-agent` repository automates technical candidate screening by analyzing GitHub profiles to identify meaningful work. Before the LLM ranks a developer's portfolio, the system applies a simple heuristic to classify GitHub projects, distinguishing collaborative open-source contributions from personal solo efforts. This classification occurs in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) and directly influences how the model prioritizes repositories during final selection.

## The Classification Heuristic in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)

The classification logic resides in the `fetch_all_github_repos` function within [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py). For every repository belonging to a candidate, the system calculates the total number of contributors and applies a binary taxonomy based on collaboration patterns.

### How Contributor Count Determines Project Type

The code explicitly counts contributors and assigns a string label according to the following rule found at lines 46-48:

```python
contributor_count = len(contributors_data)
project_type = "open_source" if contributor_count > 1 else "self_project"

```

This produces two distinct categories:

- **`open_source`** – Repositories with **more than one contributor**, indicating collaborative, community-driven development.
- **`self_project`** – Repositories where the candidate is the sole contributor, representing personal or experimental work.

These tags are stored in the repository dictionary under the key **`project_type`**, creating a structured dataset that the downstream LLM consumes without modifying the original classification.

## From Raw Data to LLM Input

After classification, the system aggregates repository metadata into a JSON payload. The `generate_projects_json` function prepares this data, ensuring each entry retains its `project_type` tag alongside other repository metrics like stars, forks, and primary language.

### The `project_type` Field

The `project_type` field serves as a categorical signal that helps the LLM understand the social context of each repository. According to the source code analysis, this field is populated during the initial GitHub API fetch and remains immutable throughout the selection process. The LLM receives this data through the Jinja template located at `prompts/templates/github_project_selection.jinja`, which formats the repository list for the model's consumption.

### Building the Selection Payload

The final payload sent to the LLM includes exactly 7 unique projects selected from the classified pool. While the system prompt instructs the model to choose exactly seven repositories, it does not grant the LLM permission to reclassify the projects—the `open_source` and `self_project` labels remain fixed based on the original contributor count analysis.

## How the LLM Uses Classification Tags

The LLM leverages the pre-computed `project_type` values as ranking signals rather than classification inputs. When evaluating a candidate's GitHub history, the model can weight collaborative projects differently from solo efforts, though the specific weighting depends on the prompt defined in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py). The [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) module handles the response parsing, ensuring the final selection returned to the user contains valid JSON with the original classification tags intact.

## Implementation Example

To classify a user's repositories programmatically, invoke the fetch function and inspect the resulting `project_type` values:

```python
from github import fetch_all_github_repos

# Fetch and classify all repos for a candidate

repos = fetch_all_github_repos("https://github.com/example_user")
for r in repos:
    print(f"{r['name']}: {r['project_type']} (contributors={r['contributor_count']})")

```

To generate the LLM-curated selection that respects these classifications:

```python
from github import generate_projects_json

# projects is the list returned by fetch_all_github_repos

top_projects = generate_projects_json(projects)  # Returns exactly 7 unique projects

for p in top_projects:
    print(p["name"], "→", p["project_type"])

```

## Summary

- Hiring Agent classifies GitHub projects in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) based solely on contributor count.
- Repositories with multiple contributors receive the **`open_source`** tag, while solo repos are labeled **`self_project`**.
- The classification is stored under the **`project_type`** key in the repository dictionary.
- The LLM receives these pre-computed tags and selects exactly **7 unique projects** without altering the original classification.
- Supporting files include `prompts/templates/github_project_selection.jinja` for prompt templating and [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) for response parsing.

## Frequently Asked Questions

### How does Hiring Agent determine if a GitHub project is open source?

Hiring Agent determines a project's status by counting contributors in the `fetch_all_github_repos` function. If a repository has more than one contributor, it is classified as `"open_source"`; otherwise, it is tagged as `"self_project"`. This logic executes in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) before any LLM processing occurs.

### Can the LLM override the project type classification?

No, the LLM cannot override the classification. The model receives the pre-computed `project_type` values as part of the input payload and uses them as static attributes when ranking repositories. The classification remains immutable from the initial GitHub API fetch through the final selection of seven projects.

### Where is the project classification logic stored in the codebase?

The classification logic is implemented in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), specifically within the `fetch_all_github_repos` function around lines 46-48. The system prompt template used to present these classifications to the LLM resides in `prompts/templates/github_project_selection.jinja`, while [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) defines the data structures that hold the profile information.