How GitHub Projects Are Classified by the LLM in Hiring Agent

Hiring Agent classifies GitHub repositories by contributor count, tagging repos with multiple contributors as "open_source" and solo repos as "self_project", then passes these pre-computed tags to an LLM to select the top 7 projects for candidate evaluation.

The interviewstreet/hiring-agent repository automates technical candidate screening by analyzing GitHub profiles to identify meaningful work. Before the LLM ranks a developer's portfolio, the system applies a simple heuristic to classify GitHub projects, distinguishing collaborative open-source contributions from personal solo efforts. This classification occurs in github.py and directly influences how the model prioritizes repositories during final selection.

The Classification Heuristic in github.py

The classification logic resides in the fetch_all_github_repos function within github.py. For every repository belonging to a candidate, the system calculates the total number of contributors and applies a binary taxonomy based on collaboration patterns.

How Contributor Count Determines Project Type

The code explicitly counts contributors and assigns a string label according to the following rule found at lines 46-48:

contributor_count = len(contributors_data)
project_type = "open_source" if contributor_count > 1 else "self_project"

This produces two distinct categories:

  • open_source – Repositories with more than one contributor, indicating collaborative, community-driven development.
  • self_project – Repositories where the candidate is the sole contributor, representing personal or experimental work.

These tags are stored in the repository dictionary under the key project_type, creating a structured dataset that the downstream LLM consumes without modifying the original classification.

From Raw Data to LLM Input

After classification, the system aggregates repository metadata into a JSON payload. The generate_projects_json function prepares this data, ensuring each entry retains its project_type tag alongside other repository metrics like stars, forks, and primary language.

The project_type Field

The project_type field serves as a categorical signal that helps the LLM understand the social context of each repository. According to the source code analysis, this field is populated during the initial GitHub API fetch and remains immutable throughout the selection process. The LLM receives this data through the Jinja template located at prompts/templates/github_project_selection.jinja, which formats the repository list for the model's consumption.

Building the Selection Payload

The final payload sent to the LLM includes exactly 7 unique projects selected from the classified pool. While the system prompt instructs the model to choose exactly seven repositories, it does not grant the LLM permission to reclassify the projects—the open_source and self_project labels remain fixed based on the original contributor count analysis.

How the LLM Uses Classification Tags

The LLM leverages the pre-computed project_type values as ranking signals rather than classification inputs. When evaluating a candidate's GitHub history, the model can weight collaborative projects differently from solo efforts, though the specific weighting depends on the prompt defined in github.py. The llm_utils.py module handles the response parsing, ensuring the final selection returned to the user contains valid JSON with the original classification tags intact.

Implementation Example

To classify a user's repositories programmatically, invoke the fetch function and inspect the resulting project_type values:

from github import fetch_all_github_repos

# Fetch and classify all repos for a candidate

repos = fetch_all_github_repos("https://github.com/example_user")
for r in repos:
    print(f"{r['name']}: {r['project_type']} (contributors={r['contributor_count']})")

To generate the LLM-curated selection that respects these classifications:

from github import generate_projects_json

# projects is the list returned by fetch_all_github_repos

top_projects = generate_projects_json(projects)  # Returns exactly 7 unique projects

for p in top_projects:
    print(p["name"], "→", p["project_type"])

Summary

  • Hiring Agent classifies GitHub projects in github.py based solely on contributor count.
  • Repositories with multiple contributors receive the open_source tag, while solo repos are labeled self_project.
  • The classification is stored under the project_type key in the repository dictionary.
  • The LLM receives these pre-computed tags and selects exactly 7 unique projects without altering the original classification.
  • Supporting files include prompts/templates/github_project_selection.jinja for prompt templating and llm_utils.py for response parsing.

Frequently Asked Questions

How does Hiring Agent determine if a GitHub project is open source?

Hiring Agent determines a project's status by counting contributors in the fetch_all_github_repos function. If a repository has more than one contributor, it is classified as "open_source"; otherwise, it is tagged as "self_project". This logic executes in github.py before any LLM processing occurs.

Can the LLM override the project type classification?

No, the LLM cannot override the classification. The model receives the pre-computed project_type values as part of the input payload and uses them as static attributes when ranking repositories. The classification remains immutable from the initial GitHub API fetch through the final selection of seven projects.

Where is the project classification logic stored in the codebase?

The classification logic is implemented in github.py, specifically within the fetch_all_github_repos function around lines 46-48. The system prompt template used to present these classifications to the LLM resides in prompts/templates/github_project_selection.jinja, while models.py defines the data structures that hold the profile information.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →