How GitHub Projects Are Classified as Open Source or Self-Hosted in Hiring Agent

The Hiring Agent classifies repositories as "open_source" or "self_project" based strictly on contributor count: repositories with two or more distinct contributors are labeled open source, while single-contributor repositories are classified as self-hosted personal projects.

The interviewstreet/hiring-agent repository automatically categorizes a candidate's GitHub repositories during the profile retrieval phase. This binary classification determines whether a project represents collaborative community work or personal experimental code, directly influencing how the LLM evaluates the candidate's technical experience and portfolio strength.

How the Classification Algorithm Works

The classification occurs within the fetch_all_github_repos function in github.py immediately after retrieving repository data from the GitHub API.

The Contributor Count Threshold

The system queries the GitHub API endpoint /repos/{owner}/{repo}/contributors to fetch the list of distinct contributors for each repository. It then applies a simple threshold-based rule:

  • Open Source: contributor_count > 1 (two or more distinct contributors)
  • Self‑Hosted: contributor_count <= 1 (only the repository owner)

This logic reflects the assumption that repositories with multiple contributors indicate shared, community-maintained code, while single-contributor repositories typically represent personal experiments, tutorials, or prototypes.

Implementation in fetch_all_github_repos

In github.py, lines 46-48 implement the core classification logic:


# github.py – fetch_all_github_repos

contributor_count = len(contributors_data)

# A project is considered open‑source if more than one person contributed.

project_type = (
    "open_source" if contributor_count > 1 else "self_project"
)

The project_type field is then stored in the repository's JSON record and used throughout the evaluation pipeline.

Aggregating Project Statistics

After classifying individual repositories, the system aggregates counts to provide high-level metrics. Lines 82-87 in github.py calculate these totals:

open_source_count = sum(
    1 for p in projects if p["project_type"] == "open_source"
)

self_project_count = sum(
    1 for p in projects if p["project_type"] == "self_project"
)

The console displays these totals during execution:


📊 Project classification: 3 open source, 5 self projects

Working with Classified Repository Data

The project_type field enables downstream filtering and analysis. Below are practical patterns for interacting with this classification data.

Fetching and Inspecting Classifications

To retrieve repositories and inspect their inferred types:

from github import fetch_all_github_repos

# Example GitHub profile URL (candidate’s public profile)

profile_url = "https://github.com/jane-doe"

# Retrieve all repositories (default max 100)

repos = fetch_all_github_repos(profile_url)

# Show each repo with its inferred type

for repo in repos:
    print(f"{repo['name']}: {repo['project_type']}")

Sample output:


awesome‑api: open_source
my‑portfolio: self_project
data‑visualizer: open_source
toy‑todo‑app: self_project

Filtering for Open Source Projects

To isolate only community-driven projects for focused analysis:

open_source_repos = [
    r for r in repos if r["project_type"] == "open_source"
]

print("Open‑source projects:")
for r in open_source_repos:
    print(f"- {r['name']} ({r['github_url']})")

Calculating Aggregate Statistics

For direct access to classification counts:

open_cnt = sum(1 for r in repos if r["project_type"] == "open_source")
self_cnt = sum(1 for r in repos if r["project_type"] == "self_project")

print(f"Open‑source: {open_cnt}, Self‑hosted: {self_cnt}")

Impact on LLM Evaluation

The classification directly influences the Hiring Agent's evaluation strategy. The github_project_selection.jinja template explicitly instructs the LLM to prioritize open-source work, using the project_type field to distinguish repository significance:

"Prioritize contributions to popular, open‑source projects …"

Similarly, the resume_evaluation_criteria.jinja template applies different scoring weights based on whether contributions are to collaborative projects versus personal repositories. The GitHubProfile dataclass in models.py structures this data for consumption by the evaluation pipeline.

Summary

  • The Hiring Agent automatically classifies GitHub repositories using contributor count data from the GitHub API.
  • Open source classification requires two or more distinct contributors, while self‑hosted (or self_project) indicates single-contributor repositories.
  • The logic resides in github.py within the fetch_all_github_repos function, specifically lines 46-48 for classification and lines 82-87 for aggregation.
  • The project_type field drives LLM prompt templates including github_project_selection.jinja and resume_evaluation_criteria.jinja, causing collaborative projects to receive higher evaluation priority.
  • The GitHubProfile model in models.py structures this data for the evaluation pipeline.

Frequently Asked Questions

How does the Hiring Agent determine if a GitHub repository is open source?

The Agent queries the GitHub API's /repos/{owner}/{repo}/contributors endpoint and counts the distinct contributors. If contributor_count exceeds one, the repository is classified as "open_source"; otherwise, it is labeled "self_project". This logic is implemented in github.py lines 46-48.

Why does the classification rely on contributor count rather than repository visibility?

Contributor count serves as a proxy for community collaboration. A repository with multiple distinct contributors indicates shared maintenance and community adoption, whereas a single-contributor repository typically represents personal experiments or work-in-progress prototypes. This distinction helps the LLM differentiate between collaborative engineering experience and solo projects.

Where is the project classification used in the evaluation pipeline?

The project_type field feeds into the github_project_selection.jinja prompt template, which explicitly directs the LLM to prioritize open-source contributions over self-hosted projects. The classification also influences scoring rules in resume_evaluation_criteria.jinja and is stored in the GitHubProfile dataclass defined in models.py.

Can the classification threshold be modified to require more contributors?

The current implementation in github.py uses a hardcoded threshold of contributor_count > 1 on line 47. To require additional contributors for open-source classification (for example, three or more), you would modify this comparison operator in the fetch_all_github_repos function and redeploy the service.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →