GitHub Profile Enrichment and Project Selection Algorithm in the Hiring-Agent Repository

The hiring-agent system extracts a candidate's GitHub username, fetches their public profile and repository metadata, filters out trivial forks, classifies remaining projects as open-source or self-authored based on contributor counts, and selects the top seven most-starred repositories to enrich résumé evaluation.

The interviewstreet/hiring-agent repository implements a sophisticated GitHub profile enrichment and project selection algorithm that automatically augments candidate résumés with live data from their GitHub accounts. By parsing varied GitHub URL formats, retrieving detailed user statistics through the GitHub API, and applying intelligent filtering logic, the system surfaces meaningful engineering contributions to support technical hiring decisions.

Username Extraction and Profile Normalization

The pipeline begins in github.py with extract_github_username (lines 16-38), which normalizes diverse input formats—whether a candidate provides https://github.com/username, github.com/username, or just the plain handle—into a clean identifier required for subsequent API calls. This robust parsing ensures the system handles résumé data inconsistencies without manual intervention.

Profile Enrichment via GitHub API

Once normalized, fetch_github_profile constructs the GitHub API endpoint https://api.github.com/users/<username> and invokes the generic _fetch_github_api helper to retrieve comprehensive account data.

Pydantic Model Validation

The API response populates a GitHubProfile Pydantic model (lines 41-70 of github.py) capturing critical fields including name, bio, public_repos, followers, and created_at. This structured validation enforces type safety and ensures consistent data handling throughout the enrichment pipeline.

JSON Serialization

For downstream processing, generate_profile_json (lines 72-78) converts the GitHubProfile instance into a plain dictionary, preparing the data for seamless integration with the scoring system and transformation modules.

Repository Retrieval and Intelligent Filtering

The core selection logic resides in fetch_all_github_repos, which orchestrates a multi-stage project discovery and classification workflow.

Bulk Repository Fetching

The function queries GET /users/<username>/repos with parameters sorting by updated and limiting results to max_repos (default 100). This retrieves the candidate's most recently active repositories while respecting API rate limits.

Fork Filtering Strategy

To eliminate low-signal repositories, the algorithm applies a specific heuristic at lines 34-36: it skips any repository where repo.get("fork") evaluates to true and repo.get("forks_count", 0) < 5. This filter removes personal forks of popular projects that demonstrate minimal independent development while preserving substantive forks that have garnered their own community interest.

Contribution Analysis and Project Classification

For each retained repository, the system performs deep contributor analysis through three sequential operations:

  1. Contributor Retrieval: fetch_repo_contributors calls /repos/<owner>/<repo>/contributors (lines 40-45) to fetch the full contributor list with commit statistics.
  2. Contribution Counting: fetch_contributions_count (lines 87-99) aggregates total repository contributions and isolates the candidate's specific commit count.
  3. Collaborative Classification: The algorithm categorizes projects based on contributor diversity:
    • "open_source": Projects with contributor_count > 1, indicating collaborative development experience
    • "self_project": Single-contributor repositories, indicating solo-authored initiatives

Star-Based Ranking

After processing all repositories, the system constructs a rich project dictionary (lines 51-71) containing metadata including name, description, language, project_type, and nested github_details (stars, forks, topics). The final list sorts descending by github_details.stars (lines 79-81), using community validation as the primary quality signal.

Top Project Selection

The final selection occurs in fetch_and_display_github_info, which truncates the sorted repository list to exactly seven entries:

projects_data = []
for project in projects[:7]:  # Keep only the 7 best-scoring repos

    project_data = { ... }
    projects_data.append(project_data)

This hard limit (implemented at lines 42-55 of github.py) ensures the enrichment process remains focused and avoids overwhelming evaluation workflows with excessive data while providing a comprehensive view of the candidate's best work. The function also generates a summary breakdown (lines 88-92) reporting total project counts and the open-source versus self-project distribution.

Integration with the Hiring Pipeline

The enrichment module integrates seamlessly with the broader evaluation system orchestrated by score.py.

Profile Discovery

The find_profile function (lines 5-11) scans the basics.profiles array of the parsed résumé JSON, identifying entries where the network field equals "Github" (case-sensitive matching) to locate the candidate's GitHub URL.

Orchestration Flow

When a GitHub profile is detected, the main scoring logic invokes the enrichment routine:

github_data = fetch_and_display_github_info(github_profile.url)

This returns a dictionary containing the enriched profile metadata, the filtered projects array (maximum 7), and total_projects count. This payload merges into the candidate evaluation pipeline, providing hiring teams with verified, quantitative data alongside traditional résumé content.

Practical Implementation Examples

Direct API Usage

from github import fetch_and_display_github_info

# Enrich a candidate's résumé

github_url = "https://github.com/exampleUser"
enriched = fetch_and_display_github_info(github_url)

print(f"Candidate: {enriched['profile']['name']}")
print(f"Bio: {enriched['profile']['bio']}")

print("\nTop projects:")
for proj in enriched['projects']:
    print(f"- {proj['name']} ({proj['project_type']}): {proj['github_details']['stars']} stars")

Understanding the Data Structure


# The returned github_data structure contains:

{
    "profile": {
        "name": "Jane Doe",
        "bio": "Full-stack developer...",
        "public_repos": 42,
        "followers": 150
    },
    "projects": [
        {
            "name": "ml-toolkit",
            "project_type": "open_source",  # or "self_project"

            "github_details": {
                "stars": 1200,
                "forks": 150,
                "language": "Python"
            }
        }
    ],
    "total_projects": 7
}

Summary

  • The algorithm normalizes GitHub URLs via extract_github_username in github.py (lines 16-38) to ensure consistent API interaction regardless of input format.
  • Fork filtering eliminates repositories with fewer than 5 forks, removing trivial copies while preserving substantive community forks that demonstrate independent development value.
  • Project classification distinguishes between "open_source" (multi-contributor) and "self_project" (solo-authored) based on contributor count analysis performed by fetch_repo_contributors.
  • Star-based ranking sorts repositories by community popularity (github_details.stars), selecting only the top seven for final enrichment to maintain evaluation focus.
  • Integration occurs through score.py (lines 5-11, 31-33), which automatically detects GitHub profiles in résumé data and triggers the complete enrichment pipeline.

Frequently Asked Questions

How does the algorithm differentiate between open-source and personal projects?

The system classifies a repository as "open_source" when fetch_repo_contributors returns more than one contributor, indicating collaborative development experience. Single-contributor repositories are labeled "self_project", reflecting solo-authored work. This classification occurs in github.py (lines 46-49) and helps hiring teams understand whether the candidate has experience working with distributed teams or primarily develops independently.

Why does the system filter out forks with fewer than 5 forks?

The algorithm skips forks where forks_count < 5 to exclude trivial personal copies of popular repositories that demonstrate minimal independent development effort. This filter (implemented at lines 34-36 of github.py) ensures only substantive forks—those that have attracted their own community interest—surface in the evaluation, providing more relevant signal about the candidate's actual maintenance capabilities and project ownership.

How many projects does the system select for résumé enrichment?

The pipeline selects exactly seven projects, as strictly implemented in fetch_and_display_github_info (lines 42-55). This hard limit balances comprehensive candidate assessment with evaluation efficiency, focusing hiring teams on the candidate's highest-quality repositories sorted by star count in descending order.

What determines which repositories appear in the final selection?

Repository selection follows a strict hierarchy: first, the system filters out trivial forks (fewer than 5 forks), then classifies by contributor type (open-source vs. self-project), and finally sorts by star count in descending order. The top seven repositories from this sorted list comprise the final enrichment payload, ensuring the highest-quality, most-community-validated projects represent the candidate's technical portfolio.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →