# GitHub Profile Enrichment and Project Selection Algorithm in the Hiring-Agent Repository

> Discover the GitHub profile enrichment and project selection algorithm used by the hiring-agent repository. Learn how it identifies and ranks candidate projects for résumé evaluation.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-06

---

**The hiring-agent system extracts a candidate's GitHub username, fetches their public profile and repository metadata, filters out trivial forks, classifies remaining projects as open-source or self-authored based on contributor counts, and selects the top seven most-starred repositories to enrich résumé evaluation.**

The `interviewstreet/hiring-agent` repository implements a sophisticated GitHub profile enrichment and project selection algorithm that automatically augments candidate résumés with live data from their GitHub accounts. By parsing varied GitHub URL formats, retrieving detailed user statistics through the GitHub API, and applying intelligent filtering logic, the system surfaces meaningful engineering contributions to support technical hiring decisions.

## Username Extraction and Profile Normalization

The pipeline begins in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) with `extract_github_username` (lines 16-38), which normalizes diverse input formats—whether a candidate provides `https://github.com/username`, `github.com/username`, or just the plain handle—into a clean identifier required for subsequent API calls. This robust parsing ensures the system handles résumé data inconsistencies without manual intervention.

## Profile Enrichment via GitHub API

Once normalized, `fetch_github_profile` constructs the GitHub API endpoint `https://api.github.com/users/<username>` and invokes the generic `_fetch_github_api` helper to retrieve comprehensive account data.

### Pydantic Model Validation

The API response populates a `GitHubProfile` Pydantic model (lines 41-70 of [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)) capturing critical fields including `name`, `bio`, `public_repos`, `followers`, and `created_at`. This structured validation enforces type safety and ensures consistent data handling throughout the enrichment pipeline.

### JSON Serialization

For downstream processing, `generate_profile_json` (lines 72-78) converts the `GitHubProfile` instance into a plain dictionary, preparing the data for seamless integration with the scoring system and transformation modules.

## Repository Retrieval and Intelligent Filtering

The core selection logic resides in `fetch_all_github_repos`, which orchestrates a multi-stage project discovery and classification workflow.

### Bulk Repository Fetching

The function queries `GET /users/<username>/repos` with parameters sorting by `updated` and limiting results to `max_repos` (default 100). This retrieves the candidate's most recently active repositories while respecting API rate limits.

### Fork Filtering Strategy

To eliminate low-signal repositories, the algorithm applies a specific heuristic at lines 34-36: it skips any repository where `repo.get("fork")` evaluates to true **and** `repo.get("forks_count", 0) < 5`. This filter removes personal forks of popular projects that demonstrate minimal independent development while preserving substantive forks that have garnered their own community interest.

### Contribution Analysis and Project Classification

For each retained repository, the system performs deep contributor analysis through three sequential operations:

1. **Contributor Retrieval**: `fetch_repo_contributors` calls `/repos/<owner>/<repo>/contributors` (lines 40-45) to fetch the full contributor list with commit statistics.
2. **Contribution Counting**: `fetch_contributions_count` (lines 87-99) aggregates total repository contributions and isolates the candidate's specific commit count.
3. **Collaborative Classification**: The algorithm categorizes projects based on contributor diversity:
   - **"open_source"**: Projects with `contributor_count > 1`, indicating collaborative development experience
   - **"self_project"**: Single-contributor repositories, indicating solo-authored initiatives

### Star-Based Ranking

After processing all repositories, the system constructs a rich project dictionary (lines 51-71) containing metadata including `name`, `description`, `language`, `project_type`, and nested `github_details` (stars, forks, topics). The final list sorts descending by `github_details.stars` (lines 79-81), using community validation as the primary quality signal.

## Top Project Selection

The final selection occurs in `fetch_and_display_github_info`, which truncates the sorted repository list to exactly seven entries:

```python
projects_data = []
for project in projects[:7]:  # Keep only the 7 best-scoring repos

    project_data = { ... }
    projects_data.append(project_data)

```

This hard limit (implemented at lines 42-55 of [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)) ensures the enrichment process remains focused and avoids overwhelming evaluation workflows with excessive data while providing a comprehensive view of the candidate's best work. The function also generates a summary breakdown (lines 88-92) reporting total project counts and the open-source versus self-project distribution.

## Integration with the Hiring Pipeline

The enrichment module integrates seamlessly with the broader evaluation system orchestrated by [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).

### Profile Discovery

The `find_profile` function (lines 5-11) scans the `basics.profiles` array of the parsed résumé JSON, identifying entries where the network field equals `"Github"` (case-sensitive matching) to locate the candidate's GitHub URL.

### Orchestration Flow

When a GitHub profile is detected, the main scoring logic invokes the enrichment routine:

```python
github_data = fetch_and_display_github_info(github_profile.url)

```

This returns a dictionary containing the enriched `profile` metadata, the filtered `projects` array (maximum 7), and `total_projects` count. This payload merges into the candidate evaluation pipeline, providing hiring teams with verified, quantitative data alongside traditional résumé content.

## Practical Implementation Examples

### Direct API Usage

```python
from github import fetch_and_display_github_info

# Enrich a candidate's résumé

github_url = "https://github.com/exampleUser"
enriched = fetch_and_display_github_info(github_url)

print(f"Candidate: {enriched['profile']['name']}")
print(f"Bio: {enriched['profile']['bio']}")

print("\nTop projects:")
for proj in enriched['projects']:
    print(f"- {proj['name']} ({proj['project_type']}): {proj['github_details']['stars']} stars")

```

### Understanding the Data Structure

```python

# The returned github_data structure contains:

{
    "profile": {
        "name": "Jane Doe",
        "bio": "Full-stack developer...",
        "public_repos": 42,
        "followers": 150
    },
    "projects": [
        {
            "name": "ml-toolkit",
            "project_type": "open_source",  # or "self_project"

            "github_details": {
                "stars": 1200,
                "forks": 150,
                "language": "Python"
            }
        }
    ],
    "total_projects": 7
}

```

## Summary

- The algorithm normalizes GitHub URLs via `extract_github_username` in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 16-38) to ensure consistent API interaction regardless of input format.
- **Fork filtering** eliminates repositories with fewer than 5 forks, removing trivial copies while preserving substantive community forks that demonstrate independent development value.
- **Project classification** distinguishes between `"open_source"` (multi-contributor) and `"self_project"` (solo-authored) based on contributor count analysis performed by `fetch_repo_contributors`.
- **Star-based ranking** sorts repositories by community popularity (`github_details.stars`), selecting only the top seven for final enrichment to maintain evaluation focus.
- Integration occurs through [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 5-11, 31-33), which automatically detects GitHub profiles in résumé data and triggers the complete enrichment pipeline.

## Frequently Asked Questions

### How does the algorithm differentiate between open-source and personal projects?

The system classifies a repository as **"open_source"** when `fetch_repo_contributors` returns more than one contributor, indicating collaborative development experience. Single-contributor repositories are labeled **"self_project"**, reflecting solo-authored work. This classification occurs in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 46-49) and helps hiring teams understand whether the candidate has experience working with distributed teams or primarily develops independently.

### Why does the system filter out forks with fewer than 5 forks?

The algorithm skips forks where `forks_count < 5` to exclude trivial personal copies of popular repositories that demonstrate minimal independent development effort. This filter (implemented at lines 34-36 of [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)) ensures only substantive forks—those that have attracted their own community interest—surface in the evaluation, providing more relevant signal about the candidate's actual maintenance capabilities and project ownership.

### How many projects does the system select for résumé enrichment?

The pipeline selects exactly **seven** projects, as strictly implemented in `fetch_and_display_github_info` (lines 42-55). This hard limit balances comprehensive candidate assessment with evaluation efficiency, focusing hiring teams on the candidate's highest-quality repositories sorted by star count in descending order.

### What determines which repositories appear in the final selection?

Repository selection follows a strict hierarchy: first, the system filters out trivial forks (fewer than 5 forks), then classifies by contributor type (open-source vs. self-project), and finally sorts by **star count** in descending order. The top seven repositories from this sorted list comprise the final enrichment payload, ensuring the highest-quality, most-community-validated projects represent the candidate's technical portfolio.