# GitHub Enrichment Pipeline Architecture in the Hiring Agent

> Discover the GitHub enrichment pipeline architecture in hiring agent. Learn how it extracts candidate data, uses LLMs for project selection, and creates a structured payload for evaluation.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: architecture
- Published: 2026-07-08

---

**The GitHub enrichment pipeline in the interviewstreet/hiring-agent repository extracts a candidate's username from their resume, fetches profile and repository data via the GitHub API, filters and classifies projects based on contributor counts, uses an LLM to select the top seven most relevant repositories, and assembles everything into a structured enrichment payload for downstream evaluation.**

The interviewstreet/hiring-agent project automates candidate evaluation by enriching resume data with public GitHub signals. Understanding the GitHub enrichment pipeline architecture reveals how raw profile URLs transform into structured, LLM-curated project portfolios that power fair, data-driven hiring decisions.

## Five-Stage Pipeline Architecture

The enrichment process implemented in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) follows a deterministic five-stage flow, moving from raw URL parsing to structured data assembly.

### Stage 1: Username Extraction

The pipeline begins with `extract_github_username()` in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 16-28), which parses any GitHub URL found in the candidate's resume to isolate the plain username handle. This normalizes various URL formats—whether the candidate provided `github.com/username` or a full profile path—into a consistent identifier for API calls.

### Stage 2: Profile Fetching

Using the extracted handle, `fetch_github_profile()` (lines 41-73) queries the GitHub Users API at `https://api.github.com/users/<username>`. The raw JSON response is validated and wrapped in a **`GitHubProfile`** Pydantic model defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), ensuring type safety for downstream consumers.

### Stage 3: Repository Retrieval

The `fetch_all_github_repos()` function (lines 18-88) orchestrates repository collection by hitting the `/users/<username>/repos` endpoint. For each repository returned, the pipeline:

- Filters out low-impact forks with fewer than **5 forks**
- Calls `fetch_repo_contributors()` to retrieve contributor metadata
- Calculates commit ratios via `fetch_contributions_count()` to determine the author's share of total activity
- Classifies the project as **`open_source`** (≥2 contributors) or **`self_project`** based on contributor thresholds

This classification distinguishes between collaborative community work and individual portfolio pieces.

### Stage 4: LLM-Driven Project Selection

Rather than returning all repositories, the pipeline uses `generate_projects_json()` (lines 34-68) to curate exactly **seven unique** projects. The function serializes the filtered repository list into JSON and passes it to the LLM via the `github_project_selection.jinja` template. The model ranks projects by relevance and impact, returning a strictly validated list. If the LLM response contains duplicates or parsing errors, the wrapper automatically deduplicates entries and falls back to the first seven items from the raw list.

### Stage 5: Payload Assembly

Finally, `fetch_and_display_github_info()` (lines 59-78) combines the profile data from `generate_profile_json()` with the LLM-selected projects into a single dictionary. This enrichment payload is consumed by downstream stages such as [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py) for candidate scoring.

## Caching and Rate Limit Protection

The pipeline includes defensive mechanisms against GitHub API constraints. The `_create_cache_filename()` helper implements a **caching layer** that stores raw API responses under the `cache/` directory when the `DEVELOPMENT_MODE` flag is enabled. This avoids redundant network calls during iterative development.

For production runs, the code monitors API rate limits in real time. When remaining requests drop below a safety threshold, the pipeline automatically sleeps until the GitHub API's reset time, preventing service interruptions during bulk candidate processing.

## Core Implementation Files

| File | Responsibility |
|------|----------------|
| [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) | Core pipeline implementation including username extraction, API orchestration, and LLM project selection |
| [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) | Pydantic schemas (`GitHubProfile`, `LLMProvider` abstractions) |
| [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) | Provider-agnostic LLM initialization and response cleanup |
| `prompts/templates/github_project_selection.jinja` | Template defining the LLM's ranking criteria and output format |
| [`prompt.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompt.py) | Central configuration for `DEFAULT_MODEL` and `LLM_PROVIDER` settings |

All LLM interactions are abstracted through [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) and [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), allowing the pipeline to switch between Ollama, Gemini, or future providers without code changes.

## Practical Code Examples

### Complete Enrichment Workflow

```python
from github import fetch_and_display_github_info

# Pass the GitHub URL found in the resume (or just the handle)

result = fetch_and_display_github_info("https://github.com/example_user")

print(result["profile"])   # GitHubProfile fields

print(result["projects"])  # List of 7 LLM-selected projects

```

### Low-Level Pipeline Access

```python
from github import (
    extract_github_username,
    fetch_github_profile,
    fetch_all_github_repos,
)

url = "https://github.com/example_user"
username = extract_github_username(url)

profile = fetch_github_profile(url)          # → GitHubProfile instance

repos   = fetch_all_github_repos(url, 50)    # → enriched repo dicts

# Inspect classification

print(repos[0]["name"], repos[0]["project_type"])

```

### Internal LLM Selection Logic

```python
from github import generate_projects_json

# Assuming `projects` is the list from fetch_all_github_repos()

selected = generate_projects_json(projects)   # → exactly 7 unique items

```

## Summary

- The pipeline extracts GitHub usernames from resume URLs using `extract_github_username()` in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)
- Profile data is fetched via the GitHub Users API and validated through Pydantic models
- Repository collection filters forks (<5) and classifies projects as `open_source` or `self_project` based on contributor counts
- An LLM selects exactly seven unique projects using the `github_project_selection.jinja` template, with automatic fallback handling
- The final payload combines profile and project data for consumption by [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)
- Built-in caching and rate-limit handling ensure reliable operation during high-volume processing

## Frequently Asked Questions

### How does the pipeline handle GitHub API rate limits?

The code monitors the API's remaining request count and reset timestamp. When approaching the limit, it pauses execution and sleeps until the reset time, ensuring the pipeline resumes automatically without manual intervention or failed requests.

### What distinguishes an open_source project from a self_project?

The classification depends on contributor count. Repositories with **two or more contributors** are labeled `open_source`, indicating collaborative community work. Projects with fewer than two contributors are classified as `self_project`, representing individual portfolio work.

### How does the LLM ensure exactly seven projects are returned?

The `generate_projects_json()` function enforces this constraint through the `github_project_selection.jinja` template, which explicitly instructs the model to return seven unique projects. The wrapper validates the response, deduplicates entries if necessary, and falls back to the first seven raw repositories if the LLM output is malformed.

### Can the pipeline run offline during development?

Yes. When `DEVELOPMENT_MODE` is enabled, the `_create_cache_filename()` mechanism stores raw GitHub API responses in the `cache/` directory. Subsequent runs load data from this cache instead of hitting the live API, enabling rapid iteration without network dependencies or rate limit concerns.