# How GitHub Profile and Repository Enrichment Works in Hiring Agent: A Technical Deep Dive

> Discover how Hiring Agent enriches GitHub profiles and repositories. This deep dive explains username extraction, API caching, repo classification, and LLM-powered project curation for hiring.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-07-05

---

**Hiring Agent enriches candidate résumés by extracting GitHub usernames from URLs, caching API responses to avoid rate limits, classifying repositories based on contributor counts, and leveraging an LLM to curate the seven most impressive projects for evaluation.**

The `interviewstreet/hiring-agent` repository automates technical candidate screening by augmenting résumés with objective, data-driven GitHub signals. The GitHub profile and repository enrichment process transforms raw profile URLs into structured metadata, distinguishing between collaborative open-source contributions and personal projects while minimizing API overhead through deterministic caching.

## Step-by-Step GitHub Enrichment Pipeline

The enrichment workflow follows a deterministic sequence designed for reproducibility and minimal API quota consumption.

### Extracting the GitHub Username

The pipeline begins by parsing the candidate's résumé for GitHub URLs or plain usernames. In [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 16‑38), the `extract_github_username` function uses regular expressions to normalize various URL formats into a canonical login handle.

```python

# github.py, lines 16-38

def extract_github_username(github_url: str) -> Optional[str]:
    ...

```

This extraction handles both full URLs (`https://github.com/torvalds`) and naked usernames, ensuring the downstream API calls receive consistent input.

### Cache-Aware API Handling

All GitHub REST requests route through `_fetch_github_api` in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 28‑63). This function implements a **file-based caching layer** that stores JSON responses under `cache/gh_githubcache_…` filenames.

```python

# github.py, lines 28-63

def _fetch_github_api(api_url, params=None):
    ...

```

The logic checks the cache directory before initiating network requests. If a cached file exists, it returns the stored data; otherwise, it executes the request with an optional `Authorization` header (when `GITHUB_TOKEN` is configured), respects GitHub's rate-limit headers, and persists the response to disk. This strategy accelerates development workflows and prevents quota exhaustion during repeated evaluations.

### Fetching Profile Metadata

Once the username is canonicalized, `fetch_github_profile` (lines 41‑78) queries the `/users/{username}` endpoint and hydrates a **Pydantic model** defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py) (lines 52‑68).

```python

# github.py, lines 41-78

def fetch_github_profile(github_url: str) -> Optional[GitHubProfile]:
    ...

```

The `GitHubProfile` model provides type safety for fields such as `name`, `bio`, `public_repos`, and `followers`, ensuring downstream components consume validated data structures.

### Repository Collection and Classification

The `fetch_all_github_repos` function (lines 18‑94) orchestrates the heavy lifting of repository analysis. It paginates through `https://api.github.com/users/{username}/repos` (sorted by recent activity) and performs three critical operations on each repository:

1. **Fork filtering** – Skips trivial forks to focus on original work.
2. **Contributor analysis** – Requests `/repos/{owner}/{repo}/contributors` to count total commits and the author's specific contribution volume.
3. **Repository classification** – Labels the repo as **open_source** (≥ 2 contributors) or **self_project** (< 2 contributors) based on community participation.

```python

# github.py, lines 18-94

def fetch_all_github_repos(github_url: str, max_repos: int = 100) -> List[Dict]:
    ...

```

The function extracts key signals including stars, primary language, topics, and commit counts, returning a uniform list of dictionaries for downstream processing.

### LLM-Powered Project Selection

Raw repository lists often contain noise from experimental or minor projects. The `generate_projects_json` function (excerpted at lines 34‑70) solves this by delegating curation to an LLM.

The system serializes the repository list to JSON and renders it via the `github_project_selection.jinja` template. The prompt strictly enforces **exactly seven unique projects**. After parsing the LLM response, the code deduplicates entries and implements fallback logic: if fewer than seven unique items are returned, it pads the selection with the highest-starred remaining repositories from the candidate's profile.

```python

# github.py, lines 34-70 (excerpt of generate_projects_json)

def generate_projects_json(projects: List[Dict]) -> List[Dict]:
    ...

```

This hybrid approach combines algorithmic filtering with semantic judgment, ensuring the final selection highlights the candidate's most impactful work.

### Data Integration and Output

The `fetch_and_display_github_info` function (lines 59‑78) serves as the orchestration layer. It aggregates the profile metadata, full repository set, and LLM-curated projects into a standardized dictionary:

- `"profile"` – Plain-dict representation of the `GitHubProfile` model.
- `"projects"` – The curated list of up to seven selected repositories.
- `"total_projects"` – Count of all fetched repositories for context.

```python

# github.py, lines 59-78

def fetch_and_display_github_info(github_url: str) -> Dict:
    ...

```

## Integration with the Résumé Scoring Workflow

In [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (the CLI entry point), the enrichment pipeline integrates seamlessly with the broader evaluation system. When processing a PDF résumé, the system scans the `basics.profiles` array for GitHub URLs. If detected, it invokes `fetch_and_display_github_info` and merges the returned dictionary into the `JSONResume` structure under the `github` key.

This enriched data then flows into the evaluator, which scores candidates on open-source engagement and project quality. The **Architecture** section of the repository's README outlines this flow as part of the end-to-end pipeline (PDF → LLM → GitHub → Evaluation).

## Practical Implementation Examples

### Stand-Alone Enrichment

You can invoke the enrichment pipeline independently for testing or custom integrations:

```python
from github import fetch_and_display_github_info

result = fetch_and_display_github_info("https://github.com/torvalds")
print(result["profile"]["name"])          # → Linus Torvalds

print("Top projects:")
for p in result["projects"]:
    print(f"- {p['name']} ({p['github_details']['stars']} ★)")

```

### Full Pipeline Integration

The following excerpt demonstrates how [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) embeds GitHub data into the evaluation workflow:

```python
from github import fetch_and_display_github_info
from pdf import PDFHandler   # parses the PDF → JSONResume

from evaluator import evaluate_resume

pdf_path = "candidate_resume.pdf"
resume_json = PDFHandler().process(pdf_path)

# Enrich with GitHub data if a profile URL is present

github_url = next(
    (p.url for p in resume_json.basics.profiles or [] if p.network == "GitHub"), None
)
if github_url:
    gh_data = fetch_and_display_github_info(github_url)
    resume_json.github = gh_data    # attaches profile & projects

evaluation = evaluate_resume(resume_json)
print(evaluation.scores.open_source.score)

```

### Command-Line Usage

The repository ships with a CLI entry point for direct execution:

```bash

# Command-line usage

$ python score.py path/to/resume.pdf

# → prints the enriched JSON and a human-readable scoring summary

```

## Summary

- **Username extraction** normalizes GitHub URLs into canonical handles using regex patterns in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py).
- **Intelligent caching** via `_fetch_github_api` eliminates redundant network calls and respects GitHub rate limits.
- **Repository classification** distinguishes **open_source** projects (≥ 2 contributors) from **self_project** work based on contributor analysis.
- **LLM curation** selects exactly seven impressive projects using the `github_project_selection.jinja` template, with algorithmic fallback to star-count ranking.
- **Seamless integration** via `fetch_and_display_github_info` attaches structured GitHub data to `JSONResume` objects for downstream scoring in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).

## Frequently Asked Questions

### How does Hiring Agent handle GitHub API rate limits?

The `_fetch_github_api` function in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) implements a file-based cache under the `cache/` directory. Before making any request, it checks for a cached response stored as `cache/gh_githubcache_…`. If found, it returns the cached JSON immediately. For fresh requests, it respects GitHub's rate-limit headers and supports authentication via the `GITHUB_TOKEN` environment variable to increase quota limits.

### What criteria determine if a repository is classified as open source?

A repository is classified as **open_source** when the contributors endpoint reveals **two or more distinct contributors**. The `fetch_all_github_repos` function queries `/repos/{owner}/{repo}/contributors` and counts unique authors. Repositories with only one contributor (typically the candidate) are labeled **self_project**, distinguishing collaborative work from personal experiments.

### Why does the enrichment process use an LLM for project selection?

The LLM serves as a semantic filter to identify the "most impressive" projects from potentially lengthy repository lists. The `generate_projects_json` function serializes repository metadata (stars, language, topics, commit counts) and prompts the model via the `github_project_selection.jinja` template to select exactly seven unique items. This combines quantitative metrics with qualitative judgment, ensuring the final selection highlights high-impact work rather than just the most recent commits.

### Can the GitHub enrichment run independently of the full scoring pipeline?

Yes. The `fetch_and_display_github_info` function in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) operates as a standalone entry point. You can import and call it directly with any GitHub URL to receive a structured dictionary containing the profile metadata and curated project list without invoking the PDF parser or evaluation engine. This modularity supports debugging, testing, and custom integrations outside the standard [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) workflow.