# How Hiring Agent Enriches Resume Data with GitHub Signals: A Technical Deep Dive

> Learn how Hiring Agent enriches resume data using GitHub signals. Discover how it fetches profile and repo data, merging signals for advanced candidate evaluation.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-06-29

---

**Hiring Agent extracts GitHub URLs from resumes, fetches profile and repository data via the GitHub API, and merges these signals into both structured CSV exports and LLM-readable markdown text to enable richer candidate evaluation.**

The `interviewstreet/hiring-agent` repository treats a candidate’s GitHub presence as a first-class signal in the hiring pipeline. By parsing profile URLs from JSON-Resume documents and integrating public repository metrics, the system transforms static resumes into dynamic technical portfolios. This enrichment process involves seven distinct stages spanning three core modules: [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py), and [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py).

## Step 1: Extracting the GitHub Username from Resume URLs

The enrichment pipeline begins in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), where the `extract_github_username()` function (lines 116–131) parses the *basics* section of a JSON-Resume to locate profile URLs. The function normalizes the URL structure and isolates the username for subsequent API calls.

```python
from github import extract_github_username

# Extract username from any standard GitHub URL format

github_url = "https://github.com/awesome-dev"
username = extract_github_username(github_url)   # → "awesome-dev"

```

This normalization supports various URL formats and ensures consistent handling before the API request stage.

## Step 2: Fetching and Caching GitHub Profile Data

Once the username is isolated, `fetch_and_display_github_info()` (lines 112–124 in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)) orchestrates the API interaction. The function calls `_fetch_github_api()` (lines 29–33) to retrieve both profile statistics and repository listings via the GitHub REST API.

Responses are cached locally under `cache/gh_githubcache_…json` to prevent rate-limiting and improve performance. The system supports optional authentication via the `GITHUB_TOKEN` environment variable for higher rate limits and access to private repository metadata.

```python
from github import fetch_and_display_github_info

# Returns a dict with keys "profile" and "projects"

github_data = fetch_and_display_github_info("https://github.com/awesome-dev")

```

## Step 3: Normalizing API Responses into Structured Data

The raw GitHub API JSON is normalized into a uniform dictionary structure with two top-level keys: `profile` (containing basic account statistics) and `projects` (containing repository details). This normalization occurs within `fetch_and_display_github_info()` and ensures downstream components receive predictable data regardless of API response variations.

The normalized structure includes fields such as `public_repos`, `followers`, `following`, `created_at`, and `bio` at the profile level, plus repository-specific metrics like stars, forks, and primary language for each project.

## Step 4: Injecting GitHub Metrics into CSV Exports

In [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) (starting at line 659), the system merges GitHub signals into structured CSV exports. When `github_data` is present, the transformer appends numeric signals to the candidate row:

- **public_repos**: Total public repositories
- **followers**: GitHub follower count
- **following**: Number of accounts followed
- **created_at**: Account creation date
- **bio**: Profile biography text

If the API fetch fails or returns null, the system applies safe defaults (0 for numeric fields, empty strings for text).

```python
def transform_resume(resume_data, github_data=None):
    csv_row = {}
    # ... other fields ...

    if github_data:
        csv_row["github_repos"]      = github_data.get("public_repos", 0)
        csv_row["github_followers"]  = github_data.get("followers", 0)
        csv_row["github_following"]  = github_data.get("following", 0)
        csv_row["github_created_at"] = github_data.get("created_at", "")
        csv_row["github_bio"]        = github_data.get("bio", "")

```

## Step 5: Converting Repository Data to LLM-Readable Markdown

For LLM evaluation, [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) provides `convert_github_data_to_text()` (lines 892–921), which renders the structured GitHub data as a markdown block. This transformation includes a header section (`=== GITHUB DATA ===`) followed by profile bullet points and a numbered list of top projects with their star counts, fork counts, and primary languages.

```python
def convert_github_data_to_text(github_data):
    github_text = "\n\n=== GITHUB DATA ===\n"
    profile = github_data["profile"]
    github_text += (
        f"- Username: {profile.get('username', 'N/A')}\n"
        f"- Public Repositories: {profile.get('public_repos', 'N/A')}\n"
        f"- Followers: {profile.get('followers', 'N/A')}\n"
        # … additional profile fields …

    )
    # Append project details

    for i, project in enumerate(github_data["projects"], 1):
        github_text += f"\n{i}. {project.get('name', 'N/A')}\n"
        github_text += f"   Stars: {project.get('github_details', {}).get('stars', 'N/A')}\n"
        github_text += f"   Language: {project.get('github_details', {}).get('language', 'N/A')}\n"
    return github_text

```

## Step 6: Merging Signals for Enhanced Candidate Scoring

The final enrichment occurs in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 174–176), where the GitHub markdown is appended to the plain-text resume before evaluation. The `_evaluate_resume()` function (line 326) receives this combined narrative, allowing the LLM to assess technical depth, code quality indicators, and open-source activity alongside traditional resume content.

```python
resume_text = render_resume_to_plain_text(resume_data)
if github_data:
    resume_text += convert_github_data_to_text(github_data)

# Pass enriched data to evaluation

score = _evaluate_resume(resume_data, github_data)

```

## Summary

- **Hiring Agent** detects GitHub URLs in JSON-Resume documents using `extract_github_username()` in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py).
- **API calls** are cached locally and support optional token-based authentication for higher rate limits.
- **Structured data** is normalized into a consistent format with `profile` and `projects` keys.
- **CSV exports** receive numeric GitHub metrics (repos, followers, account age) for spreadsheet analysis.
- **LLM evaluation** consumes markdown-formatted GitHub data appended to the resume text, enabling assessment of technical activity and repository quality.

## Frequently Asked Questions

### Where does Hiring Agent cache GitHub API responses?

Hiring Agent stores API responses in local JSON files under the `cache/` directory with the naming pattern `gh_githubcache_…json`. This caching strategy prevents redundant API calls and reduces the risk of hitting GitHub rate limits during batch resume processing.

### What happens if a candidate's GitHub profile is private or the API request fails?

When the API request fails or returns no data, the system gracefully handles the absence by applying default values. Numeric fields receive `0` and text fields receive empty strings in the CSV export, while the LLM evaluation proceeds with only the resume text—no GitHub section is appended.

### Which specific GitHub metrics are included in the CSV export?

The CSV export includes five primary signals from the `profile` object: `public_repos` (repository count), `followers` (follower count), `following` (following count), `created_at` (account creation date), and `bio` (profile biography). These metrics provide quantifiable indicators of the candidate's platform activity and tenure.

### How does the LLM use GitHub data during candidate evaluation?

The LLM receives the GitHub data as a markdown-formatted text block appended to the candidate's resume. This combined narrative allows the model to evaluate technical depth based on repository topics, assess impact through star and fork counts, and verify claimed skills against actual code languages used in public projects.