# How Hiring Agent Enriches Resume Data with GitHub Information: A Technical Deep Dive

> Learn how Hiring Agent enriches resume data with GitHub information. Fetch repo stats via GitHub API and integrate signals for comprehensive technical assessments.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: deep-dive
- Published: 2026-07-02

---

**Hiring Agent extracts GitHub profile URLs from candidate resumes, fetches repository statistics via the GitHub REST API, and integrates these signals into both structured CSV exports and LLM scoring contexts to provide a comprehensive technical assessment.**

The open-source interviewstreet/hiring-agent project treats GitHub presence as a first-class signal in the recruitment pipeline. By parsing JSON Resume data, normalizing API responses, and injecting open-source metrics into candidate evaluations, it transforms static CVs into dynamic competency profiles backed by actual code contributions.

## Extracting GitHub Usernames from Resume URLs

The enrichment process begins in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), where the system identifies GitHub URLs embedded in the *basics* section of JSON Resumes. The `extract_github_username()` function (lines 116–131) normalizes the URL and isolates the username for downstream API queries.

```python
from github import extract_github_username

github_url = "https://github.com/awesome-dev"
username = extract_github_username(github_url)   # → "awesome-dev"

```

*Source:* [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) – `extract_github_username()`【https://github.com/interviewstreet/hiring-agent/blob/main/github.py#L116-L131】

## Fetching and Caching GitHub API Data

Once extracted, the username feeds into `fetch_and_display_github_info()` (lines 112–124), which orchestrates the API call via the internal `_fetch_github_api()` method (lines 29–33). The system queries the GitHub REST API for both profile metadata and repository details, storing responses in `cache/gh_githubcache_…json` to minimize redundant network requests and avoid rate limits. An optional `GITHUB_TOKEN` environment variable enables authenticated requests for higher API quotas.

```python
from github import fetch_and_display_github_info

# Returns a dict with keys "profile" and "projects"

github_data = fetch_and_display_github_info("https://github.com/awesome-dev")

```

This function returns a normalized dictionary containing two top-level keys: `profile` (basic account statistics) and `projects` (repository details with stars, forks, and primary languages).

## Structuring Data for CSV Exports

In [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) (starting at line 659), the pipeline extracts numeric signals from the GitHub response and appends them to the candidate's CSV row. If the API call fails or returns null values, the system defaults to **0** for counters and empty strings for text fields to maintain data integrity.

The following fields are injected into the export:

- **github_repos**: Public repository count
- **github_followers**: Follower count
- **github_following**: Following count
- **github_created_at**: Account creation date (technical tenure indicator)
- **github_bio**: Self-reported expertise description

```python
def transform_resume(resume_data, github_data=None):
    csv_row = {}
    # ... other fields ...

    if github_data:
        csv_row["github_repos"]      = github_data.get("public_repos", 0)
        csv_row["github_followers"]  = github_data.get("followers", 0)
        csv_row["github_following"]  = github_data.get("following", 0)
        csv_row["github_created_at"] = github_data.get("created_at", "")
        csv_row["github_bio"]        = github_data.get("bio", "")

```

*Source:* [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) – lines 659–667【https://github.com/interviewstreet/hiring-agent/blob/main/transform.py#L659-L667】

## Converting GitHub Data to LLM-Readable Text

For AI-driven evaluation, the raw JSON must become narrative text. The `convert_github_data_to_text()` function (lines 892–921 in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py)) renders a markdown block starting with `=== GITHUB DATA ===`, followed by profile statistics and an enumerated list of top projects including star counts, fork counts, and primary languages.

```python
def convert_github_data_to_text(github_data):
    github_text = "\n\n=== GITHUB DATA ===\n"
    profile = github_data["profile"]
    github_text += (
        f"- Username: {profile.get('username', 'N/A')}\n"
        f"- Public Repositories: {profile.get('public_repos', 'N/A')}\n"
        f"- Followers: {profile.get('followers', 'N/A')}\n"
        # … other fields …

    )
    # List projects

    for i, project in enumerate(github_data["projects"], 1):
        github_text += f"\n{i}. {project.get('name', 'N/A')}\n"
        github_text += f"   Stars: {project.get('github_details', {}).get('stars', 'N/A')}\n"
        # … etc …

    return github_text

```

*Source:* [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) – `convert_github_data_to_text()`【https://github.com/interviewstreet/hiring-agent/blob/main/transform.py#L892-L921】

## Merging Signals for LLM Scoring

The final enrichment occurs in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py), where the system prepares the evaluation context. Lines 174–176 append the GitHub markdown text to the plain-text resume before invoking `_evaluate_resume()` (line 326).

```python
resume_text = render_resume_to_plain_text(resume_data)
if github_data:
    resume_text += convert_github_data_to_text(github_data)

score = _evaluate_resume(resume_data, github_data)

```

*Source:* [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) – lines 174–176【https://github.com/interviewstreet/hiring-agent/blob/main/score.py#L174-L176】 and line 326【https://github.com/interviewstreet/hiring-agent/blob/main/score.py#L326】

This merged narrative allows the LLM to evaluate technical depth, open-source impact, and coding activity alongside traditional resume qualifications.

## Summary

Hiring Agent enriches resume data with GitHub information through a seven-stage pipeline:

- **URL Extraction**: Parses GitHub URLs from the *basics* section of JSON Resumes using `extract_github_username()` in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) (lines 116–131)
- **API Integration**: Fetches profile and repository data via `fetch_and_display_github_info()`, with automatic caching to `cache/gh_githubcache_…json`
- **Data Normalization**: Structures API responses into consistent dictionaries with `profile` and `projects` keys
- **CSV Enrichment**: Injects numeric metrics (repos, followers, tenure) into export rows via [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) (line 659)
- **Text Conversion**: Renders GitHub data as markdown narratives using `convert_github_data_to_text()` (lines 892–921)
- **LLM Augmentation**: Appends GitHub context to resume text in [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) (lines 174–176) before evaluation
- **Unified Scoring**: The `_evaluate_resume()` function (line 326) processes the combined dataset for technical assessment

## Frequently Asked Questions

### How does Hiring Agent handle GitHub API rate limits?

The system implements file-based caching under `cache/gh_githubcache_…json` to store API responses and minimize redundant calls. For production deployments, set the optional `GITHUB_TOKEN` environment variable to authenticate requests and benefit from GitHub's higher rate limits for authorized users.

### What specific GitHub metrics are added to the CSV export?

According to [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) (lines 659–667), the system extracts five key fields: `public_repos` (repository count), `followers`, `following`, `created_at` (account creation date), and `bio`. These values default to 0 or empty strings if the API call fails or the data is unavailable.

### What happens if a candidate's GitHub URL is malformed or the profile is private?

The `extract_github_username()` function in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) normalizes URLs before extraction. If the API fetch fails or returns no data, [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) handles the null case gracefully by substituting default values (0 for numeric fields, empty strings for text), ensuring the pipeline continues without interruption.

### Why does Hiring Agent convert GitHub JSON data to markdown before LLM scoring?

The `convert_github_data_to_text()` function (lines 892–921 in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py)) transforms structured API responses into a narrative format prefixed with `=== GITHUB DATA ===`. This markdown block integrates seamlessly with the plain-text resume, providing the LLM with contextual signals about technical activity, popular repositories, and coding languages in a readable format that enhances evaluation accuracy.