How Hiring Agent Enriches Resume Data with GitHub Information: A Technical Deep Dive

Hiring Agent extracts GitHub profile URLs from candidate resumes, fetches repository statistics via the GitHub REST API, and integrates these signals into both structured CSV exports and LLM scoring contexts to provide a comprehensive technical assessment.

The open-source interviewstreet/hiring-agent project treats GitHub presence as a first-class signal in the recruitment pipeline. By parsing JSON Resume data, normalizing API responses, and injecting open-source metrics into candidate evaluations, it transforms static CVs into dynamic competency profiles backed by actual code contributions.

Extracting GitHub Usernames from Resume URLs

The enrichment process begins in github.py, where the system identifies GitHub URLs embedded in the basics section of JSON Resumes. The extract_github_username() function (lines 116–131) normalizes the URL and isolates the username for downstream API queries.

from github import extract_github_username

github_url = "https://github.com/awesome-dev"
username = extract_github_username(github_url)   # → "awesome-dev"

Source: github.py – extract_github_username()【https://github.com/interviewstreet/hiring-agent/blob/main/github.py#L116-L131】

Fetching and Caching GitHub API Data

Once extracted, the username feeds into fetch_and_display_github_info() (lines 112–124), which orchestrates the API call via the internal _fetch_github_api() method (lines 29–33). The system queries the GitHub REST API for both profile metadata and repository details, storing responses in cache/gh_githubcache_…json to minimize redundant network requests and avoid rate limits. An optional GITHUB_TOKEN environment variable enables authenticated requests for higher API quotas.

from github import fetch_and_display_github_info

# Returns a dict with keys "profile" and "projects"

github_data = fetch_and_display_github_info("https://github.com/awesome-dev")

This function returns a normalized dictionary containing two top-level keys: profile (basic account statistics) and projects (repository details with stars, forks, and primary languages).

Structuring Data for CSV Exports

In transform.py (starting at line 659), the pipeline extracts numeric signals from the GitHub response and appends them to the candidate's CSV row. If the API call fails or returns null values, the system defaults to 0 for counters and empty strings for text fields to maintain data integrity.

The following fields are injected into the export:

  • github_repos: Public repository count
  • github_followers: Follower count
  • github_following: Following count
  • github_created_at: Account creation date (technical tenure indicator)
  • github_bio: Self-reported expertise description
def transform_resume(resume_data, github_data=None):
    csv_row = {}
    # ... other fields ...

    if github_data:
        csv_row["github_repos"]      = github_data.get("public_repos", 0)
        csv_row["github_followers"]  = github_data.get("followers", 0)
        csv_row["github_following"]  = github_data.get("following", 0)
        csv_row["github_created_at"] = github_data.get("created_at", "")
        csv_row["github_bio"]        = github_data.get("bio", "")

Source: transform.py – lines 659–667【https://github.com/interviewstreet/hiring-agent/blob/main/transform.py#L659-L667】

Converting GitHub Data to LLM-Readable Text

For AI-driven evaluation, the raw JSON must become narrative text. The convert_github_data_to_text() function (lines 892–921 in transform.py) renders a markdown block starting with === GITHUB DATA ===, followed by profile statistics and an enumerated list of top projects including star counts, fork counts, and primary languages.

def convert_github_data_to_text(github_data):
    github_text = "\n\n=== GITHUB DATA ===\n"
    profile = github_data["profile"]
    github_text += (
        f"- Username: {profile.get('username', 'N/A')}\n"
        f"- Public Repositories: {profile.get('public_repos', 'N/A')}\n"
        f"- Followers: {profile.get('followers', 'N/A')}\n"
        # … other fields …

    )
    # List projects

    for i, project in enumerate(github_data["projects"], 1):
        github_text += f"\n{i}. {project.get('name', 'N/A')}\n"
        github_text += f"   Stars: {project.get('github_details', {}).get('stars', 'N/A')}\n"
        # … etc …

    return github_text

Source: transform.py – convert_github_data_to_text()【https://github.com/interviewstreet/hiring-agent/blob/main/transform.py#L892-L921】

Merging Signals for LLM Scoring

The final enrichment occurs in score.py, where the system prepares the evaluation context. Lines 174–176 append the GitHub markdown text to the plain-text resume before invoking _evaluate_resume() (line 326).

resume_text = render_resume_to_plain_text(resume_data)
if github_data:
    resume_text += convert_github_data_to_text(github_data)

score = _evaluate_resume(resume_data, github_data)

Source: score.py – lines 174–176【https://github.com/interviewstreet/hiring-agent/blob/main/score.py#L174-L176】 and line 326【https://github.com/interviewstreet/hiring-agent/blob/main/score.py#L326】

This merged narrative allows the LLM to evaluate technical depth, open-source impact, and coding activity alongside traditional resume qualifications.

Summary

Hiring Agent enriches resume data with GitHub information through a seven-stage pipeline:

  • URL Extraction: Parses GitHub URLs from the basics section of JSON Resumes using extract_github_username() in github.py (lines 116–131)
  • API Integration: Fetches profile and repository data via fetch_and_display_github_info(), with automatic caching to cache/gh_githubcache_…json
  • Data Normalization: Structures API responses into consistent dictionaries with profile and projects keys
  • CSV Enrichment: Injects numeric metrics (repos, followers, tenure) into export rows via transform.py (line 659)
  • Text Conversion: Renders GitHub data as markdown narratives using convert_github_data_to_text() (lines 892–921)
  • LLM Augmentation: Appends GitHub context to resume text in score.py (lines 174–176) before evaluation
  • Unified Scoring: The _evaluate_resume() function (line 326) processes the combined dataset for technical assessment

Frequently Asked Questions

How does Hiring Agent handle GitHub API rate limits?

The system implements file-based caching under cache/gh_githubcache_…json to store API responses and minimize redundant calls. For production deployments, set the optional GITHUB_TOKEN environment variable to authenticate requests and benefit from GitHub's higher rate limits for authorized users.

What specific GitHub metrics are added to the CSV export?

According to transform.py (lines 659–667), the system extracts five key fields: public_repos (repository count), followers, following, created_at (account creation date), and bio. These values default to 0 or empty strings if the API call fails or the data is unavailable.

What happens if a candidate's GitHub URL is malformed or the profile is private?

The extract_github_username() function in github.py normalizes URLs before extraction. If the API fetch fails or returns no data, transform.py handles the null case gracefully by substituting default values (0 for numeric fields, empty strings for text), ensuring the pipeline continues without interruption.

Why does Hiring Agent convert GitHub JSON data to markdown before LLM scoring?

The convert_github_data_to_text() function (lines 892–921 in transform.py) transforms structured API responses into a narrative format prefixed with === GITHUB DATA ===. This markdown block integrates seamlessly with the plain-text resume, providing the LLM with contextual signals about technical activity, popular repositories, and coding languages in a readable format that enhances evaluation accuracy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →