How GitHub Usernames Are Extracted from Resume Profiles in the Hiring Agent

The hiring-agent extracts GitHub usernames by scanning the resume's profiles section for GitHub URLs, normalizing the input, and applying regular expression patterns to isolate the username from the URL path.

The interviewstreet/hiring-agent repository automates candidate screening by parsing structured resume data to identify technical presence. Understanding how GitHub usernames are extracted from resume profiles reveals the pipeline that transforms raw profile URLs into clean, queryable data for downstream technical evaluation.

Profile Discovery in transform.py

The extraction process begins in transform.py where the fetch_profile helper function iterates through the candidate's profile objects. Each profile entry contains a network name and a url field. When the system identifies a profile whose network matches github, it retrieves the URL and prepares it for processing.

The extracted username is then stored directly in the CSV output row under the github_username column:


# transform.py (lines 544-545)

csv_row["github_username"] = (
    github_profile.username if github_profile.username else ""
)

This assignment occurs after the fetch_profile utility validates that the profile belongs to the GitHub network, ensuring only relevant URLs proceed to the extraction stage.

Username Extraction Logic in github.py

The extract_github_username function in github.py handles the actual parsing logic. This utility accepts a GitHub URL string and returns the isolated username or None if the format is unrecognizable.

URL Normalization

Before pattern matching, the function sanitizes the input to handle malformed entries. It removes all internal spaces and strips surrounding whitespace from the URL string:

github_url = github_url.replace(" ", "").strip()

This normalization step ensures that URLs containing accidental spaces or inconsistent formatting do not fail the regex matching phase.

Regex Pattern Matching

The extraction engine uses two complementary regular expressions to capture usernames from various URL formats:

  • r"https?://github\.com/([^/]+)" – Matches full URLs with optional HTTPS protocol
  • r"github\.com/([^/]+)" – Matches protocol-relative or bare domain URLs

The function iterates through these patterns using re.search and returns the first captured group (the username) upon finding a match:


# github.py (lines 124-129)

def extract_github_username(github_url: str) -> Optional[str]:
    if not github_url:
        return None
    github_url = github_url.replace(" ", "").strip()
    patterns = [
        r"https?://github\.com/([^/]+)",
        r"github\.com/([^/]+)",
    ]
    for pattern in patterns:
        match = re.search(pattern, github_url)
        if match:
            return match.group(1)
    return None

If neither pattern matches the normalized URL, the function returns None, allowing the pipeline to handle missing or invalid GitHub links gracefully.

Data Pipeline Flow

The complete extraction pipeline follows this sequence:

  1. Resume Parsing – The system reads the JSON resume and locates the basics.profiles array
  2. Network Identification – fetch_profile filters for entries where network == "github"
  3. URL Validation – The extracted URL passes to extract_github_username
  4. Normalization – Whitespace and spaces are removed from the URL string
  5. Pattern Matching – Regular expressions attempt to capture the username segment
  6. Data Persistence – The clean username (or empty string)写入 the CSV row under github_username

This architecture makes downstream scoring modules independent of URL format variations, supporting inputs like https://github.com/username, http://github.com/username, or github.com/username.

Summary

  • Profile Scanning: The fetch_profile helper in transform.py identifies GitHub entries by matching the network field against "github"
  • URL Cleaning: The extract_github_username function in github.py removes whitespace and normalizes the input before processing
  • Regex Extraction: Two patterns—https?://github\.com/([^/]+) and github\.com/([^/]+)—capture usernames from various URL formats
  • Fail-Safe Handling: The pipeline returns empty strings or None for missing or malformed URLs, preventing pipeline failures
  • CSV Integration: Extracted usernames populate the github_username column for downstream technical evaluation

Frequently Asked Questions

What file handles the initial GitHub profile detection?

The transform.py file contains the fetch_profile helper function that iterates through the resume's profiles array. It identifies GitHub entries by checking if the profile's network field matches "github", then passes the corresponding URL to the extraction utility.

Which regex patterns are used to extract the username?

The system employs two patterns in github.py: r"https?://github\.com/([^/]+)" for full URLs with HTTP/HTTPS protocols, and r"github\.com/([^/]+)" for protocol-relative or bare domain formats. Both patterns capture the first path segment as the username.

How does the system handle malformed URLs with extra spaces?

Before regex matching, the extract_github_username function calls replace(" ", "") to remove all internal spaces and strip() to eliminate surrounding whitespace. This normalization ensures that accidentally spaced URLs still parse correctly.

What happens if the URL does not match the expected patterns?

If neither regex pattern returns a match, the extract_github_username function returns None. The calling code in transform.py then converts this to an empty string for the CSV output, allowing the pipeline to continue processing without interruption.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →