# How GitHub Usernames Are Extracted from Resume Profiles in the Hiring Agent

> Discover how the hiring-agent extracts GitHub usernames from resumes. Learn about URL scanning, normalization, and regex patterns used to identify your GitHub profile.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-08

---

**The hiring-agent extracts GitHub usernames by scanning the resume's profiles section for GitHub URLs, normalizing the input, and applying regular expression patterns to isolate the username from the URL path.**

The `interviewstreet/hiring-agent` repository automates candidate screening by parsing structured resume data to identify technical presence. Understanding how GitHub usernames are extracted from resume profiles reveals the pipeline that transforms raw profile URLs into clean, queryable data for downstream technical evaluation.

## Profile Discovery in transform.py

The extraction process begins in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) where the `fetch_profile` helper function iterates through the candidate's profile objects. Each profile entry contains a `network` name and a `url` field. When the system identifies a profile whose `network` matches **github**, it retrieves the URL and prepares it for processing.

The extracted username is then stored directly in the CSV output row under the `github_username` column:

```python

# transform.py (lines 544-545)

csv_row["github_username"] = (
    github_profile.username if github_profile.username else ""
)

```

This assignment occurs after the `fetch_profile` utility validates that the profile belongs to the GitHub network, ensuring only relevant URLs proceed to the extraction stage.

## Username Extraction Logic in github.py

The `extract_github_username` function in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) handles the actual parsing logic. This utility accepts a GitHub URL string and returns the isolated username or `None` if the format is unrecognizable.

### URL Normalization

Before pattern matching, the function sanitizes the input to handle malformed entries. It removes all internal spaces and strips surrounding whitespace from the URL string:

```python
github_url = github_url.replace(" ", "").strip()

```

This normalization step ensures that URLs containing accidental spaces or inconsistent formatting do not fail the regex matching phase.

### Regex Pattern Matching

The extraction engine uses two complementary regular expressions to capture usernames from various URL formats:

- **`r"https?://github\.com/([^/]+)"`** – Matches full URLs with optional HTTPS protocol
- **`r"github\.com/([^/]+)"`** – Matches protocol-relative or bare domain URLs

The function iterates through these patterns using `re.search` and returns the first captured group (the username) upon finding a match:

```python

# github.py (lines 124-129)

def extract_github_username(github_url: str) -> Optional[str]:
    if not github_url:
        return None
    github_url = github_url.replace(" ", "").strip()
    patterns = [
        r"https?://github\.com/([^/]+)",
        r"github\.com/([^/]+)",
    ]
    for pattern in patterns:
        match = re.search(pattern, github_url)
        if match:
            return match.group(1)
    return None

```

If neither pattern matches the normalized URL, the function returns `None`, allowing the pipeline to handle missing or invalid GitHub links gracefully.

## Data Pipeline Flow

The complete extraction pipeline follows this sequence:

1. **Resume Parsing** – The system reads the JSON resume and locates the `basics.profiles` array
2. **Network Identification** – `fetch_profile` filters for entries where `network == "github"`
3. **URL Validation** – The extracted URL passes to `extract_github_username`
4. **Normalization** – Whitespace and spaces are removed from the URL string
5. **Pattern Matching** – Regular expressions attempt to capture the username segment
6. **Data Persistence** – The clean username (or empty string)写入 the CSV row under `github_username`

This architecture makes downstream scoring modules independent of URL format variations, supporting inputs like `https://github.com/username`, `http://github.com/username`, or `github.com/username`.

## Summary

- **Profile Scanning**: The `fetch_profile` helper in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) identifies GitHub entries by matching the `network` field against "github"
- **URL Cleaning**: The `extract_github_username` function in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) removes whitespace and normalizes the input before processing
- **Regex Extraction**: Two patterns—`https?://github\.com/([^/]+)` and `github\.com/([^/]+)`—capture usernames from various URL formats
- **Fail-Safe Handling**: The pipeline returns empty strings or `None` for missing or malformed URLs, preventing pipeline failures
- **CSV Integration**: Extracted usernames populate the `github_username` column for downstream technical evaluation

## Frequently Asked Questions

### What file handles the initial GitHub profile detection?

The [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) file contains the `fetch_profile` helper function that iterates through the resume's profiles array. It identifies GitHub entries by checking if the profile's `network` field matches "github", then passes the corresponding URL to the extraction utility.

### Which regex patterns are used to extract the username?

The system employs two patterns in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py): `r"https?://github\.com/([^/]+)"` for full URLs with HTTP/HTTPS protocols, and `r"github\.com/([^/]+)"` for protocol-relative or bare domain formats. Both patterns capture the first path segment as the username.

### How does the system handle malformed URLs with extra spaces?

Before regex matching, the `extract_github_username` function calls `replace(" ", "")` to remove all internal spaces and `strip()` to eliminate surrounding whitespace. This normalization ensures that accidentally spaced URLs still parse correctly.

### What happens if the URL does not match the expected patterns?

If neither regex pattern returns a match, the `extract_github_username` function returns `None`. The calling code in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) then converts this to an empty string for the CSV output, allowing the pipeline to continue processing without interruption.