How to Extract GitHub Profile URLs from Resume JSON with the Hiring-Agent Repository
The Hiring-Agent repository parses JSON Resume data using the fetch_profile function in transform.py to extract GitHub URLs from the basics.profiles array and normalizes them into usernames via extract_github_username in github.py.
When processing candidate resumes in the jsonresume.org format, the interviewstreet/hiring-agent codebase provides a robust pipeline to identify and extract GitHub profile URLs. This extraction enables automated GitHub API integration for repository analysis and candidate evaluation, transforming unstructured social profile data into actionable CSV fields.
Understanding the Resume JSON Structure
The Hiring-Agent expects resumes following the standard JSON Resume schema, specifically the basics.profiles array. This array contains social network profiles where each entry includes a network name, username, and url field.
{
"basics": {
"profiles": [
{
"network": "GitHub",
"username": "alice",
"url": "https://github.com/alice"
}
]
}
}
The extraction logic specifically targets entries where the network field matches "github" (case-insensitive) to locate the corresponding profile URL.
Extracting GitHub URLs with transform.py
The transform.py module serves as the primary orchestrator for converting resume JSON into structured CSV data. Within this module, the fetch_profile function handles the heavy lifting of social profile extraction.
The fetch_profile Function
Located at lines 531-545 in transform.py, the fetch_profile function scans the profiles array for network matches. It accepts the profiles list, a list of acceptable network aliases, and a target network identifier.
# Internal implementation reference from transform.py lines 531-545
def fetch_profile(profiles, network_aliases, target):
for profile in profiles:
if profile.get("network", "").lower() in network_aliases:
return profile.get("url", "")
return None
When processing a resume, the code calls fetch_profile(basics.profiles, ["github"], "github") to locate the GitHub entry. This approach supports multiple aliases and case-insensitive matching, ensuring variations like "GitHub", "github", or "GITHUB" are recognized correctly.
URL Validation and Storage
Upon finding a matching profile, the function validates the url field and stores it in the CSV row dictionary under the key github_url. This field becomes the canonical reference for all subsequent GitHub-related operations in the hiring pipeline.
The extraction occurs within the broader transform_resume() function, which coordinates between resume parsing, GitHub data fetching, and evaluation scoring to produce a flat CSV-ready row structure.
Normalizing URLs to Usernames with github.py
While transform.py extracts the raw URL, github.py provides utilities for URL normalization and username extraction. This separation of concerns allows the URL extraction to remain independent of GitHub API specifics.
The extract_github_username Helper
The extract_github_username function (lines 116-131 in github.py) strips GitHub URLs down to plain usernames using regex pattern matching. It handles various URL formats including https://github.com/username, http://github.com/username, and github.com/username.
from main.github import extract_github_username
# Normalize a full URL to just the username
url = "https://github.com/alice"
username = extract_github_username(url) # Returns: "alice"
The implementation uses regex patterns r"https?://github\.com/([^/]+)" and r"github\.com/([^/]+)" to extract the username segment, removing whitespace and normalizing the scheme before processing. This normalized username is then stored in the github_username CSV field for API calls.
Complete Implementation Example
To extract GitHub URLs from a resume JSON without running the full pipeline, use the following approach combining both modules:
from main.transform import transform_resume
from main.github import extract_github_username
# Sample resume data following jsonresume.org format
resume_data = {
"basics": {
"profiles": [
{
"network": "GitHub",
"username": "alice",
"url": "https://github.com/alice"
},
{
"network": "LinkedIn",
"url": "https://linkedin.com/in/alice"
}
]
}
}
# Transform into CSV row structure
csv_row = transform_resume(
file_name="alice_resume.pdf",
resume_data=resume_data,
github_data=None,
evaluation=None
)
# Extract GitHub data
github_url = csv_row.get("github_url") # "https://github.com/alice"
github_username = csv_row.get("github_username") # "alice"
# Alternative: Direct username extraction from any URL
clean_username = extract_github_username("https://github.com/alice") # "alice"
For lightweight extraction without the full CSV transformation, access the profiles array directly and apply the same matching logic used in fetch_profile:
def get_github_url(profiles):
"""Extract GitHub URL using the same logic as transform.py"""
for p in profiles:
if p.get("network", "").lower() == "github":
return p.get("url", "")
return ""
profiles = resume_data["basics"]["profiles"]
url = get_github_url(profiles)
username = extract_github_username(url) if url else None
Summary
- transform.py contains the
fetch_profilefunction (lines 531-545) that scansbasics.profilesfor GitHub entries using case-insensitive network matching. - The extracted URL is stored in the
github_urlfield of the CSV output, ready for downstream processing. - github.py provides
extract_github_username(lines 116-131) to normalize URLs to plain usernames using regex pattern matching. - The architecture separates URL extraction (
transform.py) from API preparation (github.py), enabling reuse in different contexts. - The implementation supports the standard jsonresume.org format with flexible network name matching.
Frequently Asked Questions
What JSON Resume format does the Hiring-Agent use?
The Hiring-Agent uses the standard jsonresume.org schema, specifically targeting the basics.profiles array where each object contains network, username, and url fields. This format is widely adopted across the industry for structured resume data exchange.
How does fetch_profile handle case sensitivity for network names?
The fetch_profile function in transform.py converts the network field to lowercase using .lower() before comparing against the supplied aliases (e.g., ["github"]). This ensures consistent matching regardless of whether the resume contains "GitHub", "GITHUB", or "github".
Can I extract GitHub URLs without using the full CSV pipeline?
Yes. While transform_resume() provides the complete orchestration, you can import extract_github_username directly from github.py and implement your own profile scanning logic. Simply iterate through basics.profiles and check if network.lower() == "github" to retrieve the URL without invoking the transformation pipeline.
What happens if the GitHub URL is malformed?
The extract_github_username function in github.py applies regex patterns to extract usernames and will return an empty string or None if the URL does not match expected patterns (github.com/username or https://github.com/username). The transform.py fetch logic also defaults to empty strings if no matching profile is found, ensuring the CSV generation continues without errors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →