How GitHub Profile and Repository Enrichment Works in Hiring Agent: A Technical Deep Dive
Hiring Agent enriches candidate résumés by extracting GitHub usernames from URLs, caching API responses to avoid rate limits, classifying repositories based on contributor counts, and leveraging an LLM to curate the seven most impressive projects for evaluation.
The interviewstreet/hiring-agent repository automates technical candidate screening by augmenting résumés with objective, data-driven GitHub signals. The GitHub profile and repository enrichment process transforms raw profile URLs into structured metadata, distinguishing between collaborative open-source contributions and personal projects while minimizing API overhead through deterministic caching.
Step-by-Step GitHub Enrichment Pipeline
The enrichment workflow follows a deterministic sequence designed for reproducibility and minimal API quota consumption.
Extracting the GitHub Username
The pipeline begins by parsing the candidate's résumé for GitHub URLs or plain usernames. In github.py (lines 16‑38), the extract_github_username function uses regular expressions to normalize various URL formats into a canonical login handle.
# github.py, lines 16-38
def extract_github_username(github_url: str) -> Optional[str]:
...
This extraction handles both full URLs (https://github.com/torvalds) and naked usernames, ensuring the downstream API calls receive consistent input.
Cache-Aware API Handling
All GitHub REST requests route through _fetch_github_api in github.py (lines 28‑63). This function implements a file-based caching layer that stores JSON responses under cache/gh_githubcache_… filenames.
# github.py, lines 28-63
def _fetch_github_api(api_url, params=None):
...
The logic checks the cache directory before initiating network requests. If a cached file exists, it returns the stored data; otherwise, it executes the request with an optional Authorization header (when GITHUB_TOKEN is configured), respects GitHub's rate-limit headers, and persists the response to disk. This strategy accelerates development workflows and prevents quota exhaustion during repeated evaluations.
Fetching Profile Metadata
Once the username is canonicalized, fetch_github_profile (lines 41‑78) queries the /users/{username} endpoint and hydrates a Pydantic model defined in models.py (lines 52‑68).
# github.py, lines 41-78
def fetch_github_profile(github_url: str) -> Optional[GitHubProfile]:
...
The GitHubProfile model provides type safety for fields such as name, bio, public_repos, and followers, ensuring downstream components consume validated data structures.
Repository Collection and Classification
The fetch_all_github_repos function (lines 18‑94) orchestrates the heavy lifting of repository analysis. It paginates through https://api.github.com/users/{username}/repos (sorted by recent activity) and performs three critical operations on each repository:
- Fork filtering – Skips trivial forks to focus on original work.
- Contributor analysis – Requests
/repos/{owner}/{repo}/contributorsto count total commits and the author's specific contribution volume. - Repository classification – Labels the repo as open_source (≥ 2 contributors) or self_project (< 2 contributors) based on community participation.
# github.py, lines 18-94
def fetch_all_github_repos(github_url: str, max_repos: int = 100) -> List[Dict]:
...
The function extracts key signals including stars, primary language, topics, and commit counts, returning a uniform list of dictionaries for downstream processing.
LLM-Powered Project Selection
Raw repository lists often contain noise from experimental or minor projects. The generate_projects_json function (excerpted at lines 34‑70) solves this by delegating curation to an LLM.
The system serializes the repository list to JSON and renders it via the github_project_selection.jinja template. The prompt strictly enforces exactly seven unique projects. After parsing the LLM response, the code deduplicates entries and implements fallback logic: if fewer than seven unique items are returned, it pads the selection with the highest-starred remaining repositories from the candidate's profile.
# github.py, lines 34-70 (excerpt of generate_projects_json)
def generate_projects_json(projects: List[Dict]) -> List[Dict]:
...
This hybrid approach combines algorithmic filtering with semantic judgment, ensuring the final selection highlights the candidate's most impactful work.
Data Integration and Output
The fetch_and_display_github_info function (lines 59‑78) serves as the orchestration layer. It aggregates the profile metadata, full repository set, and LLM-curated projects into a standardized dictionary:
"profile"– Plain-dict representation of theGitHubProfilemodel."projects"– The curated list of up to seven selected repositories."total_projects"– Count of all fetched repositories for context.
# github.py, lines 59-78
def fetch_and_display_github_info(github_url: str) -> Dict:
...
Integration with the Résumé Scoring Workflow
In score.py (the CLI entry point), the enrichment pipeline integrates seamlessly with the broader evaluation system. When processing a PDF résumé, the system scans the basics.profiles array for GitHub URLs. If detected, it invokes fetch_and_display_github_info and merges the returned dictionary into the JSONResume structure under the github key.
This enriched data then flows into the evaluator, which scores candidates on open-source engagement and project quality. The Architecture section of the repository's README outlines this flow as part of the end-to-end pipeline (PDF → LLM → GitHub → Evaluation).
Practical Implementation Examples
Stand-Alone Enrichment
You can invoke the enrichment pipeline independently for testing or custom integrations:
from github import fetch_and_display_github_info
result = fetch_and_display_github_info("https://github.com/torvalds")
print(result["profile"]["name"]) # → Linus Torvalds
print("Top projects:")
for p in result["projects"]:
print(f"- {p['name']} ({p['github_details']['stars']} ★)")
Full Pipeline Integration
The following excerpt demonstrates how score.py embeds GitHub data into the evaluation workflow:
from github import fetch_and_display_github_info
from pdf import PDFHandler # parses the PDF → JSONResume
from evaluator import evaluate_resume
pdf_path = "candidate_resume.pdf"
resume_json = PDFHandler().process(pdf_path)
# Enrich with GitHub data if a profile URL is present
github_url = next(
(p.url for p in resume_json.basics.profiles or [] if p.network == "GitHub"), None
)
if github_url:
gh_data = fetch_and_display_github_info(github_url)
resume_json.github = gh_data # attaches profile & projects
evaluation = evaluate_resume(resume_json)
print(evaluation.scores.open_source.score)
Command-Line Usage
The repository ships with a CLI entry point for direct execution:
# Command-line usage
$ python score.py path/to/resume.pdf
# → prints the enriched JSON and a human-readable scoring summary
Summary
- Username extraction normalizes GitHub URLs into canonical handles using regex patterns in
github.py. - Intelligent caching via
_fetch_github_apieliminates redundant network calls and respects GitHub rate limits. - Repository classification distinguishes open_source projects (≥ 2 contributors) from self_project work based on contributor analysis.
- LLM curation selects exactly seven impressive projects using the
github_project_selection.jinjatemplate, with algorithmic fallback to star-count ranking. - Seamless integration via
fetch_and_display_github_infoattaches structured GitHub data toJSONResumeobjects for downstream scoring inscore.py.
Frequently Asked Questions
How does Hiring Agent handle GitHub API rate limits?
The _fetch_github_api function in github.py implements a file-based cache under the cache/ directory. Before making any request, it checks for a cached response stored as cache/gh_githubcache_…. If found, it returns the cached JSON immediately. For fresh requests, it respects GitHub's rate-limit headers and supports authentication via the GITHUB_TOKEN environment variable to increase quota limits.
What criteria determine if a repository is classified as open source?
A repository is classified as open_source when the contributors endpoint reveals two or more distinct contributors. The fetch_all_github_repos function queries /repos/{owner}/{repo}/contributors and counts unique authors. Repositories with only one contributor (typically the candidate) are labeled self_project, distinguishing collaborative work from personal experiments.
Why does the enrichment process use an LLM for project selection?
The LLM serves as a semantic filter to identify the "most impressive" projects from potentially lengthy repository lists. The generate_projects_json function serializes repository metadata (stars, language, topics, commit counts) and prompts the model via the github_project_selection.jinja template to select exactly seven unique items. This combines quantitative metrics with qualitative judgment, ensuring the final selection highlights high-impact work rather than just the most recent commits.
Can the GitHub enrichment run independently of the full scoring pipeline?
Yes. The fetch_and_display_github_info function in github.py operates as a standalone entry point. You can import and call it directly with any GitHub URL to receive a structured dictionary containing the profile metadata and curated project list without invoking the PDF parser or evaluation engine. This modularity supports debugging, testing, and custom integrations outside the standard score.py workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →