GitHub Enrichment Pipeline Architecture in the Hiring Agent
The GitHub enrichment pipeline in the interviewstreet/hiring-agent repository extracts a candidate's username from their resume, fetches profile and repository data via the GitHub API, filters and classifies projects based on contributor counts, uses an LLM to select the top seven most relevant repositories, and assembles everything into a structured enrichment payload for downstream evaluation.
The interviewstreet/hiring-agent project automates candidate evaluation by enriching resume data with public GitHub signals. Understanding the GitHub enrichment pipeline architecture reveals how raw profile URLs transform into structured, LLM-curated project portfolios that power fair, data-driven hiring decisions.
Five-Stage Pipeline Architecture
The enrichment process implemented in github.py follows a deterministic five-stage flow, moving from raw URL parsing to structured data assembly.
Stage 1: Username Extraction
The pipeline begins with extract_github_username() in github.py (lines 16-28), which parses any GitHub URL found in the candidate's resume to isolate the plain username handle. This normalizes various URL formats—whether the candidate provided github.com/username or a full profile path—into a consistent identifier for API calls.
Stage 2: Profile Fetching
Using the extracted handle, fetch_github_profile() (lines 41-73) queries the GitHub Users API at https://api.github.com/users/<username>. The raw JSON response is validated and wrapped in a GitHubProfile Pydantic model defined in models.py, ensuring type safety for downstream consumers.
Stage 3: Repository Retrieval
The fetch_all_github_repos() function (lines 18-88) orchestrates repository collection by hitting the /users/<username>/repos endpoint. For each repository returned, the pipeline:
- Filters out low-impact forks with fewer than 5 forks
- Calls
fetch_repo_contributors()to retrieve contributor metadata - Calculates commit ratios via
fetch_contributions_count()to determine the author's share of total activity - Classifies the project as
open_source(≥2 contributors) orself_projectbased on contributor thresholds
This classification distinguishes between collaborative community work and individual portfolio pieces.
Stage 4: LLM-Driven Project Selection
Rather than returning all repositories, the pipeline uses generate_projects_json() (lines 34-68) to curate exactly seven unique projects. The function serializes the filtered repository list into JSON and passes it to the LLM via the github_project_selection.jinja template. The model ranks projects by relevance and impact, returning a strictly validated list. If the LLM response contains duplicates or parsing errors, the wrapper automatically deduplicates entries and falls back to the first seven items from the raw list.
Stage 5: Payload Assembly
Finally, fetch_and_display_github_info() (lines 59-78) combines the profile data from generate_profile_json() with the LLM-selected projects into a single dictionary. This enrichment payload is consumed by downstream stages such as evaluator.py for candidate scoring.
Caching and Rate Limit Protection
The pipeline includes defensive mechanisms against GitHub API constraints. The _create_cache_filename() helper implements a caching layer that stores raw API responses under the cache/ directory when the DEVELOPMENT_MODE flag is enabled. This avoids redundant network calls during iterative development.
For production runs, the code monitors API rate limits in real time. When remaining requests drop below a safety threshold, the pipeline automatically sleeps until the GitHub API's reset time, preventing service interruptions during bulk candidate processing.
Core Implementation Files
| File | Responsibility |
|---|---|
github.py |
Core pipeline implementation including username extraction, API orchestration, and LLM project selection |
models.py |
Pydantic schemas (GitHubProfile, LLMProvider abstractions) |
llm_utils.py |
Provider-agnostic LLM initialization and response cleanup |
prompts/templates/github_project_selection.jinja |
Template defining the LLM's ranking criteria and output format |
prompt.py |
Central configuration for DEFAULT_MODEL and LLM_PROVIDER settings |
All LLM interactions are abstracted through llm_utils.py and models.py, allowing the pipeline to switch between Ollama, Gemini, or future providers without code changes.
Practical Code Examples
Complete Enrichment Workflow
from github import fetch_and_display_github_info
# Pass the GitHub URL found in the resume (or just the handle)
result = fetch_and_display_github_info("https://github.com/example_user")
print(result["profile"]) # GitHubProfile fields
print(result["projects"]) # List of 7 LLM-selected projects
Low-Level Pipeline Access
from github import (
extract_github_username,
fetch_github_profile,
fetch_all_github_repos,
)
url = "https://github.com/example_user"
username = extract_github_username(url)
profile = fetch_github_profile(url) # → GitHubProfile instance
repos = fetch_all_github_repos(url, 50) # → enriched repo dicts
# Inspect classification
print(repos[0]["name"], repos[0]["project_type"])
Internal LLM Selection Logic
from github import generate_projects_json
# Assuming `projects` is the list from fetch_all_github_repos()
selected = generate_projects_json(projects) # → exactly 7 unique items
Summary
- The pipeline extracts GitHub usernames from resume URLs using
extract_github_username()ingithub.py - Profile data is fetched via the GitHub Users API and validated through Pydantic models
- Repository collection filters forks (<5) and classifies projects as
open_sourceorself_projectbased on contributor counts - An LLM selects exactly seven unique projects using the
github_project_selection.jinjatemplate, with automatic fallback handling - The final payload combines profile and project data for consumption by
evaluator.py - Built-in caching and rate-limit handling ensure reliable operation during high-volume processing
Frequently Asked Questions
How does the pipeline handle GitHub API rate limits?
The code monitors the API's remaining request count and reset timestamp. When approaching the limit, it pauses execution and sleeps until the reset time, ensuring the pipeline resumes automatically without manual intervention or failed requests.
What distinguishes an open_source project from a self_project?
The classification depends on contributor count. Repositories with two or more contributors are labeled open_source, indicating collaborative community work. Projects with fewer than two contributors are classified as self_project, representing individual portfolio work.
How does the LLM ensure exactly seven projects are returned?
The generate_projects_json() function enforces this constraint through the github_project_selection.jinja template, which explicitly instructs the model to return seven unique projects. The wrapper validates the response, deduplicates entries if necessary, and falls back to the first seven raw repositories if the LLM output is malformed.
Can the pipeline run offline during development?
Yes. When DEVELOPMENT_MODE is enabled, the _create_cache_filename() mechanism stores raw GitHub API responses in the cache/ directory. Subsequent runs load data from this cache instead of hitting the live API, enabling rapid iteration without network dependencies or rate limit concerns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →