How to Modify the GitHub Enrichment Logic in the Hiring Agent Repository
To modify the GitHub enrichment logic in InterviewStreet's hiring-agent repository, update the GitHubProfile Pydantic model in models.py, extend the API fetchers in github.py, and adjust the prompt rendering in transform.py to include the new data in the LLM evaluation.
The hiring-agent is an open-source evaluation tool that extracts candidate information from résumés and enriches it with public GitHub profile data before LLM scoring. Modifying the GitHub enrichment logic allows you to customize which repository metrics, profile fields, and activity indicators influence the final hiring evaluation.
Architecture of the GitHub Enrichment Pipeline
The enrichment pipeline follows a three-stage flow: data fetching, model validation, and prompt transformation. Understanding these layers is essential before modifying the logic.
Data Models and Storage
The GitHubProfile class in models.py (line 252) serves as the typed container for all GitHub data. This Pydantic model defines fields including login, name, bio, followers, public_repos, and a nested list of GitHubRepo objects. When you modify the enrichment logic, you must first update this schema to support new attributes.
Fetching and Caching Layer
Raw API calls reside in github.py, specifically within fetch_github_profile() (line 149) and fetch_github_repositories() (line 225). These functions handle GitHub REST API authentication, respect rate-limit headers, and cache raw JSON responses to disk under the cache/ directory. The caching logic spans lines 30‑70, utilizing load_cached_github_data() and write_github_cache() to minimize redundant API requests.
Transformation and Prompt Integration
The transform.py file orchestrates the merge between résumé data and GitHub enrichment. The function add_github_data_to_prompt() (line 658) constructs a markdown block containing the GitHub profile and repository list, which is then appended to the candidate evaluation prompt. This enriched prompt is subsequently consumed by score.py (line 173) for LLM evaluation.
Step-by-Step Guide to Customizing Enrichment
Follow this sequential checklist to safely modify the GitHub enrichment logic without breaking the evaluation pipeline.
-
Clone and install dependencies
git clone https://github.com/interviewstreet/hiring-agent.git cd hiring-agent pip install -r requirements.txt -
Locate the enrichment entry points
Open
github.pyand identifyfetch_github_profile(line 149) andfetch_github_repositories(line 225). These are the primary functions you will extend. -
Update the Pydantic model
Edit
GitHubProfileinmodels.pyto include new fields. For example, to add acompanyattribute:class GitHubProfile(BaseModel): login: str name: Optional[str] bio: Optional[str] followers: int following: int public_repos: int company: Optional[str] = None # New field repos: List[GitHubRepo] = [] -
Extend the API fetchers
In
github.py, populate the new field from the API response:profile = GitHubProfile( login=data["login"], name=data.get("name"), bio=data.get("bio"), followers=data["followers"], following=data["following"], public_repos=data["public_repos"], company=data.get("company"), # New mapping repos=repo_list, ) -
Adjust the caching schema
The existing
write_github_cache()function stores raw JSON, so no changes are required unless you alter the cache file structure itself. -
Modify the prompt rendering
In
transform.py, locate the enrichment block around line 658 and insert your new fields into the markdown template:def add_github_data_to_prompt(resume: dict, profile: GitHubProfile) -> str: github_text = f"GitHub Profile:\n- **User**: {profile.login}\n" if profile.company: github_text += f"- **Company**: {profile.company}\n" # Append to final prompt... return github_text -
Validate changes locally
Run the evaluator against a known GitHub profile to verify your modifications appear in the output:
python main/evaluator.py --github-url https://github.com/torvalds -
Update unit tests
Add assertions in the
tests/directory verifying that new fields persist through the transformation pipeline.
Code Examples for Common Modifications
Adding Repository Activity Metrics
To include recent push activity in the evaluation, modify the repository fetch URL in github.py to sort by latest activity:
api_url = f"https://api.github.com/users/{username}/repos?per_page=100&sort=pushed"
Then update the prompt rendering in transform.py to display the top 5 most recent projects:
recent = sorted(profile.repos, key=lambda r: r.pushed_at, reverse=True)[:5]
github_text += "\n- **Recent Projects**:\n"
for repo in recent:
github_text += f" * {repo.name} (⭐ {repo.stargazers_count}, 🍴 {repo.forks_count})\n"
Changing Cache Timeout Behavior
The default caching mechanism stores data indefinitely. To implement a time-to-live (TTL) strategy, modify the cache validation logic in github.py (lines 30‑70) to check file modification timestamps before returning cached data:
import os
import time
def load_cached_github_data(cache_path: str, max_age_hours: int = 24):
if os.path.exists(cache_path):
if (time.time() - os.path.getmtime(cache_path)) < (max_age_hours * 3600):
with open(cache_path, 'r') as f:
return json.load(f)
return None
Summary
- Source files: Modify
models.pyfor data structure,github.pyfor API fetching, andtransform.pyfor prompt integration. - Key functions:
fetch_github_profile()(line 149),fetch_github_repositories()(line 225), andadd_github_data_to_prompt()(line 658). - Caching: Raw JSON is stored under
cache/by username; no schema changes needed for simple field additions. - Validation: Always test with
python main/evaluator.py --github-url <URL>before committing changes.
Frequently Asked Questions
Where is the GitHub enrichment logic located in the hiring-agent repository?
The logic is distributed across three files: github.py handles API fetching and caching (lines 149‑225), models.py defines the GitHubProfile Pydantic schema (line 252), and transform.py merges the data into evaluation prompts (line 658).
How do I add a new field from the GitHub API to the evaluation prompt?
First, add the field to the GitHubProfile class in models.py. Then update fetch_github_profile() in github.py to extract the value from the API response. Finally, modify add_github_data_to_prompt() in transform.py to render the field in the markdown output sent to the LLM.
Can I disable the GitHub enrichment entirely?
Yes. Remove or comment out the call to add_github_data_to_prompt() in transform.py (line 658). Alternatively, pass an empty GitHubProfile object to the scoring function in score.py (line 173) to evaluate candidates based solely on résumé data.
How does the caching mechanism work for GitHub API calls?
The system caches raw API JSON responses to disk under a cache/ directory, using the GitHub username as the filename. The load_cached_github_data() and write_github_cache() functions in github.py (lines 30‑70) manage this persistence, allowing you to avoid redundant API requests and respect rate limits during development.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →