How to Modify the GitHub Enrichment Logic in the Hiring Agent Repository

To modify the GitHub enrichment logic in InterviewStreet's hiring-agent repository, update the GitHubProfile Pydantic model in models.py, extend the API fetchers in github.py, and adjust the prompt rendering in transform.py to include the new data in the LLM evaluation.

The hiring-agent is an open-source evaluation tool that extracts candidate information from résumés and enriches it with public GitHub profile data before LLM scoring. Modifying the GitHub enrichment logic allows you to customize which repository metrics, profile fields, and activity indicators influence the final hiring evaluation.

Architecture of the GitHub Enrichment Pipeline

The enrichment pipeline follows a three-stage flow: data fetching, model validation, and prompt transformation. Understanding these layers is essential before modifying the logic.

Data Models and Storage

The GitHubProfile class in models.py (line 252) serves as the typed container for all GitHub data. This Pydantic model defines fields including login, name, bio, followers, public_repos, and a nested list of GitHubRepo objects. When you modify the enrichment logic, you must first update this schema to support new attributes.

Fetching and Caching Layer

Raw API calls reside in github.py, specifically within fetch_github_profile() (line 149) and fetch_github_repositories() (line 225). These functions handle GitHub REST API authentication, respect rate-limit headers, and cache raw JSON responses to disk under the cache/ directory. The caching logic spans lines 30‑70, utilizing load_cached_github_data() and write_github_cache() to minimize redundant API requests.

Transformation and Prompt Integration

The transform.py file orchestrates the merge between résumé data and GitHub enrichment. The function add_github_data_to_prompt() (line 658) constructs a markdown block containing the GitHub profile and repository list, which is then appended to the candidate evaluation prompt. This enriched prompt is subsequently consumed by score.py (line 173) for LLM evaluation.

Step-by-Step Guide to Customizing Enrichment

Follow this sequential checklist to safely modify the GitHub enrichment logic without breaking the evaluation pipeline.

  1. Clone and install dependencies

    git clone https://github.com/interviewstreet/hiring-agent.git
    cd hiring-agent
    pip install -r requirements.txt
  2. Locate the enrichment entry points

    Open github.py and identify fetch_github_profile (line 149) and fetch_github_repositories (line 225). These are the primary functions you will extend.

  3. Update the Pydantic model

    Edit GitHubProfile in models.py to include new fields. For example, to add a company attribute:

    class GitHubProfile(BaseModel):
        login: str
        name: Optional[str]
        bio: Optional[str]
        followers: int
        following: int
        public_repos: int
        company: Optional[str] = None  # New field
    
        repos: List[GitHubRepo] = []
  4. Extend the API fetchers

    In github.py, populate the new field from the API response:

    profile = GitHubProfile(
        login=data["login"],
        name=data.get("name"),
        bio=data.get("bio"),
        followers=data["followers"],
        following=data["following"],
        public_repos=data["public_repos"],
        company=data.get("company"),  # New mapping
    
        repos=repo_list,
    )
  5. Adjust the caching schema

    The existing write_github_cache() function stores raw JSON, so no changes are required unless you alter the cache file structure itself.

  6. Modify the prompt rendering

    In transform.py, locate the enrichment block around line 658 and insert your new fields into the markdown template:

    def add_github_data_to_prompt(resume: dict, profile: GitHubProfile) -> str:
        github_text = f"GitHub Profile:\n- **User**: {profile.login}\n"
        if profile.company:
            github_text += f"- **Company**: {profile.company}\n"
        # Append to final prompt...
    
        return github_text
  7. Validate changes locally

    Run the evaluator against a known GitHub profile to verify your modifications appear in the output:

    python main/evaluator.py --github-url https://github.com/torvalds
  8. Update unit tests

    Add assertions in the tests/ directory verifying that new fields persist through the transformation pipeline.

Code Examples for Common Modifications

Adding Repository Activity Metrics

To include recent push activity in the evaluation, modify the repository fetch URL in github.py to sort by latest activity:

api_url = f"https://api.github.com/users/{username}/repos?per_page=100&sort=pushed"

Then update the prompt rendering in transform.py to display the top 5 most recent projects:

recent = sorted(profile.repos, key=lambda r: r.pushed_at, reverse=True)[:5]
github_text += "\n- **Recent Projects**:\n"
for repo in recent:
    github_text += f"  * {repo.name} (⭐ {repo.stargazers_count}, 🍴 {repo.forks_count})\n"

Changing Cache Timeout Behavior

The default caching mechanism stores data indefinitely. To implement a time-to-live (TTL) strategy, modify the cache validation logic in github.py (lines 30‑70) to check file modification timestamps before returning cached data:

import os
import time

def load_cached_github_data(cache_path: str, max_age_hours: int = 24):
    if os.path.exists(cache_path):
        if (time.time() - os.path.getmtime(cache_path)) < (max_age_hours * 3600):
            with open(cache_path, 'r') as f:
                return json.load(f)
    return None

Summary

  • Source files: Modify models.py for data structure, github.py for API fetching, and transform.py for prompt integration.
  • Key functions: fetch_github_profile() (line 149), fetch_github_repositories() (line 225), and add_github_data_to_prompt() (line 658).
  • Caching: Raw JSON is stored under cache/ by username; no schema changes needed for simple field additions.
  • Validation: Always test with python main/evaluator.py --github-url <URL> before committing changes.

Frequently Asked Questions

Where is the GitHub enrichment logic located in the hiring-agent repository?

The logic is distributed across three files: github.py handles API fetching and caching (lines 149‑225), models.py defines the GitHubProfile Pydantic schema (line 252), and transform.py merges the data into evaluation prompts (line 658).

How do I add a new field from the GitHub API to the evaluation prompt?

First, add the field to the GitHubProfile class in models.py. Then update fetch_github_profile() in github.py to extract the value from the API response. Finally, modify add_github_data_to_prompt() in transform.py to render the field in the markdown output sent to the LLM.

Can I disable the GitHub enrichment entirely?

Yes. Remove or comment out the call to add_github_data_to_prompt() in transform.py (line 658). Alternatively, pass an empty GitHubProfile object to the scoring function in score.py (line 173) to evaluate candidates based solely on résumé data.

How does the caching mechanism work for GitHub API calls?

The system caches raw API JSON responses to disk under a cache/ directory, using the GitHub username as the filename. The load_cached_github_data() and write_github_cache() functions in github.py (lines 30‑70) manage this persistence, allowing you to avoid redundant API requests and respect rate limits during development.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →