# Purpose of github.py in the hiring-agent Project: GitHub API Integration and Candidate Profiling

> Discover github.py's role in hiring-agent: extract candidate GitHub data, manage API limits, cache responses, and use LLM analysis for smart tech recruiting. Optimize your hiring!

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-06-30

---

**github.py serves as the central orchestration module that extracts candidate GitHub data, handles API rate limits, caches responses, and leverages LLM-driven analysis to select the most impressive projects for technical recruiting workflows.**

In the `interviewstreet/hiring-agent` repository, [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) acts as the primary bridge between the GitHub API and the hiring agent's evaluation pipeline. This module abstracts all GitHub-related data acquisition, transforming raw repository information into structured, ranked candidate profiles ready for downstream processing.

## Core Responsibilities of github.py

The [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) module manages the complete lifecycle of GitHub data ingestion, from URL parsing to final JSON payload generation.

### Extracting GitHub Usernames from Varied URL Formats

The `extract_github_username` function (lines 16-38) uses regex patterns to parse candidate GitHub URLs and extract clean usernames regardless of whether the input is a profile URL, repository link, or other GitHub domain variation. This ensures robust input handling when recruiters paste varied GitHub link formats.

### Intelligent API Caching and Rate-Limit Management

To prevent redundant network calls and avoid hitting GitHub's rate limits, [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) implements a sophisticated caching layer:

- **`_create_cache_filename`** (lines 18-26) builds deterministic cache paths based on API endpoints
- **`_fetch_github_api`** (lines 29-55, 61-104) reads from and writes to the local cache when `DEVELOPMENT_MODE` is enabled
- **Rate-limit introspection** (lines 55-98) inspects response headers after each request, automatically sleeping until the reset time when remaining requests drop below 10

### Profile and Repository Data Acquisition

The module fetches comprehensive candidate data through two primary functions:

- **`fetch_github_profile`** (lines 41-71) constructs the GitHub API URL, retrieves public user data, and maps it onto the `GitHubProfile` data class defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)
- **`fetch_all_github_repos`** (lines 18-94) iterates through all public repositories, filters out low-impact forks, calls `fetch_repo_contributors` for contribution statistics, and aggregates rich project metadata

### LLM-Powered Project Selection

Rather than returning raw repository lists, [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) uses intelligent ranking through:

- **`generate_projects_json`** (lines 34-70) which serializes repository data and sends it to an LLM via `llm_utils` and `TemplateManager`
- Parsing JSON responses from the LLM and enforcing unique-project rules to prevent duplicate selections
- Returning only the top 7 most technically impressive projects based on the LLM's assessment of code quality, complexity, and relevance

## Key Functions and Implementation Details

The module exposes a clean public API through `fetch_and_display_github_info` (lines 59-78), which orchestrates the entire workflow and returns a combined payload containing both the candidate profile and selected projects as JSON-serializable structures.

For development and debugging, [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) includes a CLI entry point in its `__main__` block (lines 84-97), allowing developers to run `python -m github` to test the pipeline against a hard-coded demo user.

## Integration with the hiring-agent Architecture

[`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) collaborates with several core components to function:

- **[`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)** – Provides the `GitHubProfile` dataclass that structures user information
- **[`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)** – Handles LLM provider initialization and JSON extraction from responses
- **[`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py)** – Loads the `github_project_selection.jinja` template that formats repository data for LLM analysis
- **[`prompt.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompt.py)** – Contains default model configurations and parameters referenced by the GitHub integration

## Usage Examples

Retrieve a candidate's complete GitHub profile and top projects:

```python
from github import fetch_and_display_github_info

result = fetch_and_display_github_info("https://github.com/example-candidate")
print(result["profile"])    # name, bio, followers, company, etc.

print(result["projects"])   # list of 7 LLM-selected projects with metadata

```

Run the module directly for testing:

```bash
python -m github

# Outputs formatted JSON payload for the demo user defined in __main__

```

## Summary

- **[`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)** is the central GitHub integration module in `interviewstreet/hiring-agent`, handling all API communication and data transformation
- **Caching and rate-limiting** are built-in via `_fetch_github_api` to ensure reliable operation without hitting GitHub quotas
- **Data extraction** functions parse usernames from URLs and fetch both profile metadata and complete repository histories
- **LLM integration** ranks projects automatically, selecting the most impressive 7 repositories for recruiter review
- **Clean API** exposed through `fetch_and_display_github_info` returns standardized JSON for downstream hiring workflows

## Frequently Asked Questions

### What is the main entry point for fetching GitHub data in the hiring-agent project?

The `fetch_and_display_github_info` function serves as the primary orchestration method, accepting a GitHub URL and returning a dictionary containing both the candidate profile and selected projects. This function coordinates username extraction, API calls, caching, and LLM processing into a single callable interface.

### How does github.py handle GitHub API rate limits?

The module implements automatic rate-limit detection in `_fetch_github_api` (lines 55-98). After each request, it inspects the `X-RateLimit-Remaining` header, and if fewer than 10 requests remain, calculates the sleep duration until the `X-RateLimit-Reset` timestamp and pauses execution accordingly. Additionally, responses are cached locally when `DEVELOPMENT_MODE` is true to minimize API calls.

### What data structure does github.py return for candidate profiles?

The module returns a dictionary with two keys: `"profile"` containing a `GitHubProfile` dataclass instance (defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)) with user metadata like name, bio, and follower counts, and `"projects"` containing a list of dictionaries representing the top 7 repositories selected by the LLM, including contribution statistics and repository metadata.

### Which files does github.py depend on for LLM processing?

[`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) relies on [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py) for LLM provider initialization and JSON parsing, [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py) for loading the `github_project_selection.jinja` template, and [`prompt.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompt.py) for default model configuration parameters. These dependencies enable the module to send repository data to the LLM and parse structured rankings for project selection.