# How the Hiring Agent Fetches and Classifies Data from GitHub Profiles and Repositories

> Discover how the hiring agent fetches and classifies GitHub data. It centralizes API calls, caches raw data, and uses LLMs to select seven high-impact projects.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-06-27

---

**The hiring agent centralizes GitHub API calls in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) to cache raw profile and repository data, then uses LLM-driven prompting in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) to classify and select exactly seven high-impact projects based on stars, language diversity, and relevance.**

The `interviewstreet/hiring-agent` repository automates technical candidate screening by aggregating public coding artifacts. Understanding how it fetches and classifies data from GitHub profiles and repositories reveals a resilient architecture that combines deterministic API handling with intelligent filtering to surface the most relevant engineering work.

## Centralized API Handling and Caching Strategy

All GitHub network operations are isolated in **[`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)** to ensure consistent error handling and observability. The private helper `_fetch_github_api()` manages every outgoing request to `api.github.com`, injecting an `Authorization: token …` header when the `GITHUB_TOKEN` environment variable is present.

Before making a network call, the system computes a deterministic cache key using `_create_cache_filename()`, which hashes the endpoint and parameters to produce a file path under the `cache/` directory (e.g., [`cache/gh_githubcache_users_username.json`](https://github.com/interviewstreet/hiring-agent/blob/main/cache/gh_githubcache_users_username.json)). When `DEVELOPMENT_MODE` is enabled, the function short-circuits to the cached file if it exists, eliminating redundant API traffic during local testing.

Rate-limit protection is implemented by inspecting the `X-RateLimit-Remaining`, `X-RateLimit-Limit`, and `X-RateLimit-Reset` headers after every response. If fewer than ten requests remain, the client sleeps until the reset window expires, logging the backoff duration for observability.

## Profile and Repository Retrieval

Two public functions expose GitHub data to the rest of the pipeline. **`fetch_github_profile(github_url)`** parses the username from the provided URL and queries the API:

```python
def fetch_github_profile(github_url: str) -> Optional[GitHubProfile]:
    username = github_url.rstrip("/").split("/")[-1]
    api_url = f"https://api.github.com/users/{username}"
    status, data = _fetch_github_api(api_url)
    if status != 200:
        return None
    return GitHubProfile(**data)  # Pydantic validation

```

This function validates the JSON response against the **`GitHubProfile`** Pydantic model defined in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), enforcing fields such as `name`, `bio`, `followers`, `following`, and `public_repos`.

For repository analysis, **`fetch_github_repositories(github_url)`** hits the `/users/<username>/repos` endpoint and paginates through results. It enriches each repository object with contributor statistics from `/contributors`, capturing `stargazers_count`, `forks_count`, and the usernames of top contributors. Both functions return `None` on any network or validation failure, allowing the evaluation pipeline to continue gracefully even when GitHub is unreachable.

## LLM-Powered Classification Pipeline

Once raw data is cached and validated, **[`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py)** orchestrates the classification phase. It injects the serialized `GitHubProfile` and repository list into the Jinja template located at `prompts/templates/github_project_selection.jinja`. This template encodes strict instructions for the LLM: select **exactly seven distinct, high-impact projects** and rank them according to stars, language diversity, recent commit activity, and relevance to the target job description.

The populated prompt is sent to the LLM provider (initialized via [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)), and the resulting text is parsed by **`extract_json_from_response()`**, which isolates the JSON array of selected projects from any surrounding conversational text. The structured output is then merged into the candidate’s master profile dictionary, where it feeds into downstream scoring logic.

```python
from github import fetch_github_profile, fetch_github_repositories
from transform import build_candidate_profile

# 1️⃣ Fetch raw metadata

profile = fetch_github_profile("https://github.com/example_user")
repos   = fetch_github_repositories("https://github.com/example_user")

# 2️⃣ Classify and select top projects via LLM

candidate = build_candidate_profile(
    basic_info={"name": "Jane Doe", "role": "Backend Engineer"},
    github_profile=profile,
    github_repos=repos
)

print(candidate["selected_projects"])

# → [{'name': 'distributed-system', 'stars': 1200, 'language': 'Rust'}, ...]  # Exactly 7 items

```

## Resilience and Rate Limit Management

The system is designed to degrade gracefully under API stress. Every network call in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) is wrapped in a `try/except` block that logs the exception and returns `None` rather than raising fatal errors. This ensures that a temporary GitHub outage or rate-limit exhaustion does not crash an ongoing evaluation batch.

The caching layer provides an additional resilience mechanism: once a profile or repository list is stored in `cache/`, subsequent invocations replay the data from disk, bypassing the API entirely. This is particularly useful for reprocessing candidates or debugging classification prompts without consuming additional rate-limit quota.

## Summary

- **[`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)** serves as the single source of truth for all GitHub API communication, implementing deterministic caching in `cache/` and proactive rate-limit backoff when `X-RateLimit-Remaining` drops below ten.
- **Profile retrieval** uses `fetch_github_profile()` to validate user metadata against the `GitHubProfile` Pydantic model in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), while `fetch_github_repositories()` enriches project data with contributor metrics.
- **LLM classification** in [`transform.py`](https://github.com/interviewstreet/hiring-agent/blob/main/transform.py) uses Jinja templates to prompt the model for exactly seven ranked projects, parsing the response with `extract_json_from_response()` from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py).
- **Error resilience** is achieved through `try/except` wrappers that return `None` on failure, allowing the hiring pipeline to continue operating even when GitHub data is unavailable.

## Frequently Asked Questions

### How does the hiring agent handle GitHub API rate limits?

The agent inspects the `X-RateLimit-Remaining` header after every request in `_fetch_github_api()`. When fewer than ten requests remain, it sleeps until the timestamp specified in `X-RateLimit-Reset` before continuing, preventing hard rate-limit errors that would disrupt batch processing.

### What specific criteria does the LLM use to classify repositories?

According to the prompt template in `prompts/templates/github_project_selection.jinja`, the LLM must select exactly seven unique projects and rank them based on **stargazers_count**, programming language diversity, recency of commits, and topical relevance to the job description provided in the evaluation context.

### How is the GitHub data validated before classification?

Raw API responses are validated against the **`GitHubProfile`** Pydantic model in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), which type-checks fields like `public_repos`, `followers`, and `bio`. Repository data undergoes similar structural checks, and any validation failure causes the fetch function to return `None`, signaling the pipeline to skip that data source.

### Can the agent operate offline using cached data?

Yes. When `DEVELOPMENT_MODE` is enabled, the `_fetch_github_api()` function checks for an existing file in `cache/` matching the request hash before making a network call. If the cache file exists, it is loaded and returned immediately, allowing full pipeline runs without internet connectivity or API quota consumption.