# How the Hiring Agent Analyzes GitHub Profiles: From URL to Structured JSON

> Discover how the Hiring Agent analyzes GitHub profiles, converting URLs to structured JSON by extracting data, API calls, and LLM project selection for impressive results.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-06-28

---

**The Hiring Agent converts public GitHub URLs into structured JSON by extracting usernames, fetching profile data via the GitHub API, aggregating repository statistics, and using an LLM to select the seven most impressive projects.**

The interviewstreet/hiring-agent repository provides an automated pipeline to analyze GitHub profiles and transform raw public data into evaluatable candidate intelligence. This system orchestrates multiple API calls, contribution analysis, and large language model inference to generate structured representations of developer portfolios.

## The Six-Stage Analysis Pipeline

The complete workflow is implemented in [`main/github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/github.py) and processes any public GitHub URL through six distinct stages.

### Stage 1: Username Normalization

The `extract_github_username` function (line 16) normalizes diverse GitHub URL formats—including full URLs, shorthand references, and email-style identifiers—into clean usernames using regular expressions. This resilient parsing ensures the system handles inputs like `https://github.com/torvalds`, `github.com/torvalds`, or just `torvalds` without manual intervention.

### Stage 2: Profile Retrieval

The `fetch_github_profile` function (line 41) constructs the GitHub REST endpoint `https://api.github.com/users/<username>` and delegates the HTTP request to `_fetch_github_api`. The JSON response is validated and mapped onto the Pydantic `GitHubProfile` model defined in [`main/models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/models.py), ensuring type safety and consistent data structure.

### Stage 3: Repository Aggregation

`fetch_all_github_repos` (line 18) retrieves the complete repository list via `/users/<username>/repos` with automatic pagination handling. The system filters out low-impact forks (specifically those with fewer than five stars) and enriches each repository with:

- Contributor lists via `fetch_repo_contributors`
- Commit statistics including `author_commit_count` (owner contributions) and `total_commit_count` (project-wide) via `fetch_contributions_count`

### Stage 4: Intelligent Project Selection

The raw repository list is serialized to JSON and passed to `generate_projects_json` (line 34), which loads the `github_project_selection` Jinja template from [`main/prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/prompts/template_manager.py). The LLM—initialized through utilities in [`main/llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/llm_utils.py)—processes this prompt to select exactly seven unique, high-impact projects, explicitly filtering out duplicates and generic repositories based on the system-level instructions.

### Stage 5: Data Synthesis

The `fetch_and_display_github_info` function (line 59) combines outputs from `generate_profile_json` and `generate_projects_json` into a unified dictionary containing:

- Profile metadata
- Curated project list
- `total_projects` count (fixed at 7 when successful)

### Stage 6: Resilient API Handling

All GitHub API traffic routes through `_fetch_github_api` (line 29), which implements resilient request handling:

- **Authentication**: Respects optional `GITHUB_TOKEN` environment variable for higher rate limits
- **Caching**: Stores successful responses in a local `cache/` directory when `DEVELOPMENT_MODE` is enabled in [`main/config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/config.py)
- **Rate limiting**: Monitors `X-RateLimit-Remaining` headers and proactively sleeps when approaching quota limits

## Working with the GitHub Analysis Module

Basic usage requires only a GitHub URL:

```python
from github import fetch_and_display_github_info

result = fetch_and_display_github_info("https://github.com/torvalds")
print(result["profile"]["name"])          # → Linus Torvalds

print(result["projects"][0]["name"])      # → First selected project

```

Command-line execution:

```bash
python -m main.github "https://github.com/realpython"

```

The output format follows this structure:

```json
{
  "profile": { "name": "...", "bio": "..." },
  "projects": [
    { "name": "...", "description": "...", "stars": 1500 }
  ],
  "total_projects": 7
}

```

Programmatic integration for custom workflows:

```python
from github import fetch_github_profile, fetch_all_github_repos

profile = fetch_github_profile("https://github.com/psf")
repos = fetch_all_github_repos("psf", max_repos=50)

# Process repos list for custom visualization or scoring

```

## Architecture Overview

The analysis pipeline relies on several modular components:

- **[`main/models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/models.py)**: Defines Pydantic data models (`GitHubProfile`, `JSONResume`) for validation
- **[`main/llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/llm_utils.py)**: Handles LLM provider initialization (Ollama/Gemini) and JSON extraction
- **[`main/prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/prompts/template_manager.py)**: Loads Jinja templates including `github_project_selection`
- **[`main/config.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/config.py)**: Contains `DEVELOPMENT_MODE` flag controlling caching behavior

## Summary

- The Hiring Agent analyzes GitHub profiles through a six-stage pipeline implemented in [`main/github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/github.py)
- **Username extraction** handles multiple URL formats via `extract_github_username` (line 16)
- **Profile data** is fetched from the GitHub REST API and validated against Pydantic models
- **Repository aggregation** filters low-value forks (< 5 stars) and computes contribution statistics
- **LLM selection** uses the `github_project_selection` template to identify the top seven unique projects
- **Rate limiting** and caching in `_fetch_github_api` (line 29) ensure reliable operation against API quotas

## Frequently Asked Questions

### How does the Hiring Agent handle different GitHub URL formats?

The `extract_github_username` function at line 16 in [`main/github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/github.py) uses regular expressions to normalize various inputs—including full HTTPS URLs, shorthand references, and email-style identifiers—into clean GitHub usernames before API calls are made.

### What criteria does the LLM use to select the top seven projects?

The system prompts the LLM with the `github_project_selection` template to identify the seven most impressive and unique repositories, explicitly excluding duplicate or generic projects while considering factors like star count, contribution depth, and project uniqueness.

### How does the system prevent GitHub API rate limit errors?

All requests pass through `_fetch_github_api` at line 29, which monitors `X-RateLimit-Remaining` headers and proactively sleeps when quotas are low, supports optional `GITHUB_TOKEN` authentication for higher limits, and caches responses locally when `DEVELOPMENT_MODE` is enabled.

### Can I analyze GitHub profiles programmatically without the LLM selection?

Yes, you can import `fetch_github_profile` and `fetch_all_github_repos` directly from [`main/github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/main/github.py) to retrieve raw profile data and repository lists, bypassing the LLM-based project ranking stage for custom analysis workflows.