# How the LLM Selects Top 7 GitHub Projects from Repositories in InterviewStreet's Hiring Agent

> Discover how InterviewStreet's LLM selects top 7 GitHub projects by analyzing JSON metadata and Python code. Learn about the deduplication and fallback logic.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: internals
- Published: 2026-07-08

---

**The LLM selects exactly seven unique projects by analyzing enriched repository metadata serialized in JSON, while the surrounding Python code in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) enforces the limit through system prompts, deduplication logic, and fallback mechanisms.**

The `interviewstreet/hiring-agent` repository automates candidate technical screening by leveraging a Large Language Model (LLM) to curate the most impressive GitHub projects. This article explains how the system selects exactly seven unique repositories from a candidate's public profile, combining intelligent prompt engineering with robust post-processing safeguards defined in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py).

## Gathering and Enriching Repository Metadata

### Fetching Public Repositories

The selection process begins in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py) with the `fetch_all_github_repos` function (lines 31-76), which retrieves all public repositories for a candidate. This function filters out low-impact forks and constructs a comprehensive list of project dictionaries containing detailed GitHub statistics including stars, forks, topics, and contributor counts.

### Serializing Projects for the LLM

The `generate_projects_json` function (lines 34-44) converts the enriched project list into a JSON string. This serialized data feeds into the `github_project_selection` prompt template managed by `TemplateManager` (located in [`prompts/template_manager.py`](https://github.com/interviewstreet/hiring-agent/blob/main/prompts/template_manager.py)), preparing the structured context required for the LLM analysis.

## LLM Prompt Engineering and Invocation

### System Instruction Constraints

The system message explicitly instructs the LLM to **"select exactly 7 UNIQUE projects – no duplicates allowed"** (lines 81-84 in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py)). This hard constraint ensures the model understands the strict cardinality requirement before processing the repository metadata.

### Invoking the Language Model

The code instantiates the LLM provider using `initialize_llm_provider` (from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)) and invokes it via `provider.chat(**chat_params)` (lines 90-92). The chat parameters include the constrained system instruction, the serialized project JSON, and model-specific options typically defaulting to OpenAI GPT-4.

```python
chat_params = {
    "model": DEFAULT_MODEL,
    "messages": [
        {
            "role": "system",
            "content": (
                "You are an expert technical recruiter analyzing GitHub repositories "
                "to identify the most impressive projects. CRITICAL: You must select "
                "exactly 7 UNIQUE projects - no duplicates allowed. Each project must be "
                "different from the others."
            ),
        },
        {"role": "user", "content": prompt},
    ],
    "options": model_params,
}
response = provider.chat(**chat_params)

```

## Response Processing and Uniqueness Enforcement

### Parsing the LLM Output

After receiving the response, `extract_json_from_response` (from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)) cleans the raw text and parses it into a Python list of project objects (lines 95-99). The code then iterates through the LLM-selected projects, maintaining a `seen_names` set to build a `unique_projects` list (lines 100-108).

### Guaranteeing Seven Unique Selections

If the LLM returns fewer than seven unique projects, the system fills the gap by walking the original `projects_data` list and appending any missing project names until exactly seven distinct entries are collected (lines 110-122). This ensures the final output always meets the cardinality requirement regardless of LLM variance.

```python

# ---- post‑processing ----

selected_projects = json.loads(extract_json_from_response(response["message"]["content"]))
unique_projects = []
seen_names = set()
for proj in selected_projects:
    name = proj.get("name")
    if name and name not in seen_names:
        unique_projects.append(proj)
        seen_names.add(name)

# Fill missing slots if LLM gave <7 uniques

for proj in projects_data:
    if len(unique_projects) >= 7:
        break
    name = proj.get("name")
    if name and name not in seen_names:
        unique_projects.append(proj)
        seen_names.add(name)

```

## Fallback Mechanisms for Robustness

### Handling Malformed JSON Responses

If the LLM response cannot be parsed as JSON, the function immediately falls back to returning the first seven projects from the original list using `projects_data[:7]` (lines 132-138). This safety net prevents pipeline failures from invalid LLM outputs.

```python
except json.JSONDecodeError:
    print("🔄 Falling back to first 7 projects")
    return projects_data[:7]

```

### General Exception Handling

A broader try-except block catches any unexpected exceptions during the selection process, defaulting to the first seven repositories to ensure continuous operation (lines 138-146).

## High-Level Usage Example

To execute the selection pipeline for a specific GitHub profile:

```python
from github import fetch_and_display_github_info

# Fetch profile + projects, letting the LLM pick the top 7

result = fetch_and_display_github_info("https://github.com/example-user")
print(result["projects"])   # → list of 7 curated project dicts

```

## Summary

- The system enforces **exactly seven unique projects** through a combination of LLM instructions and Python post-processing.
- Repository metadata is gathered via `fetch_all_github_repos` and serialized using `generate_projects_json`.
- The `github_project_selection` prompt template structures the data for the LLM.
- Deduplication uses a `seen_names` set to ensure uniqueness across the final selection.
- Robust fallback logic returns the first seven repositories if the LLM response is malformed or incomplete.

## Frequently Asked Questions

### What happens if the LLM selects fewer than 7 projects?

The code detects incomplete selections and fills remaining slots by iterating through the original `projects_data` list, appending projects not already in the `seen_names` set until the count reaches seven distinct entries.

### How does the system prevent duplicate project selections?

The code maintains a `seen_names` set during post-processing, verifying each project name against previous selections before adding it to the `unique_projects` list, ensuring strict uniqueness in the final output.

### Which file contains the core selection logic?

All selection logic resides in [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py), specifically within the functions handling LLM invocation, response parsing, and the deduplication loops between lines 81 and 146.

### What model does the hiring agent use for project selection?

The system defaults to OpenAI GPT-4 (or the configured `DEFAULT_MODEL`) invoked through `provider.chat()` after initializing the provider via `initialize_llm_provider` from [`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py).