How the LLM Selects Top 7 GitHub Projects from Repositories in InterviewStreet's Hiring Agent

The LLM selects exactly seven unique projects by analyzing enriched repository metadata serialized in JSON, while the surrounding Python code in github.py enforces the limit through system prompts, deduplication logic, and fallback mechanisms.

The interviewstreet/hiring-agent repository automates candidate technical screening by leveraging a Large Language Model (LLM) to curate the most impressive GitHub projects. This article explains how the system selects exactly seven unique repositories from a candidate's public profile, combining intelligent prompt engineering with robust post-processing safeguards defined in github.py.

Gathering and Enriching Repository Metadata

Fetching Public Repositories

The selection process begins in github.py with the fetch_all_github_repos function (lines 31-76), which retrieves all public repositories for a candidate. This function filters out low-impact forks and constructs a comprehensive list of project dictionaries containing detailed GitHub statistics including stars, forks, topics, and contributor counts.

Serializing Projects for the LLM

The generate_projects_json function (lines 34-44) converts the enriched project list into a JSON string. This serialized data feeds into the github_project_selection prompt template managed by TemplateManager (located in prompts/template_manager.py), preparing the structured context required for the LLM analysis.

LLM Prompt Engineering and Invocation

System Instruction Constraints

The system message explicitly instructs the LLM to "select exactly 7 UNIQUE projects – no duplicates allowed" (lines 81-84 in github.py). This hard constraint ensures the model understands the strict cardinality requirement before processing the repository metadata.

Invoking the Language Model

The code instantiates the LLM provider using initialize_llm_provider (from llm_utils.py) and invokes it via provider.chat(**chat_params) (lines 90-92). The chat parameters include the constrained system instruction, the serialized project JSON, and model-specific options typically defaulting to OpenAI GPT-4.

chat_params = {
    "model": DEFAULT_MODEL,
    "messages": [
        {
            "role": "system",
            "content": (
                "You are an expert technical recruiter analyzing GitHub repositories "
                "to identify the most impressive projects. CRITICAL: You must select "
                "exactly 7 UNIQUE projects - no duplicates allowed. Each project must be "
                "different from the others."
            ),
        },
        {"role": "user", "content": prompt},
    ],
    "options": model_params,
}
response = provider.chat(**chat_params)

Response Processing and Uniqueness Enforcement

Parsing the LLM Output

After receiving the response, extract_json_from_response (from llm_utils.py) cleans the raw text and parses it into a Python list of project objects (lines 95-99). The code then iterates through the LLM-selected projects, maintaining a seen_names set to build a unique_projects list (lines 100-108).

Guaranteeing Seven Unique Selections

If the LLM returns fewer than seven unique projects, the system fills the gap by walking the original projects_data list and appending any missing project names until exactly seven distinct entries are collected (lines 110-122). This ensures the final output always meets the cardinality requirement regardless of LLM variance.


# ---- post‑processing ----

selected_projects = json.loads(extract_json_from_response(response["message"]["content"]))
unique_projects = []
seen_names = set()
for proj in selected_projects:
    name = proj.get("name")
    if name and name not in seen_names:
        unique_projects.append(proj)
        seen_names.add(name)

# Fill missing slots if LLM gave <7 uniques

for proj in projects_data:
    if len(unique_projects) >= 7:
        break
    name = proj.get("name")
    if name and name not in seen_names:
        unique_projects.append(proj)
        seen_names.add(name)

Fallback Mechanisms for Robustness

Handling Malformed JSON Responses

If the LLM response cannot be parsed as JSON, the function immediately falls back to returning the first seven projects from the original list using projects_data[:7] (lines 132-138). This safety net prevents pipeline failures from invalid LLM outputs.

except json.JSONDecodeError:
    print("🔄 Falling back to first 7 projects")
    return projects_data[:7]

General Exception Handling

A broader try-except block catches any unexpected exceptions during the selection process, defaulting to the first seven repositories to ensure continuous operation (lines 138-146).

High-Level Usage Example

To execute the selection pipeline for a specific GitHub profile:

from github import fetch_and_display_github_info

# Fetch profile + projects, letting the LLM pick the top 7

result = fetch_and_display_github_info("https://github.com/example-user")
print(result["projects"])   # → list of 7 curated project dicts

Summary

  • The system enforces exactly seven unique projects through a combination of LLM instructions and Python post-processing.
  • Repository metadata is gathered via fetch_all_github_repos and serialized using generate_projects_json.
  • The github_project_selection prompt template structures the data for the LLM.
  • Deduplication uses a seen_names set to ensure uniqueness across the final selection.
  • Robust fallback logic returns the first seven repositories if the LLM response is malformed or incomplete.

Frequently Asked Questions

What happens if the LLM selects fewer than 7 projects?

The code detects incomplete selections and fills remaining slots by iterating through the original projects_data list, appending projects not already in the seen_names set until the count reaches seven distinct entries.

How does the system prevent duplicate project selections?

The code maintains a seen_names set during post-processing, verifying each project name against previous selections before adding it to the unique_projects list, ensuring strict uniqueness in the final output.

Which file contains the core selection logic?

All selection logic resides in github.py, specifically within the functions handling LLM invocation, response parsing, and the deduplication loops between lines 81 and 146.

What model does the hiring agent use for project selection?

The system defaults to OpenAI GPT-4 (or the configured DEFAULT_MODEL) invoked through provider.chat() after initializing the provider via initialize_llm_provider from llm_utils.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →