How the LLM Selects Top 7 GitHub Projects from Repositories in InterviewStreet's Hiring Agent
The LLM selects exactly seven unique projects by analyzing enriched repository metadata serialized in JSON, while the surrounding Python code in github.py enforces the limit through system prompts, deduplication logic, and fallback mechanisms.
The interviewstreet/hiring-agent repository automates candidate technical screening by leveraging a Large Language Model (LLM) to curate the most impressive GitHub projects. This article explains how the system selects exactly seven unique repositories from a candidate's public profile, combining intelligent prompt engineering with robust post-processing safeguards defined in github.py.
Gathering and Enriching Repository Metadata
Fetching Public Repositories
The selection process begins in github.py with the fetch_all_github_repos function (lines 31-76), which retrieves all public repositories for a candidate. This function filters out low-impact forks and constructs a comprehensive list of project dictionaries containing detailed GitHub statistics including stars, forks, topics, and contributor counts.
Serializing Projects for the LLM
The generate_projects_json function (lines 34-44) converts the enriched project list into a JSON string. This serialized data feeds into the github_project_selection prompt template managed by TemplateManager (located in prompts/template_manager.py), preparing the structured context required for the LLM analysis.
LLM Prompt Engineering and Invocation
System Instruction Constraints
The system message explicitly instructs the LLM to "select exactly 7 UNIQUE projects – no duplicates allowed" (lines 81-84 in github.py). This hard constraint ensures the model understands the strict cardinality requirement before processing the repository metadata.
Invoking the Language Model
The code instantiates the LLM provider using initialize_llm_provider (from llm_utils.py) and invokes it via provider.chat(**chat_params) (lines 90-92). The chat parameters include the constrained system instruction, the serialized project JSON, and model-specific options typically defaulting to OpenAI GPT-4.
chat_params = {
"model": DEFAULT_MODEL,
"messages": [
{
"role": "system",
"content": (
"You are an expert technical recruiter analyzing GitHub repositories "
"to identify the most impressive projects. CRITICAL: You must select "
"exactly 7 UNIQUE projects - no duplicates allowed. Each project must be "
"different from the others."
),
},
{"role": "user", "content": prompt},
],
"options": model_params,
}
response = provider.chat(**chat_params)
Response Processing and Uniqueness Enforcement
Parsing the LLM Output
After receiving the response, extract_json_from_response (from llm_utils.py) cleans the raw text and parses it into a Python list of project objects (lines 95-99). The code then iterates through the LLM-selected projects, maintaining a seen_names set to build a unique_projects list (lines 100-108).
Guaranteeing Seven Unique Selections
If the LLM returns fewer than seven unique projects, the system fills the gap by walking the original projects_data list and appending any missing project names until exactly seven distinct entries are collected (lines 110-122). This ensures the final output always meets the cardinality requirement regardless of LLM variance.
# ---- post‑processing ----
selected_projects = json.loads(extract_json_from_response(response["message"]["content"]))
unique_projects = []
seen_names = set()
for proj in selected_projects:
name = proj.get("name")
if name and name not in seen_names:
unique_projects.append(proj)
seen_names.add(name)
# Fill missing slots if LLM gave <7 uniques
for proj in projects_data:
if len(unique_projects) >= 7:
break
name = proj.get("name")
if name and name not in seen_names:
unique_projects.append(proj)
seen_names.add(name)
Fallback Mechanisms for Robustness
Handling Malformed JSON Responses
If the LLM response cannot be parsed as JSON, the function immediately falls back to returning the first seven projects from the original list using projects_data[:7] (lines 132-138). This safety net prevents pipeline failures from invalid LLM outputs.
except json.JSONDecodeError:
print("🔄 Falling back to first 7 projects")
return projects_data[:7]
General Exception Handling
A broader try-except block catches any unexpected exceptions during the selection process, defaulting to the first seven repositories to ensure continuous operation (lines 138-146).
High-Level Usage Example
To execute the selection pipeline for a specific GitHub profile:
from github import fetch_and_display_github_info
# Fetch profile + projects, letting the LLM pick the top 7
result = fetch_and_display_github_info("https://github.com/example-user")
print(result["projects"]) # → list of 7 curated project dicts
Summary
- The system enforces exactly seven unique projects through a combination of LLM instructions and Python post-processing.
- Repository metadata is gathered via
fetch_all_github_reposand serialized usinggenerate_projects_json. - The
github_project_selectionprompt template structures the data for the LLM. - Deduplication uses a
seen_namesset to ensure uniqueness across the final selection. - Robust fallback logic returns the first seven repositories if the LLM response is malformed or incomplete.
Frequently Asked Questions
What happens if the LLM selects fewer than 7 projects?
The code detects incomplete selections and fills remaining slots by iterating through the original projects_data list, appending projects not already in the seen_names set until the count reaches seven distinct entries.
How does the system prevent duplicate project selections?
The code maintains a seen_names set during post-processing, verifying each project name against previous selections before adding it to the unique_projects list, ensuring strict uniqueness in the final output.
Which file contains the core selection logic?
All selection logic resides in github.py, specifically within the functions handling LLM invocation, response parsing, and the deduplication loops between lines 81 and 146.
What model does the hiring agent use for project selection?
The system defaults to OpenAI GPT-4 (or the configured DEFAULT_MODEL) invoked through provider.chat() after initializing the provider via initialize_llm_provider from llm_utils.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →