How the Hiring Agent Selects the Top 7 GitHub Projects for Evaluation

The Hiring Agent filters repositories by a minimum author commit count of 4, then applies an LLM-driven ranking based on contribution volume, project popularity, and technical complexity to identify exactly 7 unique projects for evaluation.

The interviewstreet/hiring-agent repository automates technical candidate screening by extracting and analyzing GitHub portfolios. Understanding how the system selects the top 7 GitHub projects for evaluation reveals a multi-stage pipeline that combines deterministic thresholds with AI-powered qualitative assessment.

Data Collection and Metric Extraction

The selection process begins in github.py, where the fetch_all_github_repos function gathers every public repository associated with the extracted username. For each repository, the system computes author_commit_count—the definitive metric representing commits authored by the candidate—alongside auxiliary data including stars, forks, primary language, and topics.

The generate_projects_json function (lines 34-57) transforms this raw data into a structured JSON list. This normalization ensures the downstream ranking logic receives consistent inputs containing both quantitative activity metrics and qualitative project metadata.

Hard Contribution Thresholds

Before any ranking occurs, the system enforces a strict eligibility rule: only projects with author_commit_count >= 4 qualify for consideration. This threshold eliminates forks, template repositories, and superficial contributions.

Enforcement occurs at two independent layers for redundancy:

  • Template layer: The Jinja prompt explicitly filters out repositories below this threshold (lines 44-48 in prompts/templates/github_project_selection.jinja)
  • LLM system prompt: The system message reiterates the constraint (lines 55-57), instructing the model to disregard low-participation projects even if present in the input data

The Prioritization Hierarchy

Once filtered, repositories are evaluated against a strictly ordered list of criteria defined in the template under "Selection Criteria (in order of importance)" (lines 14-22):

  1. Highest author commit count (≥ 15) — Indicates substantial, sustained involvement
  2. Moderate author commit count (5-14) — Signifies meaningful contribution without requiring project ownership
  3. Contributions to popular open-source projects (≥ 1,000 stars) — Recognizes impact on widely-used libraries
  4. Technical complexity, real-world impact, code quality, community engagement, modern tech stack, originality — Captures qualitative engineering excellence

This hierarchy ensures the LLM prioritizes demonstrated expertise and ownership over superficial popularity metrics.

LLM-Driven Ranking and Enforcement

The prepared projects_data JSON is injected into the github_project_selection template. The LLM invocation (lines 78-86 in github.py) includes a system message explicitly demanding exactly 7 unique projects, preventing variable list sizes or duplicate entries.

The template constrains the LLM to apply the prioritization hierarchy while respecting the hard commit threshold, combining deterministic rules with the model's ability to assess qualitative factors like architectural sophistication.

Post-Processing Safeguards

After receiving the LLM response, the code at lines 96-115 in github.py implements defensive validation:

  • Deduplication: Removes duplicate project IDs from the returned array
  • Completeness check: Verifies at least 7 unique projects are present
  • Fallback mechanism: If the LLM returns fewer than 7 valid projects, the system automatically substitutes the top entries from a list sorted by author_commit_count

This ensures the final output always contains exactly 7 projects even if the LLM output is incomplete.

Implementation Example

To retrieve the curated list of 7 projects for a candidate:

from hiring_agent.github import fetch_and_display_github_info

# 1️⃣ Provide a GitHub profile URL (e.g., from a resume)

profile_url = "https://github.com/exampleUser"

# 2️⃣ Run the full enrichment pipeline

result = fetch_and_display_github_info(profile_url)

# 3️⃣ Access the exactly-7 selected projects

top_projects = result["projects"]  # List of dicts

for proj in top_projects:
    print(f"{proj['name']} – {proj['author_commit_count']} commits")

This end-to-end helper handles username extraction (extract_github_username), repository fetching (fetch_all_github_repos), JSON generation (generate_projects_json), and the LLM ranking pipeline automatically.

Summary

  • Minimum threshold: Projects require at least 4 author commits to qualify for evaluation
  • Ranking priority: Commit count (highest first), then project popularity (≥1,000 stars), then technical quality metrics
  • Dual enforcement: Thresholds are validated both in template logic and LLM system prompts
  • Guaranteed output: Post-processing deduplicates and falls back to commit-count sorting to ensure exactly 7 projects are always returned
  • Key files: github.py handles the pipeline, github_project_selection.jinja encodes the ranking rules

Frequently Asked Questions

What is the minimum number of commits required for a project to be considered?

A project must have an author_commit_count of at least 4 to be eligible for selection. This threshold is enforced in both the Jinja template (github_project_selection.jinja, lines 44-48) and the LLM system prompt (lines 55-57) to ensure low-effort contributions are filtered out before ranking occurs.

How does the system ensure exactly 7 projects are always returned?

The LLM is explicitly instructed to return exactly 7 unique projects in its system message (github.py, lines 78-86). Additionally, post-processing logic (lines 96-115) deduplicates results and implements a fallback mechanism that selects the top 7 projects by author commit count if the LLM returns fewer than 7 valid entries.

What factors besides commit count influence project selection?

While commit count is the primary signal (prioritizing projects with ≥15 commits, then 5-14), the LLM also considers contributions to popular open-source projects (≥1,000 stars) and qualitative factors including technical complexity, real-world impact, code quality, community engagement, modern tech stack, and originality.

Where are the selection criteria defined in the codebase?

The prioritization hierarchy and hard thresholds are defined in prompts/templates/github_project_selection.jinja (lines 14-22 for criteria ordering, lines 44-58 for threshold enforcement). The orchestration logic that invokes the LLM and handles post-processing resides in github.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →