How the Deduplication Algorithm Works for Search Results in Llama-GitHub
The deduplication algorithm in llama-github removes duplicate entries from code, issue, and repository search pipelines by tracking unique identifiers (URLs or full names) in a seen set while applying quality thresholds to retain high-value results.
The llama-github library implements a Retrieval-Augmented Generation (RAG) system that aggregates data from multiple GitHub search endpoints to provide context for LLM queries. To prevent redundant information from inflating context windows and degrading answer quality, the deduplication algorithm for search results processes each result set to ensure only unique, high-quality entries proceed to the ranking stage.
How the Deduplication Algorithm Works
The deduplication logic resides in llama_github/github_rag.py and operates independently across three primary search pipelines: code search, issue search, and repository search. Each pipeline uses a distinct unique identifier to detect duplicates, ensuring that the same file, issue, or repository does not appear multiple times in the final context.
Code Search Deduplication
For code search results (lines 61-70), the algorithm uses the raw GitHub file url as the primary key. However, unlike simple set-based filtering, this pipeline implements a quality gate that evaluates popularity and ranking position before admitting an entry into the deduplicated list.
At line 66, the algorithm calculates a composite score:
value = d["url"]
if (value not in seen and
d["stargazers_count"] + config.get("code_search_max_hits") - d["index"]
>= config.get("min_stars_to_keep_result")):
seen.add(value)
unique_list.append(d)
The score combines the repository's star count with the hit's index position (higher-ranked results have lower index values). Only URLs that pass the min_stars_to_keep_result threshold (default 20 stars) are retained, ensuring that popular, authoritative code examples survive the deduplication process.
Issue Search Deduplication
Issue search deduplication occurs at lines 103-113 and relies on the API url field from the GitHub REST response. The algorithm stores the raw api.github.com URL in the seen set to detect duplicates, then transforms it to a human-readable github.com URL for downstream processing at lines 114-116:
api_url = d["url"]
if api_url not in seen:
seen.add(api_url)
html_url = api_url.replace('api.github.com/repos', 'github.com').replace('issues/', 'issues/')
d["url"] = html_url
unique_list.append(d)
This transformation happens post-deduplication, ensuring that the same issue accessed through different API parameters does not create multiple entries while preserving the canonical web URL for LLM context.
Repository Search Deduplication
For repository searches (lines 222-231), the algorithm uses the repository full_name (e.g., owner/repo) as the unique identifier. This approach prevents the same repository from being fetched multiple times when it matches different query keywords:
value = d.full_name
if value not in seen:
seen.add(value)
unique_list.append(d)
This deduplication step occurs before the system initiates parallel fetches for README content and repository structure data, preventing redundant API calls and processing overhead.
Google Search Handling
The Google search pipeline (lines 44-62) does not implement explicit deduplication within llama-github. According to the source code, the Jina-AI endpoint used for this pipeline already returns a unique list of URLs, so the code simply aggregates these results without additional filtering.
Step-by-Step Deduplication Process
Regardless of the search type, the deduplication algorithm follows a consistent five-step pattern implemented in the retrieval methods:
-
Collect raw results – Each async task (
code_search_retrieval,issue_search_retrieval,repo_search_retrieval) queries the GitHub API and accumulates hits in a temporaryresultlist. -
Initialize tracking structures – The algorithm creates a
seenset to store unique identifiers and aunique_listto hold filtered results:seen = set() unique_list = [] -
Iterate and filter – For each dictionary
din the raw results, the algorithm extracts the appropriate identifier (URL orfull_name) and checks membership inseen. For code search, it additionally validates the star-count threshold. -
Replace the original list – After processing all items, the algorithm substitutes the raw list with the deduplicated version:
result = unique_list -
Proceed to RAG processing – The deduplicated list passes to
rag_processor.py, where entries are ranked, chunked, and formatted for LLM prompting.
Configuration and Quality Thresholds
The deduplication behavior is tunable via llama_github/config/config.json. Two key parameters influence which code search results survive filtering:
min_stars_to_keep_result(default: 20) – The minimum star-count threshold for code search hits. Results below this value are excluded during deduplication.code_search_max_hits– Used in the scoring calculation to weight the index position of results.
These settings allow operators to balance between result diversity and quality, ensuring that highly-starred repositories receive priority in the final context window.
Code Examples
The following examples demonstrate how to invoke the RAG system and observe deduplication in action.
Basic Usage of GithubRAG
from llama_github.github_rag import GithubRAG
# Initialise with a personal access token
rag = GithubRAG(github_access_token="YOUR_TOKEN", simple_mode=False)
# Query triggers internal deduplication automatically
answer = rag.answer_with_context(
query="How does llama-github handle duplicate search results?",
simple_mode=False,
)
print(answer)
The answer_with_context method calls async_retrieve_context, which executes the four search tasks and applies the deduplication logic described above.
Inspecting Deduplicated Code Search Results
import asyncio
async def demo_code_dedup():
rag = GithubRAG(github_access_token="YOUR_TOKEN")
# Direct low-level method to see raw vs deduped results
raw = await rag.github_api_handler.search_code('language:python "def"')
print(f"Raw hits: {len(raw)}")
# RAG method applies deduplication at lines 62-69
deduped = await rag.code_search_retrieval(query="How to read a file in Python?")
print(f"Deduped hits: {len(deduped)}")
# Verify uniqueness
urls = [hit["url"] for hit in deduped]
print(f"Unique URLs: {len(set(urls))}")
asyncio.run(demo_code_dedup())
Verifying Repository Deduplication
import asyncio
async def demo_repo_dedup():
rag = GithubRAG(github_access_token="YOUR_TOKEN")
# Search that may return the same repo multiple times
repo_results = await rag.repo_search_retrieval(query="LLM")
names = [r["full_name"] for r in repo_results]
# Should equal len(repo_results) if deduplication worked
print("Unique repo count:", len(set(names)))
asyncio.run(demo_repo_dedup())
This example targets the repo_search_retrieval method at lines 224-230 of github_rag.py.
Summary
- The deduplication algorithm in
llama_github/github_rag.pyprocesses code, issue, and repository search results independently using type-specific unique identifiers. - Code search uses file URLs combined with a star-count quality threshold to prioritize popular repositories.
- Issue search uses API URLs for deduplication, then transforms them to human-readable formats at lines 114-116.
- Repository search uses the
full_namefield to prevent duplicate repository processing. - A simple
seenset pattern provides deterministic O(n) deduplication performance across all pipelines. - Configuration via
config.jsonallows tuning of quality thresholds without code changes.
Frequently Asked Questions
What identifier does the deduplication algorithm use for code search results?
The algorithm uses the raw GitHub file url field as the unique identifier for code search results. This URL represents the direct link to the source file in the repository. The deduplication check occurs at line 66 of github_rag.py, where the URL is verified against a seen set before being added to the unique results list.
How does the algorithm prioritize which duplicate entries to keep?
For code search, the algorithm implements a quality gate that prioritizes entries from popular repositories or higher-ranked search positions. It calculates a score using stargazers_count + code_search_max_hits - index, which must exceed the min_stars_to_keep_result threshold (default 20 stars). This ensures that if duplicate URLs exist, the instance with better repository metrics survives the filter.
Does the deduplication work across different search types?
No, the deduplication operates independently within each search pipeline. The algorithm processes code search, issue search, and repository search separately, meaning the same repository could appear once in the code results and once in the repository results. Cross-type deduplication is not implemented, as each result type serves a distinct purpose in the RAG context.
Where is the deduplication threshold configured?
Quality thresholds are defined in llama_github/config/config.json. The key parameter min_stars_to_keep_result controls the minimum star count required for code search results to survive deduplication, while code_search_max_hits influences the scoring calculation. These settings allow users to adjust the quality bar without modifying the source code in github_rag.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →