# How the Deduplication Algorithm Works for Search Results in Llama-GitHub

> Discover how the deduplication algorithm in Llama-GitHub removes duplicate search results. Learn how it uses unique identifiers and quality thresholds to ensure valuable code, issue, and repository listings.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: internals
- Published: 2026-03-04

---

**The deduplication algorithm in llama-github removes duplicate entries from code, issue, and repository search pipelines by tracking unique identifiers (URLs or full names) in a `seen` set while applying quality thresholds to retain high-value results.**

The llama-github library implements a Retrieval-Augmented Generation (RAG) system that aggregates data from multiple GitHub search endpoints to provide context for LLM queries. To prevent redundant information from inflating context windows and degrading answer quality, the deduplication algorithm for search results processes each result set to ensure only unique, high-quality entries proceed to the ranking stage.

## How the Deduplication Algorithm Works

The deduplication logic resides in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) and operates independently across three primary search pipelines: code search, issue search, and repository search. Each pipeline uses a distinct unique identifier to detect duplicates, ensuring that the same file, issue, or repository does not appear multiple times in the final context.

### Code Search Deduplication

For code search results (lines **61-70**), the algorithm uses the raw GitHub file **`url`** as the primary key. However, unlike simple set-based filtering, this pipeline implements a **quality gate** that evaluates popularity and ranking position before admitting an entry into the deduplicated list.

At line **66**, the algorithm calculates a composite score:

```python
value = d["url"]
if (value not in seen and
    d["stargazers_count"] + config.get("code_search_max_hits") - d["index"]
    >= config.get("min_stars_to_keep_result")):
    seen.add(value)
    unique_list.append(d)

```

The score combines the repository's star count with the hit's index position (higher-ranked results have lower index values). Only URLs that pass the `min_stars_to_keep_result` threshold (default 20 stars) are retained, ensuring that popular, authoritative code examples survive the deduplication process.

### Issue Search Deduplication

Issue search deduplication occurs at lines **103-113** and relies on the API **`url`** field from the GitHub REST response. The algorithm stores the raw `api.github.com` URL in the `seen` set to detect duplicates, then transforms it to a human-readable `github.com` URL for downstream processing at lines **114-116**:

```python
api_url = d["url"]
if api_url not in seen:
    seen.add(api_url)
    html_url = api_url.replace('api.github.com/repos', 'github.com').replace('issues/', 'issues/')
    d["url"] = html_url
    unique_list.append(d)

```

This transformation happens post-deduplication, ensuring that the same issue accessed through different API parameters does not create multiple entries while preserving the canonical web URL for LLM context.

### Repository Search Deduplication

For repository searches (lines **222-231**), the algorithm uses the repository **`full_name`** (e.g., `owner/repo`) as the unique identifier. This approach prevents the same repository from being fetched multiple times when it matches different query keywords:

```python
value = d.full_name
if value not in seen:
    seen.add(value)
    unique_list.append(d)

```

This deduplication step occurs before the system initiates parallel fetches for README content and repository structure data, preventing redundant API calls and processing overhead.

### Google Search Handling

The Google search pipeline (lines **44-62**) does not implement explicit deduplication within llama-github. According to the source code, the Jina-AI endpoint used for this pipeline already returns a unique list of URLs, so the code simply aggregates these results without additional filtering.

## Step-by-Step Deduplication Process

Regardless of the search type, the deduplication algorithm follows a consistent five-step pattern implemented in the retrieval methods:

1. **Collect raw results** – Each async task (`code_search_retrieval`, `issue_search_retrieval`, `repo_search_retrieval`) queries the GitHub API and accumulates hits in a temporary `result` list.

2. **Initialize tracking structures** – The algorithm creates a `seen` set to store unique identifiers and a `unique_list` to hold filtered results:
   
   ```python
   seen = set()
   unique_list = []
   ```

3. **Iterate and filter** – For each dictionary `d` in the raw results, the algorithm extracts the appropriate identifier (URL or `full_name`) and checks membership in `seen`. For code search, it additionally validates the star-count threshold.

4. **Replace the original list** – After processing all items, the algorithm substitutes the raw list with the deduplicated version:
   
   ```python
   result = unique_list
   ```

5. **Proceed to RAG processing** – The deduplicated list passes to [`rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/rag_processor.py), where entries are ranked, chunked, and formatted for LLM prompting.

## Configuration and Quality Thresholds

The deduplication behavior is tunable via [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json). Two key parameters influence which code search results survive filtering:

- **`min_stars_to_keep_result`** (default: 20) – The minimum star-count threshold for code search hits. Results below this value are excluded during deduplication.
- **`code_search_max_hits`** – Used in the scoring calculation to weight the index position of results.

These settings allow operators to balance between result diversity and quality, ensuring that highly-starred repositories receive priority in the final context window.

## Code Examples

The following examples demonstrate how to invoke the RAG system and observe deduplication in action.

### Basic Usage of GithubRAG

```python
from llama_github.github_rag import GithubRAG

# Initialise with a personal access token

rag = GithubRAG(github_access_token="YOUR_TOKEN", simple_mode=False)

# Query triggers internal deduplication automatically

answer = rag.answer_with_context(
    query="How does llama-github handle duplicate search results?",
    simple_mode=False,
)

print(answer)

```

The `answer_with_context` method calls `async_retrieve_context`, which executes the four search tasks and applies the deduplication logic described above.

### Inspecting Deduplicated Code Search Results

```python
import asyncio

async def demo_code_dedup():
    rag = GithubRAG(github_access_token="YOUR_TOKEN")
    # Direct low-level method to see raw vs deduped results

    raw = await rag.github_api_handler.search_code('language:python "def"')
    print(f"Raw hits: {len(raw)}")

    # RAG method applies deduplication at lines 62-69

    deduped = await rag.code_search_retrieval(query="How to read a file in Python?")
    print(f"Deduped hits: {len(deduped)}")
    
    # Verify uniqueness

    urls = [hit["url"] for hit in deduped]
    print(f"Unique URLs: {len(set(urls))}")

asyncio.run(demo_code_dedup())

```

### Verifying Repository Deduplication

```python
import asyncio

async def demo_repo_dedup():
    rag = GithubRAG(github_access_token="YOUR_TOKEN")
    # Search that may return the same repo multiple times

    repo_results = await rag.repo_search_retrieval(query="LLM")
    names = [r["full_name"] for r in repo_results]
    
    # Should equal len(repo_results) if deduplication worked

    print("Unique repo count:", len(set(names)))

asyncio.run(demo_repo_dedup())

```

This example targets the `repo_search_retrieval` method at lines **224-230** of [`github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/github_rag.py).

## Summary

- The deduplication algorithm in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) processes **code**, **issue**, and **repository** search results independently using type-specific unique identifiers.
- **Code search** uses file URLs combined with a star-count quality threshold to prioritize popular repositories.
- **Issue search** uses API URLs for deduplication, then transforms them to human-readable formats at lines **114-116**.
- **Repository search** uses the `full_name` field to prevent duplicate repository processing.
- A simple `seen` set pattern provides deterministic O(n) deduplication performance across all pipelines.
- Configuration via [`config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config.json) allows tuning of quality thresholds without code changes.

## Frequently Asked Questions

### What identifier does the deduplication algorithm use for code search results?

The algorithm uses the raw GitHub file **`url`** field as the unique identifier for code search results. This URL represents the direct link to the source file in the repository. The deduplication check occurs at line **66** of [`github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/github_rag.py), where the URL is verified against a `seen` set before being added to the unique results list.

### How does the algorithm prioritize which duplicate entries to keep?

For code search, the algorithm implements a **quality gate** that prioritizes entries from popular repositories or higher-ranked search positions. It calculates a score using `stargazers_count + code_search_max_hits - index`, which must exceed the `min_stars_to_keep_result` threshold (default 20 stars). This ensures that if duplicate URLs exist, the instance with better repository metrics survives the filter.

### Does the deduplication work across different search types?

No, the deduplication operates **independently** within each search pipeline. The algorithm processes code search, issue search, and repository search separately, meaning the same repository could appear once in the code results and once in the repository results. Cross-type deduplication is not implemented, as each result type serves a distinct purpose in the RAG context.

### Where is the deduplication threshold configured?

Quality thresholds are defined in **[`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json)**. The key parameter `min_stars_to_keep_result` controls the minimum star count required for code search results to survive deduplication, while `code_search_max_hits` influences the scoring calculation. These settings allow users to adjust the quality bar without modifying the source code in [`github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/github_rag.py).