# How to Troubleshoot Context Retrieval Returning Empty Results in Llama-GitHub

> Fix empty context retrieval in Llama-GitHub by troubleshooting search API hits and ranking logic. Learn quick solutions for zero results.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: troubleshooting-guide
- Published: 2026-03-04

---

**Empty context retrieval in Llama-GitHub typically stems from upstream search APIs returning zero hits or early-exit guards in the ranking logic that trigger when insufficient candidates are available.**

Llama-GitHub is an open-source Retrieval-Augmented Generation (RAG) framework that aggregates code, issues, and repository metadata from GitHub and Google Search. When `async_retrieve_context` returns an empty list, the failure usually occurs in one of three pipeline stages: search collection, context arrangement, or ranking and selection.

## Understanding the RAG Pipeline Architecture

The retrieval flow follows three distinct stages, each capable of producing empty results that propagate downstream.

### Stage 1: Search and Collection

The `GithubRAG.async_retrieve_context` method in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) (lines 15-98) orchestrates parallel searches across Google, GitHub code, issues, and repositories. Each search returns a raw list of dictionaries containing `content`, `url`, and metadata. If any underlying API returns no hits—due to rate limiting, invalid tokens, or overly specific queries—that branch contributes an empty list.

### Stage 2: Context Arrangement

The `RAGProcessor.arrange_context` method in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) (lines 368-383) normalizes every raw result into a uniform `{context, url}` structure and appends metadata like repository descriptions and star counts. This method concatenates only non-empty results; if **all** search results are empty, it returns an empty `context` list, causing downstream steps to skip processing.

### Stage 3: Ranking and Selection

The `RAGProcessor.retrieve_topn_contexts` method in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) (lines 394-473) performs three operations: reranking with a dedicated model, optional embedding similarity filtering, and final LLM-based relevance scoring. This method contains two critical early-exit guards:

- **Guard A**: If `len(selected_contexts) < top_n * 2`, the method returns a truncated list (lines 30-32 within the method implementation)
- **Guard B**: If an exception occurs during model scoring, the method logs the error and returns an empty list

Both guards trigger when the input `context_list` is empty or when models fail to load.

## Common Root Causes of Empty Context Retrieval

Understanding where the pipeline breaks requires checking specific failure modes:

**Zero Hits from Search Back-ends**
When `context_list` is empty before ranking, all search back-ends returned zero hits. This occurs with rate-limited GitHub APIs, invalid personal access tokens, or queries that exceed length limits. Check the debug logs from `async_retrieve_context` for entries like `logger.debug(f"Google search: {str(len(task_google_search.result()))}")`.

**Guard A Early Return**
If `retrieve_topn_contexts` exits at Guard A, the `selected_contexts` list after reranking contains fewer than `top_n * 2` items. For example, with `top_n=4`, the system requires at least 8 candidates. This happens when the reranker scores a tiny list or when search limits are set too low in the configuration.

**Ranking Exceptions**
Missing or incompatible model files trigger exceptions inside the ranking loop. If `self.llm_manager.get_rerank_model()` returns an object without `compute_score` or `encode` methods, numpy operations fail. The `try/except` block (lines 70-72) catches these and returns an empty list after logging `"Error retrieving top n context:"`.

**Empty Content Fields**
If the `_arrange_code_search_result` method (lines 212-228) or its counterparts for issues and repositories process results with missing `content` keys, they build strings containing only metadata. Empty source content results in empty strings that later filtering removes from the candidate pool.

## Step-by-Step Debugging Checklist

Follow this sequence to isolate the failure point:

1. **Enable verbose logging** – Set `LLAMA_GITHUB_LOG_LEVEL=DEBUG` or modify [`llama_github/logger.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/logger.py) to use `logging.DEBUG`.

2. **Inspect search task outputs** – Run a single query and verify printed lengths:
   ```text
   Google search: 12
   Code search: 5
   Issue search: 0
   Repo search: 3
   ```

   Any zero indicates a failed search branch.

3. **Verify `context_list` before ranking** – Insert `logger.debug(context_list)` immediately after line 89 of [`github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/github_rag.py) to confirm data reaches the processor.

4. **Validate model instantiation** – Ensure `self.llm_manager.get_rerank_model()` and `self.llm_manager.get_embedding_model()` return objects with functional `compute_score` and `encode` methods. Mock objects in tests must return realistic dummy arrays.

5. **Check Guard A thresholds** – After the rerank step, log `logger.debug(f"selected_contexts={len(selected_contexts)} need>={top_n*2}")`. If the number falls below the threshold, increase `code_search_max_hits` or `issue_search_max_hits` in [`config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config.json), or lower `"top_n_contexts"`.

6. **Expose hidden exceptions** – Temporarily replace the `except` block’s `return []` with `raise` in [`rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/rag_processor.py) to surface full stack traces, then restore the original handling once identified.

## Code Examples for Diagnosis and Testing

### Normal Usage Pattern

This example demonstrates standard retrieval with debugging hooks:

```python
import asyncio
from llama_github.github_rag import GithubRAG

async def demo():
    rag = GithubRAG(
        github_access_token="YOUR_TOKEN",  # Requires repo scope

        openai_api_key="OPENAI_KEY",
        simple_mode=False                  # Full pipeline: code, issue, repo, google

    )
    contexts = await rag.async_retrieve_context(
        query="How can I stream a large CSV file with pandas?"
    )
    print(f"Retrieved {len(contexts)} contexts")
    for i, ctx in enumerate(contexts[:3], 1):
        print(f"--- Context {i} ---")
        print(ctx[:500])

asyncio.run(demo())

```

If this prints `Retrieved 0 contexts`, apply the checklist above to locate the failure stage.

### Reproducing Guard A Behavior

This unit test reproduces the early-exit condition when insufficient contexts are available:

```python
import pytest
import asyncio
from unittest.mock import MagicMock
from llama_github.rag_processing.rag_processor import RAGProcessor

@pytest.fixture
def processor():
    mock_api = MagicMock()
    mock_manager = MagicMock()
    mock_handler = MagicMock()
    return RAGProcessor(
        github_api_handler=mock_api,
        llm_manager=mock_manager,
        llm_handler=mock_handler
    )

@pytest.mark.asyncio
async def test_empty_contexts_trigger_guard_a(processor):
    # Simulate reranker returning only two scores

    mock_reranker = MagicMock()
    mock_reranker.compute_score.return_value = [0.9, 0.1]
    processor.llm_manager.get_rerank_model.return_value = mock_reranker
    
    # Provide only two contexts (< top_n * 2 where top_n = 4)

    context_list = [
        {"context": "good context", "url": "u1"},
        {"context": "bad context", "url": "u2"},
    ]
    
    result = await processor.retrieve_topn_contexts(
        context_list, "query", top_n=4
    )
    assert len(result) == 2  # Guard A returns available contexts

```

### Bypassing Guard A with Sufficient Candidates

To force the full ranking path without early exit, provide adequate candidate volume:

```python
import numpy as np
from unittest.mock import MagicMock

# Generate ten dummy contexts

context_list = [
    {"context": f"context {i}", "url": f"url{i}"} 
    for i in range(10)
]

# Mock embedding model with consistent 5-dimensional vectors

mock_emb = MagicMock()
mock_emb.encode.side_effect = lambda txt: np.ones(5)
processor.llm_manager.get_embedding_model.return_value = mock_emb

# Execute with top_n=4 (requires 8+ candidates)

top_contexts = await processor.retrieve_topn_contexts(
    context_list, "query", top_n=4
)
assert len(top_contexts) == 4

```

## Summary

- **Llama-GitHub** retrieves context through three stages: search collection (`async_retrieve_context`), arrangement (`arrange_context`), and ranking (`retrieve_topn_contexts`).
- **Empty results** usually indicate zero hits from GitHub/Google APIs or early-exit **Guard A** when `selected_contexts < top_n * 2`.
- **Debugging** requires enabling `DEBUG` logging in [`logger.py`](https://github.com/jetxu-llm/llama-github/blob/main/logger.py), inspecting search task lengths, and verifying model objects expose `compute_score` and `encode` methods.
- **Configuration** adjustments to `top_n_contexts` or `*_search_max_hits` in [`config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config.json) resolve threshold-related empty returns.

## Frequently Asked Questions

### Why does `async_retrieve_context` return an empty list even with a valid GitHub token?

Valid tokens can still encounter **rate limits** or **scope restrictions**. Verify the token has `repo` scope and check debug logs for specific search branch failures. If GitHub code search returns zero results while Google search returns data, the token may lack permissions for private repositories or the query may match no public code.

### What is Guard A in `retrieve_topn_contexts` and how does it affect results?

**Guard A** is a protective threshold in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py) that triggers when the reranked context list contains fewer than `top_n * 2` items. When activated, the method returns all available contexts immediately without applying embedding similarity or LLM scoring. This prevents over-filtering small candidate pools but can return fewer results than requested.

### How can I verify if the reranker model is loading correctly?

Check that `LLMManager.get_rerank_model()` returns an object with a callable `compute_score` method. In unit tests, mocks must return list objects of the same length as the input contexts. Runtime failures typically log `"Error retrieving top n context:"` with a stack trace pointing to numpy operations or missing model files in [`llama_github/llm_integration/initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/initial_load.py).

### Where should I adjust the number of contexts retrieved?

Modify [`llama_github/config/config.json`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/config/config.json) to change `"top_n_contexts"` (controls final output count) or increase `"code_search_max_hits"`, `"issue_search_max_hits"`, and `"repo_search_max_hits"` to feed more candidates into the ranking stage. If you need fewer strict filters, reduce `top_n` to prevent Guard A from truncating the pipeline.