How to Troubleshoot Context Retrieval Returning Empty Results in Llama-GitHub

Empty context retrieval in Llama-GitHub typically stems from upstream search APIs returning zero hits or early-exit guards in the ranking logic that trigger when insufficient candidates are available.

Llama-GitHub is an open-source Retrieval-Augmented Generation (RAG) framework that aggregates code, issues, and repository metadata from GitHub and Google Search. When async_retrieve_context returns an empty list, the failure usually occurs in one of three pipeline stages: search collection, context arrangement, or ranking and selection.

Understanding the RAG Pipeline Architecture

The retrieval flow follows three distinct stages, each capable of producing empty results that propagate downstream.

Stage 1: Search and Collection

The GithubRAG.async_retrieve_context method in llama_github/github_rag.py (lines 15-98) orchestrates parallel searches across Google, GitHub code, issues, and repositories. Each search returns a raw list of dictionaries containing content, url, and metadata. If any underlying API returns no hits—due to rate limiting, invalid tokens, or overly specific queries—that branch contributes an empty list.

Stage 2: Context Arrangement

The RAGProcessor.arrange_context method in llama_github/rag_processing/rag_processor.py (lines 368-383) normalizes every raw result into a uniform {context, url} structure and appends metadata like repository descriptions and star counts. This method concatenates only non-empty results; if all search results are empty, it returns an empty context list, causing downstream steps to skip processing.

Stage 3: Ranking and Selection

The RAGProcessor.retrieve_topn_contexts method in llama_github/rag_processing/rag_processor.py (lines 394-473) performs three operations: reranking with a dedicated model, optional embedding similarity filtering, and final LLM-based relevance scoring. This method contains two critical early-exit guards:

  • Guard A: If len(selected_contexts) < top_n * 2, the method returns a truncated list (lines 30-32 within the method implementation)
  • Guard B: If an exception occurs during model scoring, the method logs the error and returns an empty list

Both guards trigger when the input context_list is empty or when models fail to load.

Common Root Causes of Empty Context Retrieval

Understanding where the pipeline breaks requires checking specific failure modes:

Zero Hits from Search Back-ends When context_list is empty before ranking, all search back-ends returned zero hits. This occurs with rate-limited GitHub APIs, invalid personal access tokens, or queries that exceed length limits. Check the debug logs from async_retrieve_context for entries like logger.debug(f"Google search: {str(len(task_google_search.result()))}").

Guard A Early Return If retrieve_topn_contexts exits at Guard A, the selected_contexts list after reranking contains fewer than top_n * 2 items. For example, with top_n=4, the system requires at least 8 candidates. This happens when the reranker scores a tiny list or when search limits are set too low in the configuration.

Ranking Exceptions Missing or incompatible model files trigger exceptions inside the ranking loop. If self.llm_manager.get_rerank_model() returns an object without compute_score or encode methods, numpy operations fail. The try/except block (lines 70-72) catches these and returns an empty list after logging "Error retrieving top n context:".

Empty Content Fields If the _arrange_code_search_result method (lines 212-228) or its counterparts for issues and repositories process results with missing content keys, they build strings containing only metadata. Empty source content results in empty strings that later filtering removes from the candidate pool.

Step-by-Step Debugging Checklist

Follow this sequence to isolate the failure point:

  1. Enable verbose logging – Set LLAMA_GITHUB_LOG_LEVEL=DEBUG or modify llama_github/logger.py to use logging.DEBUG.

  2. Inspect search task outputs – Run a single query and verify printed lengths:

    Google search: 12
    Code search: 5
    Issue search: 0
    Repo search: 3

    Any zero indicates a failed search branch.

  3. Verify context_list before ranking – Insert logger.debug(context_list) immediately after line 89 of github_rag.py to confirm data reaches the processor.

  4. Validate model instantiation – Ensure self.llm_manager.get_rerank_model() and self.llm_manager.get_embedding_model() return objects with functional compute_score and encode methods. Mock objects in tests must return realistic dummy arrays.

  5. Check Guard A thresholds – After the rerank step, log logger.debug(f"selected_contexts={len(selected_contexts)} need>={top_n*2}"). If the number falls below the threshold, increase code_search_max_hits or issue_search_max_hits in config.json, or lower "top_n_contexts".

  6. Expose hidden exceptions – Temporarily replace the except block’s return [] with raise in rag_processor.py to surface full stack traces, then restore the original handling once identified.

Code Examples for Diagnosis and Testing

Normal Usage Pattern

This example demonstrates standard retrieval with debugging hooks:

import asyncio
from llama_github.github_rag import GithubRAG

async def demo():
    rag = GithubRAG(
        github_access_token="YOUR_TOKEN",  # Requires repo scope

        openai_api_key="OPENAI_KEY",
        simple_mode=False                  # Full pipeline: code, issue, repo, google

    )
    contexts = await rag.async_retrieve_context(
        query="How can I stream a large CSV file with pandas?"
    )
    print(f"Retrieved {len(contexts)} contexts")
    for i, ctx in enumerate(contexts[:3], 1):
        print(f"--- Context {i} ---")
        print(ctx[:500])

asyncio.run(demo())

If this prints Retrieved 0 contexts, apply the checklist above to locate the failure stage.

Reproducing Guard A Behavior

This unit test reproduces the early-exit condition when insufficient contexts are available:

import pytest
import asyncio
from unittest.mock import MagicMock
from llama_github.rag_processing.rag_processor import RAGProcessor

@pytest.fixture
def processor():
    mock_api = MagicMock()
    mock_manager = MagicMock()
    mock_handler = MagicMock()
    return RAGProcessor(
        github_api_handler=mock_api,
        llm_manager=mock_manager,
        llm_handler=mock_handler
    )

@pytest.mark.asyncio
async def test_empty_contexts_trigger_guard_a(processor):
    # Simulate reranker returning only two scores

    mock_reranker = MagicMock()
    mock_reranker.compute_score.return_value = [0.9, 0.1]
    processor.llm_manager.get_rerank_model.return_value = mock_reranker
    
    # Provide only two contexts (< top_n * 2 where top_n = 4)

    context_list = [
        {"context": "good context", "url": "u1"},
        {"context": "bad context", "url": "u2"},
    ]
    
    result = await processor.retrieve_topn_contexts(
        context_list, "query", top_n=4
    )
    assert len(result) == 2  # Guard A returns available contexts

Bypassing Guard A with Sufficient Candidates

To force the full ranking path without early exit, provide adequate candidate volume:

import numpy as np
from unittest.mock import MagicMock

# Generate ten dummy contexts

context_list = [
    {"context": f"context {i}", "url": f"url{i}"} 
    for i in range(10)
]

# Mock embedding model with consistent 5-dimensional vectors

mock_emb = MagicMock()
mock_emb.encode.side_effect = lambda txt: np.ones(5)
processor.llm_manager.get_embedding_model.return_value = mock_emb

# Execute with top_n=4 (requires 8+ candidates)

top_contexts = await processor.retrieve_topn_contexts(
    context_list, "query", top_n=4
)
assert len(top_contexts) == 4

Summary

  • Llama-GitHub retrieves context through three stages: search collection (async_retrieve_context), arrangement (arrange_context), and ranking (retrieve_topn_contexts).
  • Empty results usually indicate zero hits from GitHub/Google APIs or early-exit Guard A when selected_contexts < top_n * 2.
  • Debugging requires enabling DEBUG logging in logger.py, inspecting search task lengths, and verifying model objects expose compute_score and encode methods.
  • Configuration adjustments to top_n_contexts or *_search_max_hits in config.json resolve threshold-related empty returns.

Frequently Asked Questions

Why does async_retrieve_context return an empty list even with a valid GitHub token?

Valid tokens can still encounter rate limits or scope restrictions. Verify the token has repo scope and check debug logs for specific search branch failures. If GitHub code search returns zero results while Google search returns data, the token may lack permissions for private repositories or the query may match no public code.

What is Guard A in retrieve_topn_contexts and how does it affect results?

Guard A is a protective threshold in llama_github/rag_processing/rag_processor.py that triggers when the reranked context list contains fewer than top_n * 2 items. When activated, the method returns all available contexts immediately without applying embedding similarity or LLM scoring. This prevents over-filtering small candidate pools but can return fewer results than requested.

How can I verify if the reranker model is loading correctly?

Check that LLMManager.get_rerank_model() returns an object with a callable compute_score method. In unit tests, mocks must return list objects of the same length as the input contexts. Runtime failures typically log "Error retrieving top n context:" with a stack trace pointing to numpy operations or missing model files in llama_github/llm_integration/initial_load.py.

Where should I adjust the number of contexts retrieved?

Modify llama_github/config/config.json to change "top_n_contexts" (controls final output count) or increase "code_search_max_hits", "issue_search_max_hits", and "repo_search_max_hits" to feed more candidates into the ranking stage. If you need fewer strict filters, reduce top_n to prevent Guard A from truncating the pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →