# How the Agentic RAG Architecture Works in llama-github: A Technical Deep Dive

> Discover the Agentic RAG architecture in llama-github. Learn how LLM query planning, parallel API searches, and multi-stage ranking retrieve and synthesize code context from GitHub repositories.

- Repository: [Jet Xu/llama-github](https://github.com/jetxu-llm/llama-github)
- Tags: deep-dive
- Published: 2026-03-04

---

**The Agentic RAG architecture in llama-github orchestrates LLM-driven query planning, parallel GitHub API searches, and multi-stage ranking to retrieve and synthesize context from repositories.**

The **llama-github** library implements a sophisticated Agentic Retrieval-Augmented Generation (RAG) pipeline that treats the LLM as an active participant in the retrieval process. Unlike traditional RAG systems that rely on static retrieval strategies, this architecture dynamically generates search criteria, evaluates result relevance through multiple scoring layers, and synthesizes answers from GitHub source code, issues, and repositories. The system is designed around three core Python modules that coordinate authentication, retrieval, and language model interactions.

## Core Components of the Agentic RAG System

### GithubRAG: The High-Level Orchestrator

The **`GithubRAG`** class in [`llama_github/github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/github_rag.py) serves as the primary entry point for the entire pipeline. During initialization (lines 71‑108), it constructs four essential subsystems: a `GitHubAuthManager` for API authentication, a `RepositoryPool` for caching, an `LLMManager` for model handling, and a `RAGProcessor` for search execution. The orchestrator exposes two main async interfaces: `async_retrieve_context` for raw context retrieval and `async_answer_with_context` for complete answer generation.

### RAGProcessor: Search Strategy and Ranking Engine

Located in [`llama_github/rag_processing/rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/rag_processing/rag_processor.py), the **`RAGProcessor`** contains the agentic logic. It implements `analyze_question` (lines 51‑66) to generate LLM-driven search strategies, methods like `get_code_search_criteria` to produce structured Pydantic search queries, and `retrieve_topn_contexts` for multi-stage result ranking. This component also handles content chunking via `RecursiveCharacterTextSplitter` and arranges raw GitHub API responses into uniform context objects.

### LLMManager and LLMHandler: Unified Model Abstraction

The **`LLMManager`** ([`llama_github/llm_integration/initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/initial_load.py)) and **`LLMHandler`** ([`llama_github/llm_integration/llm_handler.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/llm_handler.py)) provide a unified interface for cloud and local models. The manager initializes embedding models, rerankers, and inference endpoints for OpenAI, Mistral, or HuggingFace deployments. The handler exposes `ainvoke`, an async method that standardizes prompt templating and structured output parsing across all supported providers.

## The Agentic RAG Pipeline Step-by-Step

### 1. System Initialization and Construction

When instantiating `GithubRAG`, the `__init__` method wires together all dependencies. The system loads configuration from [`config.json`](https://github.com/jetxu-llm/llama-github/blob/main/config.json), which stores prompt templates, model names, and thresholds like `top_n_contexts` and `min_stars_to_keep_result`. The `RAGProcessor` receives injected instances of `GitHubAPIHandler` and `LLMManager`, establishing the dependency chain for subsequent operations.

### 2. Query Analysis and Strategy Generation

Upon receiving a user query, the pipeline enters an agentic planning phase. The `analyze_question` method calls the LLM with the `always_answer_prompt` from configuration, returning a structured tuple: `[question, draft_answer, code_logic, issue_logic]`. This LLM-driven analysis determines what to search and how to combine different GitHub data sources (code, issues, repositories).

### 3. Parallel Retrieval Execution

The `async_retrieve_context` method (lines 154‑188 of [`github_rag.py`](https://github.com/jetxu-llm/llama-github/blob/main/github_rag.py)) executes four concurrent tasks using `asyncio.gather`:

- **`analyze_question`** – Generates the search strategy.
- **`google_search_retrieval`** – Performs web search limited to `site:github.com`.
- **`code_search_retrieval`** – Queries the GitHub Code Search API.
- **`issue_search_retrieval`** – Queries the GitHub Issues API.
- **`repo_search_retrieval`** – Queries the GitHub Repository API.

Each search method receives criteria generated by dedicated LLM calls (e.g., `get_repo_search_criteria`), which return structured Pydantic models like `_GitHubRepoSearchCriteria` containing search terms and necessity scores.

### 4. Context Chunking and Arrangement

Raw API results undergo processing in `RAGProcessor.arrange_context` and its helper methods (`_arrange_code_search_result`, `_arrange_issue_search_result`). Long files and issue bodies are split into token-aware chunks using `RecursiveCharacterTextSplitter`. The system enriches each chunk with metadata including repository stars, URL, programming language, and timestamps, creating uniform context dictionaries for downstream ranking.

### 5. Multi-Stage Ranking and Scoring

The `retrieve_topn_contexts` method implements a three-layer ranking system:

1. **Reranker Scoring** – A sentence-pair classification model scores each `(query, document)` pair.
2. **Embedding Similarity** – The top `3 × n` candidates are re-scored using cosine similarity from the embedding model (`llm_manager.get_embedding_model()`).
3. **LLM Relevance Scoring** – A lightweight LLM call (`_ContextRelevanceScore`) computes a context relevance score.

The final ranking uses a composite formula: `llm_score × cos_sim × rerank_score`. The top `n` contexts are returned to the caller.

### 6. Final Answer Synthesis

In `async_answer_with_context`, the system retrieves contexts (if not provided) and invokes `LLMHandler.ainvoke` with the `general_prompt` template. The structured prompt includes the retrieved chunks and their metadata, instructing the LLM to generate a cited answer that references specific GitHub URLs.

## What Makes the Architecture "Agentic"

**LLM-Driven Planning** – The initial `analyze_question` call determines search strategy, effectively letting the model decide what information is needed before executing API calls.

**Adaptive Search Depth** – The `necessity_score` in repository search criteria allows the system to skip expensive repo-level searches when the LLM judges them unnecessary.

**Dynamic Reranking** – Rather than relying solely on static embeddings, the system employs a dedicated reranker model and a relevance-scoring LLM to judge context quality, creating a feedback loop where the model evaluates its own retrieval choices.

**Structured Decision Making** – Search criteria are generated as structured Pydantic objects (`_GitHubCodeSearchCriteria`, etc.), enabling the LLM to output precise API parameters rather than free-text queries.

## Implementation Examples

### Initialize the Full Agentic Pipeline

```python
from llama_github.github_rag import GithubRAG

rag = GithubRAG(
    github_access_token="ghp_YOUR_TOKEN",
    openai_api_key="sk-YOUR_KEY",
    simple_mode=False  # Enable full embedding + rerank pipeline

)

```

### Retrieve Context Only

```python
contexts = await rag.async_retrieve_context(
    "How does React handle concurrent rendering?"
)

# Returns list of dicts: {"context": "<chunk>", "url": "<github url>"}

```

### Generate a Complete Answer

```python
answer = await rag.async_answer_with_context(
    query="Explain the RAGProcessor chunking strategy"
)
print(answer)

```

### Enable Simple Mode for Fast Queries

```python
rag_simple = GithubRAG(
    github_access_token="ghp_TOKEN",
    simple_mode=True  # Uses only Google/Jina search, no embeddings

)
result = rag_simple.retrieve_context("Latest LangChain version")

```

## Summary

- **Agentic Planning** – The LLM generates dynamic search strategies via `analyze_question` before executing any API calls.
- **Parallel Architecture** – Four concurrent search streams (code, issues, repos, web) maximize retrieval coverage.
- **Multi-Stage Ranking** – Results are scored by reranker models, embedding similarity, and LLM relevance judgments before final selection.
- **Modular Design** – `GithubRAG` orchestrates while `RAGProcessor` handles logic and `LLMManager` abstracts model providers.
- **Configurable Modes** – Full mode uses embeddings and rerankers; simple mode provides fast, lightweight retrieval via web search only.

## Frequently Asked Questions

### What is the difference between simple mode and full Agentic RAG?

**Simple mode** disables the embedding model and reranker, performing only Google/Jina web searches limited to GitHub domains. This provides faster responses with lower computational cost. **Full mode** activates the complete pipeline including LLM-driven strategy generation, GitHub API searches, token-aware chunking, and the three-stage ranking system (reranker + embeddings + LLM scoring).

### How does the ranking system combine multiple scoring methods?

The `retrieve_topn_contexts` method in [`rag_processor.py`](https://github.com/jetxu-llm/llama-github/blob/main/rag_processor.py) combines scores multiplicatively: `final_score = llm_relevance_score × cosine_similarity × reranker_score`. The reranker acts as a broad filter, the embedding model provides semantic similarity, and the LLM provides a final relevance judgment, ensuring high-precision context selection.

### Which LLM providers are supported by the LLMManager?

According to [`llama_github/llm_integration/initial_load.py`](https://github.com/jetxu-llm/llama-github/blob/main/llama_github/llm_integration/initial_load.py), the system supports **OpenAI** (GPT models), **Mistral AI**, and **local HuggingFace** models. The `LLMHandler` unifies these behind a single `ainvoke` interface, handling prompt templating and structured output parsing consistently across providers.

### How does the system handle large code files?

The `RAGProcessor._split_content_into_chunks` method uses `RecursiveCharacterTextSplitter` with token-aware chunking to break large files into manageable segments. Each chunk retains metadata linking it back to the source file URL, repository stars, and language, ensuring that even large files contribute relevant, attributable context to the final answer.