How the Agentic RAG Architecture Works in llama-github: A Technical Deep Dive
The Agentic RAG architecture in llama-github orchestrates LLM-driven query planning, parallel GitHub API searches, and multi-stage ranking to retrieve and synthesize context from repositories.
The llama-github library implements a sophisticated Agentic Retrieval-Augmented Generation (RAG) pipeline that treats the LLM as an active participant in the retrieval process. Unlike traditional RAG systems that rely on static retrieval strategies, this architecture dynamically generates search criteria, evaluates result relevance through multiple scoring layers, and synthesizes answers from GitHub source code, issues, and repositories. The system is designed around three core Python modules that coordinate authentication, retrieval, and language model interactions.
Core Components of the Agentic RAG System
GithubRAG: The High-Level Orchestrator
The GithubRAG class in llama_github/github_rag.py serves as the primary entry point for the entire pipeline. During initialization (lines 71‑108), it constructs four essential subsystems: a GitHubAuthManager for API authentication, a RepositoryPool for caching, an LLMManager for model handling, and a RAGProcessor for search execution. The orchestrator exposes two main async interfaces: async_retrieve_context for raw context retrieval and async_answer_with_context for complete answer generation.
RAGProcessor: Search Strategy and Ranking Engine
Located in llama_github/rag_processing/rag_processor.py, the RAGProcessor contains the agentic logic. It implements analyze_question (lines 51‑66) to generate LLM-driven search strategies, methods like get_code_search_criteria to produce structured Pydantic search queries, and retrieve_topn_contexts for multi-stage result ranking. This component also handles content chunking via RecursiveCharacterTextSplitter and arranges raw GitHub API responses into uniform context objects.
LLMManager and LLMHandler: Unified Model Abstraction
The LLMManager (llama_github/llm_integration/initial_load.py) and LLMHandler (llama_github/llm_integration/llm_handler.py) provide a unified interface for cloud and local models. The manager initializes embedding models, rerankers, and inference endpoints for OpenAI, Mistral, or HuggingFace deployments. The handler exposes ainvoke, an async method that standardizes prompt templating and structured output parsing across all supported providers.
The Agentic RAG Pipeline Step-by-Step
1. System Initialization and Construction
When instantiating GithubRAG, the __init__ method wires together all dependencies. The system loads configuration from config.json, which stores prompt templates, model names, and thresholds like top_n_contexts and min_stars_to_keep_result. The RAGProcessor receives injected instances of GitHubAPIHandler and LLMManager, establishing the dependency chain for subsequent operations.
2. Query Analysis and Strategy Generation
Upon receiving a user query, the pipeline enters an agentic planning phase. The analyze_question method calls the LLM with the always_answer_prompt from configuration, returning a structured tuple: [question, draft_answer, code_logic, issue_logic]. This LLM-driven analysis determines what to search and how to combine different GitHub data sources (code, issues, repositories).
3. Parallel Retrieval Execution
The async_retrieve_context method (lines 154‑188 of github_rag.py) executes four concurrent tasks using asyncio.gather:
analyze_question– Generates the search strategy.google_search_retrieval– Performs web search limited tosite:github.com.code_search_retrieval– Queries the GitHub Code Search API.issue_search_retrieval– Queries the GitHub Issues API.repo_search_retrieval– Queries the GitHub Repository API.
Each search method receives criteria generated by dedicated LLM calls (e.g., get_repo_search_criteria), which return structured Pydantic models like _GitHubRepoSearchCriteria containing search terms and necessity scores.
4. Context Chunking and Arrangement
Raw API results undergo processing in RAGProcessor.arrange_context and its helper methods (_arrange_code_search_result, _arrange_issue_search_result). Long files and issue bodies are split into token-aware chunks using RecursiveCharacterTextSplitter. The system enriches each chunk with metadata including repository stars, URL, programming language, and timestamps, creating uniform context dictionaries for downstream ranking.
5. Multi-Stage Ranking and Scoring
The retrieve_topn_contexts method implements a three-layer ranking system:
- Reranker Scoring – A sentence-pair classification model scores each
(query, document)pair. - Embedding Similarity – The top
3 × ncandidates are re-scored using cosine similarity from the embedding model (llm_manager.get_embedding_model()). - LLM Relevance Scoring – A lightweight LLM call (
_ContextRelevanceScore) computes a context relevance score.
The final ranking uses a composite formula: llm_score × cos_sim × rerank_score. The top n contexts are returned to the caller.
6. Final Answer Synthesis
In async_answer_with_context, the system retrieves contexts (if not provided) and invokes LLMHandler.ainvoke with the general_prompt template. The structured prompt includes the retrieved chunks and their metadata, instructing the LLM to generate a cited answer that references specific GitHub URLs.
What Makes the Architecture "Agentic"
LLM-Driven Planning – The initial analyze_question call determines search strategy, effectively letting the model decide what information is needed before executing API calls.
Adaptive Search Depth – The necessity_score in repository search criteria allows the system to skip expensive repo-level searches when the LLM judges them unnecessary.
Dynamic Reranking – Rather than relying solely on static embeddings, the system employs a dedicated reranker model and a relevance-scoring LLM to judge context quality, creating a feedback loop where the model evaluates its own retrieval choices.
Structured Decision Making – Search criteria are generated as structured Pydantic objects (_GitHubCodeSearchCriteria, etc.), enabling the LLM to output precise API parameters rather than free-text queries.
Implementation Examples
Initialize the Full Agentic Pipeline
from llama_github.github_rag import GithubRAG
rag = GithubRAG(
github_access_token="ghp_YOUR_TOKEN",
openai_api_key="sk-YOUR_KEY",
simple_mode=False # Enable full embedding + rerank pipeline
)
Retrieve Context Only
contexts = await rag.async_retrieve_context(
"How does React handle concurrent rendering?"
)
# Returns list of dicts: {"context": "<chunk>", "url": "<github url>"}
Generate a Complete Answer
answer = await rag.async_answer_with_context(
query="Explain the RAGProcessor chunking strategy"
)
print(answer)
Enable Simple Mode for Fast Queries
rag_simple = GithubRAG(
github_access_token="ghp_TOKEN",
simple_mode=True # Uses only Google/Jina search, no embeddings
)
result = rag_simple.retrieve_context("Latest LangChain version")
Summary
- Agentic Planning – The LLM generates dynamic search strategies via
analyze_questionbefore executing any API calls. - Parallel Architecture – Four concurrent search streams (code, issues, repos, web) maximize retrieval coverage.
- Multi-Stage Ranking – Results are scored by reranker models, embedding similarity, and LLM relevance judgments before final selection.
- Modular Design –
GithubRAGorchestrates whileRAGProcessorhandles logic andLLMManagerabstracts model providers. - Configurable Modes – Full mode uses embeddings and rerankers; simple mode provides fast, lightweight retrieval via web search only.
Frequently Asked Questions
What is the difference between simple mode and full Agentic RAG?
Simple mode disables the embedding model and reranker, performing only Google/Jina web searches limited to GitHub domains. This provides faster responses with lower computational cost. Full mode activates the complete pipeline including LLM-driven strategy generation, GitHub API searches, token-aware chunking, and the three-stage ranking system (reranker + embeddings + LLM scoring).
How does the ranking system combine multiple scoring methods?
The retrieve_topn_contexts method in rag_processor.py combines scores multiplicatively: final_score = llm_relevance_score × cosine_similarity × reranker_score. The reranker acts as a broad filter, the embedding model provides semantic similarity, and the LLM provides a final relevance judgment, ensuring high-precision context selection.
Which LLM providers are supported by the LLMManager?
According to llama_github/llm_integration/initial_load.py, the system supports OpenAI (GPT models), Mistral AI, and local HuggingFace models. The LLMHandler unifies these behind a single ainvoke interface, handling prompt templating and structured output parsing consistently across providers.
How does the system handle large code files?
The RAGProcessor._split_content_into_chunks method uses RecursiveCharacterTextSplitter with token-aware chunking to break large files into manageable segments. Each chunk retains metadata linking it back to the source file URL, repository stars, and language, ensuring that even large files contribute relevant, attributable context to the final answer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →