RAG Implementation Architecture in CodeWiki: A Deep Dive into the Modular Retrieval Pipeline

CodeWiki implements a modular Retrieval-Augmented Generation (RAG) architecture that combines FAISS vector storage, Ollama embeddings, and Google Gemini 2.5-Pro to answer questions about code repositories with retrieved context.

The quangdungluong/codewiki repository provides a production-ready implementation of the RAG implementation architecture in CodeWiki, designed specifically for querying software repositories. This system ingests source code, indexes it into a searchable vector database, and augments LLM responses with relevant retrieved snippets to provide accurate, context-aware answers about codebases.

Document Ingestion and Indexing Pipeline

The ingestion layer transforms raw repository files into searchable vector embeddings through a multi-stage pipeline implemented across several utility modules.

Repository Discovery and Download

The process begins in utils.repository_structure.RepositoryStructureFetcher, which determines whether the target is a local path or remote GitHub URL via fetch_repository_structure. For remote repositories, utils.repo_downloader.RepoDownloader clones the code into ./.cache/repos/<repo-name> for local processing.

Document Processing and Chunking

Once local, utils.document_pipeline.RecursiveDocumentReader walks the directory tree, applying inclusion and exclusion filters to identify relevant source files. It creates adalflow.core.types.Document objects containing raw text and metadata including file_path, is_implementation, type, is_code, and token_count.

The DocumentTransformer then processes these documents using a two-stage approach:

  • TextSplitter divides content into 500-word chunks with 100-word overlap to maintain context boundaries
  • OllamaDocumentProcessor prepares chunks for embedding

Embedding and Vector Storage

The pipeline uses adalflow.Embedder backed by Ollama with the nomic-embed-text model to generate vector representations. utils.localdb_manager.LocalDBManager persists these embeddings to a local file-based store (*.pkl) via save_state, enabling fast subsequent loads through load_state. The prepare_db method orchestrates the entire flow from repository fetch to persistent storage.

Retrieval Layer Architecture

When a user submits a query, the retrieval layer transforms the question into vector space and identifies relevant code contexts.

In api.rag.RAG.prepare_retriever, the system initializes a FAISSRetriever from adalflow.components.retriever.faiss_retriever. The retriever uses the same Ollama embedder (query_embedder) to convert the raw query string into an embedding vector.

The retriever performs a nearest-neighbor search against the indexed vector store with top_k=20, returning the most relevant document chunks based on cosine similarity.

Context Assembly

After retrieval, the system reattaches the original Document objects to the RetrieverOutput via retrieved_documents[0].documents. This ensures the generation layer receives not just similarity scores but the full text and metadata of the retrieved code snippets.

Generation and Memory Management

The generation layer synthesizes answers by combining retrieved contexts with conversational history and system instructions.

Prompt Construction with RAG_TEMPLATE

The RAG class in api/rag.py constructs prompts using RAG_TEMPLATE, which stitches together:

  • system_prompt: Guidelines for answer formatting and code explanation style
  • conversation_history: Prior dialog turns stored in memory
  • contexts: The retrieved code chunks from FAISS
  • input_str: The current user query

LLM Integration with Gemini 2.5-Pro

The RAG.generator is an adalflow.Generator configured to call Google Gemini 2.5-Pro via model_kwargs={"model": "gemini-2.5-pro", ...}. A DataClassParser enforces the output schema defined by RAGAnswer, which includes fields for rationale and answer.

The RAG.call method orchestrates the full inference flow: embedding the query, retrieving contexts, building the prompt, invoking the LLM, and parsing the structured response. Error handling ensures that if any step fails, a fallback RAGAnswer with an error message is returned.

Conversational Memory

The system maintains dialog context through Memory and CustomConversation classes. Each interaction is recorded as a DialogTurn containing the UserQuery and AssistantResponse. This history is injected into subsequent prompts via conversation_history, enabling multi-turn conversations about the codebase.

End-to-End Data Flow

The complete RAG pipeline follows this sequence:

  1. Ingestion: RepositoryStructureFetcher → RepoDownloader → RecursiveDocumentReader → DocumentTransformer → LocalDBManager
  2. Indexing: TextSplitter (500-word chunks, 100-word overlap) → OllamaDocumentProcessor (nomic-embed-text) → FAISS vector store (*.pkl)
  3. Query: User input → query_embedder (Ollama) → FAISSRetriever (top_k=20) → context assembly
  4. Generation: RAG_TEMPLATE (system + history + contexts) → adalflow.Generator (Gemini 2.5-Pro) → RAGAnswer parsing
  5. Memory: DialogTurn recording → CustomConversation → next-turn context injection

Key Source Files and Components

File Purpose
api/rag.py Core RAG orchestration – embedder, retriever, generator, and memory management
utils/localdb_manager.py Vector database lifecycle – creation, persistence, and loading of FAISS stores
utils/document_pipeline.py Document reading, chunking, and embedding preparation
api/repository_structure.py Repository metadata and structure fetching
utils/repo_downloader.py GitHub repository cloning and local caching
api/models.py Pydantic models (WikiCacheData, WikiStructureModel, RAGAnswer) used throughout the pipeline

Summary

  • CodeWiki implements a three-layer RAG architecture comprising ingestion/indexing, retrieval, and generation with memory.
  • The system uses Ollama (nomic-embed-text) for embeddings and FAISS for vector storage, persisted locally via LocalDBManager.
  • Retrieval employs FAISSRetriever with cosine similarity search (top_k=20) to find relevant code chunks.
  • Generation combines retrieved contexts with conversation history via RAG_TEMPLATE and synthesizes answers using Google Gemini 2.5-Pro.
  • The modular design allows easy swapping of embedders, vector stores, or LLMs without disrupting the pipeline.

Frequently Asked Questions

What embedding model does CodeWiki use for RAG?

CodeWiki uses the nomic-embed-text model via Ollama for generating embeddings. This is implemented in utils/document_pipeline.py through the OllamaDocumentProcessor class and the adalflow.Embedder component, ensuring consistent embedding generation for both document indexing and query retrieval.

How does CodeWiki handle large code repositories?

The system handles large repositories through chunking and persistent caching. The TextSplitter in utils/document_pipeline.py splits documents into 500-word chunks with 100-word overlap, while LocalDBManager persists the FAISS vector store to disk (*.pkl files). This allows the system to process repositories incrementally and load pre-computed embeddings on subsequent runs without re-indexing.

Can CodeWiki work with local repositories?

Yes, CodeWiki supports both local and remote repositories. The RepositoryStructureFetcher.fetch_repository_structure method in api/repository_structure.py detects whether the input is a local path or a GitHub URL. For local repositories, the system skips the download step and directly processes the files from the specified directory path.

What LLM powers the generation layer in CodeWiki?

The generation layer uses Google Gemini 2.5-Pro, configured through the adalflow.Generator in api/rag.py. The RAG.generator component sends constructed prompts containing system instructions, retrieved contexts, and conversation history to the Gemini API, with output parsing enforced by the DataClassParser for structured RAGAnswer responses.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →