# RAG Implementation Architecture in CodeWiki: A Deep Dive into the Modular Retrieval Pipeline

> Explore the RAG implementation architecture in CodeWiki. Learn how FAISS Ollama and Gemini 2.5 Pro power its modular retrieval pipeline for code repository Q&A.

- Repository: [Luong Quang Dung/codewiki](https://github.com/quangdungluong/codewiki)
- Tags: architecture
- Published: 2026-02-16

---

**CodeWiki implements a modular Retrieval-Augmented Generation (RAG) architecture that combines FAISS vector storage, Ollama embeddings, and Google Gemini 2.5-Pro to answer questions about code repositories with retrieved context.**

The `quangdungluong/codewiki` repository provides a production-ready implementation of the RAG implementation architecture in CodeWiki, designed specifically for querying software repositories. This system ingests source code, indexes it into a searchable vector database, and augments LLM responses with relevant retrieved snippets to provide accurate, context-aware answers about codebases.

## Document Ingestion and Indexing Pipeline

The ingestion layer transforms raw repository files into searchable vector embeddings through a multi-stage pipeline implemented across several utility modules.

### Repository Discovery and Download

The process begins in `utils.repository_structure.RepositoryStructureFetcher`, which determines whether the target is a local path or remote GitHub URL via `fetch_repository_structure`. For remote repositories, `utils.repo_downloader.RepoDownloader` clones the code into `./.cache/repos/<repo-name>` for local processing.

### Document Processing and Chunking

Once local, `utils.document_pipeline.RecursiveDocumentReader` walks the directory tree, applying inclusion and exclusion filters to identify relevant source files. It creates `adalflow.core.types.Document` objects containing raw text and metadata including `file_path`, `is_implementation`, `type`, `is_code`, and `token_count`.

The `DocumentTransformer` then processes these documents using a two-stage approach:

- `TextSplitter` divides content into 500-word chunks with 100-word overlap to maintain context boundaries
- `OllamaDocumentProcessor` prepares chunks for embedding

### Embedding and Vector Storage

The pipeline uses `adalflow.Embedder` backed by **Ollama** with the `nomic-embed-text` model to generate vector representations. `utils.localdb_manager.LocalDBManager` persists these embeddings to a local file-based store (`*.pkl`) via `save_state`, enabling fast subsequent loads through `load_state`. The `prepare_db` method orchestrates the entire flow from repository fetch to persistent storage.

## Retrieval Layer Architecture

When a user submits a query, the retrieval layer transforms the question into vector space and identifies relevant code contexts.

### Query Embedding and FAISS Search

In `api.rag.RAG.prepare_retriever`, the system initializes a `FAISSRetriever` from `adalflow.components.retriever.faiss_retriever`. The retriever uses the same Ollama embedder (`query_embedder`) to convert the raw query string into an embedding vector.

The retriever performs a nearest-neighbor search against the indexed vector store with `top_k=20`, returning the most relevant document chunks based on cosine similarity.

### Context Assembly

After retrieval, the system reattaches the original `Document` objects to the `RetrieverOutput` via `retrieved_documents[0].documents`. This ensures the generation layer receives not just similarity scores but the full text and metadata of the retrieved code snippets.

## Generation and Memory Management

The generation layer synthesizes answers by combining retrieved contexts with conversational history and system instructions.

### Prompt Construction with RAG_TEMPLATE

The `RAG` class in [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py) constructs prompts using `RAG_TEMPLATE`, which stitches together:

- `system_prompt`: Guidelines for answer formatting and code explanation style
- `conversation_history`: Prior dialog turns stored in memory
- `contexts`: The retrieved code chunks from FAISS
- `input_str`: The current user query

### LLM Integration with Gemini 2.5-Pro

The `RAG.generator` is an `adalflow.Generator` configured to call **Google Gemini 2.5-Pro** via `model_kwargs={"model": "gemini-2.5-pro", ...}`. A `DataClassParser` enforces the output schema defined by `RAGAnswer`, which includes fields for `rationale` and `answer`.

The `RAG.call` method orchestrates the full inference flow: embedding the query, retrieving contexts, building the prompt, invoking the LLM, and parsing the structured response. Error handling ensures that if any step fails, a fallback `RAGAnswer` with an error message is returned.

### Conversational Memory

The system maintains dialog context through `Memory` and `CustomConversation` classes. Each interaction is recorded as a `DialogTurn` containing the `UserQuery` and `AssistantResponse`. This history is injected into subsequent prompts via `conversation_history`, enabling multi-turn conversations about the codebase.

## End-to-End Data Flow

The complete RAG pipeline follows this sequence:

1. **Ingestion**: `RepositoryStructureFetcher` → `RepoDownloader` → `RecursiveDocumentReader` → `DocumentTransformer` → `LocalDBManager`
2. **Indexing**: `TextSplitter` (500-word chunks, 100-word overlap) → `OllamaDocumentProcessor` (`nomic-embed-text`) → FAISS vector store (`*.pkl`)
3. **Query**: User input → `query_embedder` (Ollama) → `FAISSRetriever` (`top_k=20`) → context assembly
4. **Generation**: `RAG_TEMPLATE` (system + history + contexts) → `adalflow.Generator` (Gemini 2.5-Pro) → `RAGAnswer` parsing
5. **Memory**: `DialogTurn` recording → `CustomConversation` → next-turn context injection

## Key Source Files and Components

| File | Purpose |
|------|---------|
| [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py) | Core RAG orchestration – embedder, retriever, generator, and memory management |
| [`utils/localdb_manager.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/localdb_manager.py) | Vector database lifecycle – creation, persistence, and loading of FAISS stores |
| [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) | Document reading, chunking, and embedding preparation |
| [`api/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/api/repository_structure.py) | Repository metadata and structure fetching |
| [`utils/repo_downloader.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repo_downloader.py) | GitHub repository cloning and local caching |
| [`api/models.py`](https://github.com/quangdungluong/codewiki/blob/main/api/models.py) | Pydantic models (`WikiCacheData`, `WikiStructureModel`, `RAGAnswer`) used throughout the pipeline |

## Summary

- CodeWiki implements a **three-layer RAG architecture** comprising ingestion/indexing, retrieval, and generation with memory.
- The system uses **Ollama** (`nomic-embed-text`) for embeddings and **FAISS** for vector storage, persisted locally via `LocalDBManager`.
- Retrieval employs **FAISSRetriever** with cosine similarity search (`top_k=20`) to find relevant code chunks.
- Generation combines retrieved contexts with conversation history via **RAG_TEMPLATE** and synthesizes answers using **Google Gemini 2.5-Pro**.
- The modular design allows easy swapping of embedders, vector stores, or LLMs without disrupting the pipeline.

## Frequently Asked Questions

### What embedding model does CodeWiki use for RAG?

CodeWiki uses the `nomic-embed-text` model via **Ollama** for generating embeddings. This is implemented in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) through the `OllamaDocumentProcessor` class and the `adalflow.Embedder` component, ensuring consistent embedding generation for both document indexing and query retrieval.

### How does CodeWiki handle large code repositories?

The system handles large repositories through **chunking** and **persistent caching**. The `TextSplitter` in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) splits documents into 500-word chunks with 100-word overlap, while `LocalDBManager` persists the FAISS vector store to disk (`*.pkl` files). This allows the system to process repositories incrementally and load pre-computed embeddings on subsequent runs without re-indexing.

### Can CodeWiki work with local repositories?

Yes, CodeWiki supports both local and remote repositories. The `RepositoryStructureFetcher.fetch_repository_structure` method in [`api/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/api/repository_structure.py) detects whether the input is a local path or a GitHub URL. For local repositories, the system skips the download step and directly processes the files from the specified directory path.

### What LLM powers the generation layer in CodeWiki?

The generation layer uses **Google Gemini 2.5-Pro**, configured through the `adalflow.Generator` in [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py). The `RAG.generator` component sends constructed prompts containing system instructions, retrieved contexts, and conversation history to the Gemini API, with output parsing enforced by the `DataClassParser` for structured `RAGAnswer` responses.