# What AI Models Does CodeWiki Use for Code Analysis? A Technical Deep Dive

> Discover the AI models powering CodeWiki's code analysis. We use Google Gemini 2.5 Pro and Nomic Embed Text via Ollama in a RAG pipeline for in-depth repository insights.

- Repository: [Luong Quang Dung/codewiki](https://github.com/quangdungluong/codewiki)
- Tags: deep-dive
- Published: 2026-02-16

---

**CodeWiki uses Google Gemini 2.5 Pro for large language model tasks and Nomic Embed Text via Ollama for vector embeddings, combining them in a RAG pipeline to analyze code repositories.**

The open-source project `quangdungluong/codewiki` implements a dual-model architecture that separates natural language generation from semantic code retrieval. This design leverages the strengths of each AI model: Gemini handles complex reasoning and explanation generation, while Nomic Embed Text creates dense vector representations for fast similarity search across repository files.

## Overview of CodeWiki's AI Architecture

CodeWiki's backend relies on two distinct AI services working in tandem. The **Google Gemini 2.5 Pro** model serves as the primary LLM for generating explanations, diagrams, and chat responses. The **Nomic Embed Text** model, running locally via Ollama, handles all embedding operations for the vector database.

This separation allows the system to optimize for both reasoning quality and retrieval speed. The architecture is provider-agnostic, with configuration managed through the `provider` and `model` fields in [`api/models.py`](https://github.com/quangdungluong/codewiki/blob/main/api/models.py), enabling swaps to alternative services like OpenAI or OpenRouter without structural changes.

## Google Gemini 2.5 Pro: The LLM Engine

### Role in Code Analysis

Gemini 2.5 Pro powers the natural language capabilities of CodeWiki. It generates structured explanations of code, creates architectural diagrams, and responds to user queries in the chat interface. The model processes repository context retrieved from the vector store and synthesizes coherent, technical answers.

### Implementation in gemini_service.py

The `GeminiService` class in [`api/services/gemini_service.py`](https://github.com/quangdungluong/codewiki/blob/main/api/services/gemini_service.py) encapsulates all interactions with the Google API. It initializes the client with the `gemini-2.5-pro` model and handles streaming responses for real-time chat.

```python

# api/services/gemini_service.py

from utils.format_message import format_message
from utils.logger import logger

class GeminiService:
    def __init__(self):
        self.client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
        self.model = "gemini-2.5-pro"          # ← Gemini LLM

    async def generate(self, system_prompt: str, data: dict):
        user_message = format_message(data)
        response = await self.client.aio.models.generate_content_stream(
            model=self.model,
            config=types.GenerateContentConfig(
                system_instruction=system_prompt,
                response_mime_type="application/json",
                max_output_tokens=12000,
            ),
            contents=user_message,
        )
        async for chunk in response:
            if chunk.text is not None:
                yield chunk.text

```

### Repository Structure Generation

Beyond chat, Gemini 2.5 Pro also drives the [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py) module, which generates structured XML descriptions of repository layouts. This allows CodeWiki to create high-level architectural overviews before diving into specific file implementations.

## Nomic Embed Text: The Vector Embedding Model

### Purpose and Provider

For semantic search capabilities, CodeWiki uses **Nomic Embed Text**, a high-quality open-source embedding model distributed through Ollama. This model converts code snippets and documentation into 768-dimensional vectors that capture semantic meaning, enabling similarity-based retrieval across the repository.

### Integration with Ollama

The embedding pipeline runs locally via the Ollama client, ensuring data privacy and reducing latency for retrieval operations. The `OllamaClient` from the AdalFlow library interfaces with the local instance to generate embeddings.

```python

# utils/document_pipeline.py (excerpt)

self.embedder = adalflow.Embedder(
    model_client=OllamaClient(),               # ← Ollama client

    model_kwargs={"model": "nomic-embed-text"} # ← embedding model

)

```

These embeddings feed into a FAISS vector store, which performs fast approximate nearest neighbor searches to retrieve relevant code context for the LLM.

## How the Models Work Together in RAG

CodeWiki implements a **Retrieval-Augmented Generation (RAG)** pipeline that combines both AI models. When a user submits a query, the system first uses Nomic Embed Text to retrieve relevant code snippets from the vector database. It then passes this retrieved context to Gemini 2.5 Pro, which generates a comprehensive, context-aware answer.

The [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py) file explicitly wires these components together:

```python

# api/rag.py

self.generator = adalflow.Generator(
    template=RAG_TEMPLATE,
    model_client=GoogleGenAIClient(),
    model_kwargs={"model": "gemini-2.5-pro", "temperature": 0.7},
    output_processors=data_parser,
)

self.embedder = adalflow.Embedder(
    model_client=OllamaClient(),
    model_kwargs={"model": "nomic-embed-text"},
)

```

This architecture ensures that responses are grounded in actual repository content while maintaining the conversational fluency of a large language model.

## Configuration and Model Selection

While CodeWiki defaults to Google Gemini for LLM tasks and Ollama for embeddings, the system is designed to be provider-agnostic. The [`api/models.py`](https://github.com/quangdungluong/codewiki/blob/main/api/models.py) file defines configuration fields for `provider` and `model`, allowing developers to swap in alternative services such as OpenAI, OpenRouter, or different Ollama-hosted models without modifying the core logic.

The default configuration sets the provider to `"google"` and the model to `"gemini-2.5-pro"`, but these values can be overridden through environment variables or request parameters to customize the AI models CodeWiki uses for code analysis.

## Summary

- **CodeWiki employs a dual-model architecture** combining Google Gemini 2.5 Pro for natural language generation and Nomic Embed Text via Ollama for vector embeddings.
- **Gemini 2.5 Pro** handles chat responses, code explanations, and diagram generation through the `GeminiService` class in [`api/services/gemini_service.py`](https://github.com/quangdungluong/codewiki/blob/main/api/services/gemini_service.py).
- **Nomic Embed Text** creates semantic embeddings for repository files, enabling fast similarity search via FAISS in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py).
- **The RAG pipeline** in [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py) integrates both models, retrieving relevant code context with embeddings before generating answers with the LLM.
- **Provider-agnostic configuration** in [`api/models.py`](https://github.com/quangdungluong/codewiki/blob/main/api/models.py) allows swapping AI models without structural changes to the codebase.

## Frequently Asked Questions

### Does CodeWiki support other LLM providers besides Google Gemini?

Yes, while CodeWiki defaults to Google Gemini 2.5 Pro, the architecture is provider-agnostic. The [`api/models.py`](https://github.com/quangdungluong/codewiki/blob/main/api/models.py) configuration allows specifying different providers such as OpenAI, OpenRouter, or Ollama-hosted models. The `GeminiService` class can be extended or replaced to accommodate alternative LLM clients without modifying the core RAG logic.

### Why does CodeWiki use Ollama for embeddings instead of Gemini?

CodeWiki uses Ollama to run the **nomic-embed-text** model locally for several technical advantages. Local embedding generation ensures data privacy by keeping repository code on the infrastructure, reduces latency for vector operations, and eliminates API costs for high-volume embedding tasks. The FAISS vector store in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) relies on these local embeddings for fast similarity search.

### Can I switch to a different embedding model in CodeWiki?

Yes, you can configure a different embedding model by modifying the `model_kwargs` in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) or [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py). The system uses AdalFlow's `Embedder` component with an `OllamaClient`, so any model available in your Ollama instance—such as `mxbai-embed-large` or `all-minilm`—can be substituted for `nomic-embed-text` without changing the retrieval logic.

### What is the context window size for the Gemini 2.5 Pro model in CodeWiki?

According to the implementation in [`api/services/gemini_service.py`](https://github.com/quangdungluong/codewiki/blob/main/api/services/gemini_service.py), CodeWiki configures the Gemini 2.5 Pro model with a **maximum output token limit of 12,000** (`max_output_tokens=12000`). This large context window enables the system to generate comprehensive code explanations, detailed architectural diagrams, and multi-file analysis responses while processing extensive retrieved context from the vector store.