What AI Models Does CodeWiki Use for Code Analysis? A Technical Deep Dive
CodeWiki uses Google Gemini 2.5 Pro for large language model tasks and Nomic Embed Text via Ollama for vector embeddings, combining them in a RAG pipeline to analyze code repositories.
The open-source project quangdungluong/codewiki implements a dual-model architecture that separates natural language generation from semantic code retrieval. This design leverages the strengths of each AI model: Gemini handles complex reasoning and explanation generation, while Nomic Embed Text creates dense vector representations for fast similarity search across repository files.
Overview of CodeWiki's AI Architecture
CodeWiki's backend relies on two distinct AI services working in tandem. The Google Gemini 2.5 Pro model serves as the primary LLM for generating explanations, diagrams, and chat responses. The Nomic Embed Text model, running locally via Ollama, handles all embedding operations for the vector database.
This separation allows the system to optimize for both reasoning quality and retrieval speed. The architecture is provider-agnostic, with configuration managed through the provider and model fields in api/models.py, enabling swaps to alternative services like OpenAI or OpenRouter without structural changes.
Google Gemini 2.5 Pro: The LLM Engine
Role in Code Analysis
Gemini 2.5 Pro powers the natural language capabilities of CodeWiki. It generates structured explanations of code, creates architectural diagrams, and responds to user queries in the chat interface. The model processes repository context retrieved from the vector store and synthesizes coherent, technical answers.
Implementation in gemini_service.py
The GeminiService class in api/services/gemini_service.py encapsulates all interactions with the Google API. It initializes the client with the gemini-2.5-pro model and handles streaming responses for real-time chat.
# api/services/gemini_service.py
from utils.format_message import format_message
from utils.logger import logger
class GeminiService:
def __init__(self):
self.client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
self.model = "gemini-2.5-pro" # ← Gemini LLM
async def generate(self, system_prompt: str, data: dict):
user_message = format_message(data)
response = await self.client.aio.models.generate_content_stream(
model=self.model,
config=types.GenerateContentConfig(
system_instruction=system_prompt,
response_mime_type="application/json",
max_output_tokens=12000,
),
contents=user_message,
)
async for chunk in response:
if chunk.text is not None:
yield chunk.text
Repository Structure Generation
Beyond chat, Gemini 2.5 Pro also drives the utils/repository_structure.py module, which generates structured XML descriptions of repository layouts. This allows CodeWiki to create high-level architectural overviews before diving into specific file implementations.
Nomic Embed Text: The Vector Embedding Model
Purpose and Provider
For semantic search capabilities, CodeWiki uses Nomic Embed Text, a high-quality open-source embedding model distributed through Ollama. This model converts code snippets and documentation into 768-dimensional vectors that capture semantic meaning, enabling similarity-based retrieval across the repository.
Integration with Ollama
The embedding pipeline runs locally via the Ollama client, ensuring data privacy and reducing latency for retrieval operations. The OllamaClient from the AdalFlow library interfaces with the local instance to generate embeddings.
# utils/document_pipeline.py (excerpt)
self.embedder = adalflow.Embedder(
model_client=OllamaClient(), # ← Ollama client
model_kwargs={"model": "nomic-embed-text"} # ← embedding model
)
These embeddings feed into a FAISS vector store, which performs fast approximate nearest neighbor searches to retrieve relevant code context for the LLM.
How the Models Work Together in RAG
CodeWiki implements a Retrieval-Augmented Generation (RAG) pipeline that combines both AI models. When a user submits a query, the system first uses Nomic Embed Text to retrieve relevant code snippets from the vector database. It then passes this retrieved context to Gemini 2.5 Pro, which generates a comprehensive, context-aware answer.
The api/rag.py file explicitly wires these components together:
# api/rag.py
self.generator = adalflow.Generator(
template=RAG_TEMPLATE,
model_client=GoogleGenAIClient(),
model_kwargs={"model": "gemini-2.5-pro", "temperature": 0.7},
output_processors=data_parser,
)
self.embedder = adalflow.Embedder(
model_client=OllamaClient(),
model_kwargs={"model": "nomic-embed-text"},
)
This architecture ensures that responses are grounded in actual repository content while maintaining the conversational fluency of a large language model.
Configuration and Model Selection
While CodeWiki defaults to Google Gemini for LLM tasks and Ollama for embeddings, the system is designed to be provider-agnostic. The api/models.py file defines configuration fields for provider and model, allowing developers to swap in alternative services such as OpenAI, OpenRouter, or different Ollama-hosted models without modifying the core logic.
The default configuration sets the provider to "google" and the model to "gemini-2.5-pro", but these values can be overridden through environment variables or request parameters to customize the AI models CodeWiki uses for code analysis.
Summary
- CodeWiki employs a dual-model architecture combining Google Gemini 2.5 Pro for natural language generation and Nomic Embed Text via Ollama for vector embeddings.
- Gemini 2.5 Pro handles chat responses, code explanations, and diagram generation through the
GeminiServiceclass inapi/services/gemini_service.py. - Nomic Embed Text creates semantic embeddings for repository files, enabling fast similarity search via FAISS in
utils/document_pipeline.py. - The RAG pipeline in
api/rag.pyintegrates both models, retrieving relevant code context with embeddings before generating answers with the LLM. - Provider-agnostic configuration in
api/models.pyallows swapping AI models without structural changes to the codebase.
Frequently Asked Questions
Does CodeWiki support other LLM providers besides Google Gemini?
Yes, while CodeWiki defaults to Google Gemini 2.5 Pro, the architecture is provider-agnostic. The api/models.py configuration allows specifying different providers such as OpenAI, OpenRouter, or Ollama-hosted models. The GeminiService class can be extended or replaced to accommodate alternative LLM clients without modifying the core RAG logic.
Why does CodeWiki use Ollama for embeddings instead of Gemini?
CodeWiki uses Ollama to run the nomic-embed-text model locally for several technical advantages. Local embedding generation ensures data privacy by keeping repository code on the infrastructure, reduces latency for vector operations, and eliminates API costs for high-volume embedding tasks. The FAISS vector store in utils/document_pipeline.py relies on these local embeddings for fast similarity search.
Can I switch to a different embedding model in CodeWiki?
Yes, you can configure a different embedding model by modifying the model_kwargs in utils/document_pipeline.py or api/rag.py. The system uses AdalFlow's Embedder component with an OllamaClient, so any model available in your Ollama instance—such as mxbai-embed-large or all-minilm—can be substituted for nomic-embed-text without changing the retrieval logic.
What is the context window size for the Gemini 2.5 Pro model in CodeWiki?
According to the implementation in api/services/gemini_service.py, CodeWiki configures the Gemini 2.5 Pro model with a maximum output token limit of 12,000 (max_output_tokens=12000). This large context window enables the system to generate comprehensive code explanations, detailed architectural diagrams, and multi-file analysis responses while processing extensive retrieved context from the vector store.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →