Architecture of the FAISS Retriever for Code Search in Codewiki

The FAISS retriever in codewiki uses the adalflow library to embed queries via Ollama's nomic-embed-text model, searches a pre-built FAISS index of repository code vectors, and returns the top-k most similar documents for RAG-based code search.

The FAISS retriever for code search in the codewiki repository enables semantic retrieval of code snippets by leveraging Facebook AI Similarity Search (FAISS) for high-performance vector similarity operations. This architecture integrates the external adalflow library to handle embedding generation and nearest-neighbor search against a pre-computed index of repository code vectors.

Core Components of the FAISS Retriever

The retriever is instantiated from adalflow.components.retriever.faiss_retriever and configured in api/rag.py. It consists of four primary components that work together to enable efficient semantic code search.

Embedding Function and Vector Generation

The embedding function (embedder) converts query strings into dense vectors using the Ollama client with the nomic-embed-text model. In api/rag.py, this is configured as:

self.query_embedder = adalflow.Embedder(
    model_client=OllamaClient(),
    model_kwargs={"model": "nomic-embed-text"}
)

This embedder processes queries on-the-fly during the retrieval phase, ensuring that user questions are transformed into the same vector space as the pre-computed code embeddings.

Document Store and Vector Mapping

The document store consists of a list of transformed Document objects created by LocalDBManager.prepare_db in utils/localdb_manager.py. Each document carries a pre-computed vector accessible via the .vector attribute.

The vector mapping function (document_map_func) extracts these vectors during index construction. In codewiki, this is implemented as a simple lambda:

document_map_func=lambda doc: doc.vector

FAISS Index Structure

Internally, FAISSRetriever builds a FAISS index (such as IndexFlatL2 or an IVF variant) from the supplied document vectors. This index enables fast nearest-neighbor search in high-dimensional space, typically returning results in milliseconds even with thousands of code snippets.

Data Flow and Retrieval Pipeline

The architecture follows a classic embed-index-search pattern that separates ingestion from query-time operations.

Repository Ingestion and Document Preparation

When a user submits a repository URL, LocalDBManager.prepare_db in utils/localdb_manager.py parses the codebase using RepositoryParser, chunks files into meaningful segments, and computes embeddings for each chunk using the same nomic-embed-text model. These transformed documents, complete with their vectors, are stored in memory for the retriever to consume.

During a search operation in RAG.call (defined in api/rag.py), the retriever executes three steps:

  1. Embeds the query string via self.embedder
  2. Searches the FAISS index for the top_k closest vectors (defaulting to 20)
  3. Returns a RetrieverOutput object containing doc_indices, scores, and placeholder documents

Document Resolution and RAG Integration

After retrieval, RAG.call resolves the document IDs to full Document objects by indexing into self.transformed_docs. These retrieved code snippets are then formatted and injected into the LLM prompt (using Gemini-2.5-pro) to generate contextually grounded answers about the codebase.

Implementation in Codewiki

The integration points in api/rag.py demonstrate how the abstract FAISS retriever is concretely instantiated for code search workflows.

Initializing the Retriever in api/rag.py

The RAG.prepare_retriever method constructs the retriever with all necessary dependencies:

from adalflow.components.retriever.faiss_retriever import FAISSRetriever

def prepare_retriever(self, repo_url: str, type: str = "github", access_token: str = None):
    # Load and embed all repo documents

    self.transformed_docs = self.db_manager.prepare_db(
        repo_url=repo_url, access_token=access_token
    )
    # Build the FAISS retriever

    self.retriever = FAISSRetriever(
        embedder=self.query_embedder,   # embeds queries on‑the‑fly

        top_k=20,                      # number of docs to return

        documents=self.transformed_docs,
        document_map_func=lambda doc: doc.vector,  # extract stored vectors

    )

Performing Code Search Queries

The retrieval operation is invoked within RAG.call, which handles the end-to-end search process:

def call(self, query: str):
    # Run the FAISS index lookup

    retrieved_documents: List[RetrieverOutput] = self.retriever(query)

    # Resolve IDs → full Document objects

    retrieved_documents[0].documents = [
        self.transformed_docs[doc_id]
        for doc_id in retrieved_documents[0].doc_indices
    ]
    return retrieved_documents

This pattern allows the system to separate the high-performance vector search (handled by FAISS) from the document enrichment and LLM prompting logic.

Key Files and Their Roles

File Role
api/rag.py Instantiates the FAISSRetriever, prepares the document store, and runs queries.
utils/localdb_manager.py Loads a repository, parses files into Document objects, and computes their embeddings/vectors.
utils/repository_parser.py Helper for extracting source code and metadata from a GitHub repo.
api/main.py & other FastAPI routes Expose the RAG service that ultimately calls the FAISSRetriever.

These files together compose the end‑to‑end pipeline: repo → documents → vectors → FAISS index → retrieval → LLM answer.

Summary

  • The FAISS retriever for code search in codewiki is provided by the adalflow library and imported in api/rag.py.
  • It uses Ollama with the nomic-embed-text model to embed queries and pre-computes vectors for repository code chunks.
  • The retriever builds a FAISS index (such as IndexFlatL2) from document vectors to enable fast nearest-neighbor search.
  • By default, the system retrieves the top 20 most similar code snippets, which are then resolved to full documents and fed into a Gemini-2.5-pro LLM prompt.

Frequently Asked Questions

The retriever uses the nomic-embed-text model served via Ollama to generate dense vector representations. This model is configured in api/rag.py through the adalflow.Embedder class, which handles both query embedding at search time and document embedding during the ingestion phase.

How does the FAISS retriever handle repository ingestion?

Repository ingestion is handled by LocalDBManager.prepare_db in utils/localdb_manager.py. This process parses the GitHub repository using RepositoryParser, chunks the source code into meaningful segments, computes embeddings for each chunk using the same Ollama embedder, and returns a list of Document objects with pre-computed .vector attributes ready for FAISS indexing.

What is the default top_k value for retrieved documents?

The default top_k value is 20, meaning the retriever returns the 20 most similar code snippets by default. This is configured during the instantiation of FAISSRetriever in the prepare_retriever method of api/rag.py, though this parameter can be adjusted based on specific retrieval requirements.

Can the FAISS retriever work with other embedding providers besides Ollama?

Yes, the architecture supports swapping the embedding provider. The FAISSRetriever accepts any embedder object that follows the adalflow Embedder interface. While codewiki currently uses Ollama with nomic-embed-text, you could configure the embedder to use OpenAI, Cohere, or other providers by changing the model_client and model_kwargs passed to adalflow.Embedder in api/rag.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →