# How Roo Code's Codebase Indexing System Enables Semantic Search

> Discover how Roo Code's codebase indexing enables semantic search. Our system embeds code into vectors, allowing natural language queries to find relevant snippets instantly.

- Repository: [Roo Code/Roo-Code](https://github.com/RooCodeInc/Roo-Code)
- Tags: internals
- Published: 2026-04-26

---

**Roo Code indexes your entire repository by parsing abstract syntax trees into structured code blocks, embedding them as high-dimensional vectors, and querying them with natural language to surface semantically relevant code snippets.**

The Roo Code extension transforms static codebases into intelligent, queryable knowledge graphs through its sophisticated codebase indexing system. By combining Tree-sitter parsing with vector embeddings, Roo Code enables AI assistants to locate functionality based on meaning rather than literal text matches. This architecture splits the process into two tightly-coupled stages: index creation and semantic query execution.

## The Two-Stage Architecture

Roo Code's semantic search relies on a pipeline that first builds a comprehensive vector index and then queries it using natural language. This separation ensures that searches are meaning-based rather than literal-text, guaranteeing high recall for synonyms, paraphrases, or incomplete code fragments.

### Stage 1: AST-Aware Index Creation

The indexing process begins in [`src/services/code-index/processors/parser.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/services/code-index/processors/parser.ts), where the `CodeParser` class processes each supported source file. For every file discovered while walking the workspace (respecting `.rooignore`), the system calls `codeParser.parseFile(filePath)` to extract searchable code blocks.

The parser leverages Tree-sitter grammars through `loadRequiredLanguageParsers`, which lazily loads language-specific parsers to walk the syntax tree and capture named nodes such as functions, classes, and methods. Each captured block is transformed into a deterministic **segment hash** using the formula `sha256(filePath-start-end-len-preview)` and a **file hash** for deduplication, ensuring that re-indexing the same code does not create duplicate vectors.

When a language lacks Tree-sitter support or when captures exceed `MAX_BLOCK_CHARS * MAX_CHARS_TOLERANCE_FACTOR`, the system employs a **fallback chunker** that splits content using `_chunkLeafNodeByLines` or `_chunkTextByLines`. These methods respect whole-line boundaries while enforcing minimum and maximum character limits to ensure each chunk carries sufficient context for reliable similarity matching.

Markdown files receive special treatment in [`src/services/tree-sitter/markdownParser.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/services/tree-sitter/markdownParser.ts), where headers are parsed into separate blocks typed as `markdown_header_hN`. This allows documentation sections to become first-class searchable entities alongside actual code.

Once parsed, the resulting `CodeBlock[]` objects are fed to the configured **embedding provider** (OpenAI, Ollama, Gemini, etc.) to produce high-dimensional vectors stored in the vector database.

### Stage 2: Vector-Based Semantic Retrieval

The **`codebase_search`** tool serves as the public entry-point for every search operation. Defined in [`src/core/prompts/tools/native-tools/codebase_search.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/core/prompts/tools/native-tools/codebase_search.ts), this tool receives natural-language queries and forwards them to the embedding service.

Using the `EmbeddingModelProfile` configuration from [`packages/types/src/embedding.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/packages/types/src/embedding.ts), the system embeds the query and retrieves the nearest code-block vectors from the store. Because the index stores structured blocks complete with identifiers, types, and line ranges, returned hits are already scoped to precise locations in the source tree, allowing the UI to display results as "function X in [`src/foo/bar.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/foo/bar.ts)".

## From Source Code to Vector Embeddings

The architectural walk-through reveals seven distinct phases that transform raw files into queryable vectors:

1. **File discovery** – The indexer walks the workspace respecting `.rooignore` and processes every file whose extension appears in `scannerExtensions`.

2. **Language-specific parsing** – `loadRequiredLanguageParsers` dynamically loads Tree-sitter grammars to capture named AST nodes for each supported language.

3. **Intelligent chunking** – Large captures trigger fallback chunkers that split by line length while preserving boundaries and avoiding tiny remainders.

4. **Deterministic hashing** – Each chunk receives a unique segment hash to prevent embedding duplication and keep the vector store lean.

5. **Markdown extraction** – Headers become separate searchable blocks, enabling documentation queries at the same granularity as code.

6. **Vector generation** – `CodeBlock` objects are passed to the embedding provider specified in the user's configuration.

7. **Semantic lookup** – The `codebase_search` tool embeds natural language queries and retrieves the nearest vectors, returning metadata-rich results to the LLM for answer generation.

## Provider-Agnostic Configuration

The embedding layer abstracts provider-specific details through the `EmbeddingModelProfile` interface. This configuration, defined in [`packages/types/src/embedding.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/packages/types/src/embedding.ts), supplies the vector dimension and optional `scoreThreshold` for each supported model:

```typescript
import type { EmbeddingModelProfiles } from "./packages/types/src/embedding";

export const embeddingProfiles: EmbeddingModelProfiles = {
  openai: {
    "text-embedding-3-large": { dimension: 3072, scoreThreshold: 0.78 },
  },
  ollama: {
    "mistral": { dimension: 4096 },
  },
};

```

This abstraction allows Roo Code to switch between OpenAI, Gemini, Ollama, and other providers without modifying the core indexing logic. Each profile defines the expected vector dimension and confidence thresholds used to filter low-similarity matches.

## Working with the Indexing System

### Indexing a Single File

To programmatically parse a file into searchable blocks, use the `codeParser` singleton:

```typescript
import { codeParser } from "./src/services/code-index/processors/parser";

async function indexFile(filePath: string) {
  const blocks = await codeParser.parseFile(filePath);
  // `blocks` is an array of CodeBlock objects ready for embedding
  console.log(`Found ${blocks.length} indexable blocks in ${filePath}`);
}

```

### Performing a Semantic Search

The `codebase_search` tool can be invoked directly when building custom agentic workflows:

```typescript
// This is the shape the LLM calls internally; you can invoke it directly:
await openAi.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "How does Roo Code handle user authentication?" }],
  tools: [{ /* codebase_search tool definition */ }],
  tool_choice: { type: "function", function: { name: "codebase_search" } },
});

```

This tool leverages the semantic index to find meaning-related code even when exact keyword matches are absent.

## Summary

- **Structural awareness** via Tree-sitter parsing allows Roo Code to distinguish between functions, classes, and documentation headers, outperforming naive text tokenization.
- **Deterministic segment hashing** prevents embedding duplication and maintains a lean vector store across re-indexing operations.
- **Adaptive chunking** strategies ensure each vector represents a meaningful code unit, splitting large files while respecting line boundaries.
- **Provider-agnostic embeddings** enable flexibility across OpenAI, Ollama, Gemini, and other services through configurable `EmbeddingModelProfile` definitions.
- The **`codebase_search`** tool exposes semantic lookup capabilities directly to the LLM, enabling natural language queries against the indexed repository.

## Frequently Asked Questions

### How does Roo Code handle programming languages without Tree-sitter support?

When a language lacks a Tree-sitter query or returns empty captures, Roo Code falls back to line-based chunking via `_chunkLeafNodeByLines` or `_chunkTextByLines`. These methods split files by character limits while respecting line boundaries, ensuring every file type remains searchable even without AST parsing.

### What prevents duplicate embeddings when files are re-indexed?

Each code block receives a deterministic **segment hash** calculated as `sha256(filePath-start-end-len-preview)`. This hash, combined with a file-level hash, allows the system to detect existing entries during re-indexing, preventing duplicate vectors in the database and maintaining storage efficiency.

### Why does Roo Code parse Markdown headers separately?

Markdown headers are extracted in [`src/services/tree-sitter/markdownParser.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/src/services/tree-sitter/markdownParser.ts) and treated as distinct block types (`markdown_header_hN`). This architectural choice makes documentation sections searchable at the same granularity as code symbols, allowing the semantic search to match queries against README sections, API documentation, and inline comments.

### How is the semantic similarity threshold configured?

The `EmbeddingModelProfile` interface in [`packages/types/src/embedding.ts`](https://github.com/RooCodeInc/Roo-Code/blob/main/packages/types/src/embedding.ts) defines an optional `scoreThreshold` property for each model configuration. This threshold filters low-confidence vector matches before returning results to the `codebase_search` tool, with specific values like `0.78` tuned for models such as OpenAI's `text-embedding-3-large`.