# What Code Splitters Are Available in Claude Context and When to Use Each One

> Explore Claude Context code splitters. Learn when to use the AST based splitter for semantic boundaries or the LangChain splitter for fallback.

- Repository: [Zilliz/claude-context](https://github.com/zilliztech/claude-context)
- Tags: deep-dive
- Published: 2026-04-22

---

**Claude Context provides two interchangeable code splitters: the AST-based `AstCodeSplitter` (default, 2500 bytes) for semantic boundaries and the `LangChainCodeSplitter` (1000 bytes) for language-agnostic fallback.**

Claude Context, an open-source semantic code search tool from Zilliz, indexes your codebase by splitting source files into searchable chunks. The choice of **code splitter** directly impacts search accuracy and recall. According to the `zilliztech/claude-context` source code, you have two production-ready implementations that share a common `Splitter` interface defined in [`packages/core/src/splitter/index.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/splitter/index.ts).

## The Two Available Code Splitters

Claude Context ships with **AST** and **LangChain** splitters, selectable via the `SplitterType` enum (`AST` or `LANGCHAIN`). Both implement the same interface: `split(code, language, filePath)`, `setChunkSize()`, and `setChunkOverlap()`.

### AST Splitter: Semantic-Aware Chunking

The **`AstCodeSplitter`** class in [`packages/core/src/splitter/ast-splitter.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/splitter/ast-splitter.ts) parses source code into an Abstract Syntax Tree using **tree-sitter**, then extracts logical units like functions, classes, and methods as discrete chunks.

**Default configuration:**
- Chunk size: **2500 bytes**
- Chunk overlap: **300 bytes**

**Key implementation details from [`ast-splitter.ts`](https://github.com/zilliztech/claude-context/blob/main/ast-splitter.ts):**

1. **Language support map** (lines 86-105): Maps language identifiers to tree-sitter parsers and `SPLITTABLE_NODE_TYPES` (functions, classes, etc.)
2. **Fallback logic** (lines 44-51): If `getLanguageConfig` returns `null`, automatically delegates to `LangChainCodeSplitter`
3. **Chunk extraction** (lines 109-138): Walks the AST, creating chunks for each splittable node type
4. **Refinement** (lines 64-77): Oversized chunks are further split using character-based logic

**When to use the AST splitter:**

- **Supported languages with robust tree-sitter grammars**: JavaScript, TypeScript, Python, Java, C++, Go, Rust, C#, Scala
- **When semantic boundaries matter**: retrieving whole functions or class definitions rather than arbitrary text windows
- **Performance-critical indexing**: parses once, avoids redundant character-splitting
- **When you want automatic fallback**: built-in delegation to LangChain when parsing fails

### LangChain Splitter: Flexible Language-Agnostic Chunking

The **`LangChainCodeSplitter`** class in [`packages/core/src/splitter/langchain-splitter.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/splitter/langchain-splitter.ts) uses LangChain's `RecursiveCharacterTextSplitter` with language-specific optimizations and generic fallback.

**Default configuration:**
- Chunk size: **1000 bytes**
- Chunk overlap: **200 bytes**

**Key implementation details from [`langchain-splitter.ts`](https://github.com/zilliztech/claude-context/blob/main/langchain-splitter.ts):**

1. **Language mapping** (lines 64-89): Translates language strings to LangChain-supported identifiers (e.g., `'javascript' → 'js'`)
2. **Primary splitting** (lines 16-48): Uses `RecursiveCharacterTextSplitter.fromLanguage()` when mapped; otherwise generic `RecursiveCharacterTextSplitter`
3. **Line-number estimation** (lines 15-31): When using the generic splitter, estimates start/end lines by locating the chunk in original source

**When to use the LangChain splitter:**

- **Unsupported or niche languages**: DSLs, proprietary formats, or languages without tree-sitter grammars
- **Smaller chunk requirements**: config files, microservices, or when higher recall on short snippets is needed
- **Language-agnostic preference**: simpler mental model, no AST parsing overhead
- **Custom overlap tuning**: when you need precise control over chunk boundaries for specific retrieval patterns

## Selecting and Configuring Your Splitter

### Runtime Selection via Environment Variable

As demonstrated in [`examples/basic-usage/index.ts`](https://github.com/zilliztech/claude-context/blob/main/examples/basic-usage/index.ts) (lines 23-49), choose the splitter at runtime:

```typescript
import { AstCodeSplitter, LangChainCodeSplitter } from '@zilliz/claude-context-core';

const splitterType = process.env.SPLITTER_TYPE?.toLowerCase() || 'ast';

const splitter = splitterType === 'langchain'
  ? new LangChainCodeSplitter(1000, 200)
  : new AstCodeSplitter(2500, 300);

```

Pass this to your `Context` instance:

```typescript
import { Context, MilvusVectorDatabase } from '@zilliz/claude-context-core';

const vectorDb = new MilvusVectorDatabase({ address: 'localhost:19530' });
const ctx = new Context({ 
  vectorDatabase: vectorDb, 
  codeSplitter: splitter 
});

await ctx.indexCodebase('/path/to/project');

```

### VS Code Extension Configuration

The VS Code extension persists splitter settings via [`configManager.ts`](https://github.com/zilliztech/claude-context/blob/main/configManager.ts) (lines 345-381):

1. Open **Settings → Extensions → Claude Context → Splitter**
2. Choose **AST** or **LangChain**
3. Adjust **Chunk Size** and **Chunk Overlap** as needed

The extension reads these values via `getSplitterConfig()` and instantiates the appropriate splitter in [`extension.ts`](https://github.com/zilliztech/claude-context/blob/main/extension.ts) (lines 155-170).

## Quick Reference: Choosing Your Splitter

| Situation | Recommended Splitter | Rationale |
|-----------|---------------------|-----------|
| JavaScript, TypeScript, Python, Java, C++, Go, Rust, C#, Scala | **AST splitter** | Native tree-sitter support, semantic boundaries |
| Need whole functions/classes in results | **AST splitter** | AST extraction preserves logical units |
| Performance-critical indexing | **AST splitter** | Single parse, minimal overhead |
| Automatic fallback on parse failure | **AST splitter** | Built-in delegation to LangChain |
| Niche languages, DSLs, or proprietary formats | **LangChain splitter** | Language-agnostic, no grammar required |
| Config files, microservices, short snippets | **LangChain splitter** | Smaller default chunk size (1000 bytes) |
| Precise overlap control for retrieval tuning | **LangChain splitter** | Easier to customize chunk boundaries |

## Summary

- **Two code splitters** ship with Claude Context: `AstCodeSplitter` and `LangChainCodeSplitter`, both implementing the common `Splitter` interface from [`packages/core/src/splitter/index.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/splitter/index.ts)

- **AST splitter** (default: 2500 bytes, 300 overlap) uses tree-sitter for semantic, language-aware chunking with automatic fallback—ideal for supported languages requiring whole-function or class-level retrieval

- **LangChain splitter** (default: 1000 bytes, 200 overlap) provides language-agnostic, character-based splitting with smaller chunks—optimal for niche languages, config files, or fine-grained recall tuning

- **Selection methods** include environment variable (`SPLITTER_TYPE`), direct instantiation in code, or VS Code extension settings managed via [`configManager.ts`](https://github.com/zilliztech/claude-context/blob/main/configManager.ts)

## Frequently Asked Questions

### What happens if the AST splitter encounters an unsupported language?

The AST splitter automatically falls back to the LangChain splitter. In [`packages/core/src/splitter/ast-splitter.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/splitter/ast-splitter.ts) (lines 44-51), if `getLanguageConfig` returns `null` for the given language, the code delegates splitting to a `LangChainCodeSplitter` instance, ensuring indexing continues without error.

### Can I use different splitters for different files in the same project?

Currently, Claude Context uses a single splitter instance per `Context`. As shown in [`examples/basic-usage/index.ts`](https://github.com/zilliztech/claude-context/blob/main/examples/basic-usage/index.ts) (lines 44-49), you select one splitter at initialization time. To use different strategies for different file types, you would need to create separate `Context` instances or implement a custom `Splitter` that delegates based on file extension.

### How do chunk size and overlap affect search quality?

**Chunk size** determines the maximum amount of code per searchable unit. Larger sizes (AST default: 2500 bytes) preserve more context and reduce embedding fragmentation, improving coherence for semantic search. Smaller sizes (LangChain default: 1000 bytes) increase the number of chunks, potentially improving recall for specific short queries but risking context loss.

**Chunk overlap** (AST: 300 bytes, LangChain: 200 bytes) ensures adjacent chunks share content, preventing relevant code from being split across chunk boundaries. Increase overlap when you notice search missing code that spans logical boundaries; decrease it to reduce storage overhead.

### Does the AST splitter support all tree-sitter languages?

No. The AST splitter's language support is defined in [`packages/core/src/splitter/ast-splitter.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/splitter/ast-splitter.ts) (lines 86-105), which maps specific language identifiers to tree-sitter parsers and `SPLITTABLE_NODE_TYPES`. Supported languages include JavaScript, TypeScript, Python, Java, C++, Go, Rust, C#, and Scala. Languages outside this set trigger the fallback to LangChain splitting.