What Code Splitters Are Available in Claude Context and When to Use Each One

Claude Context provides two interchangeable code splitters: the AST-based AstCodeSplitter (default, 2500 bytes) for semantic boundaries and the LangChainCodeSplitter (1000 bytes) for language-agnostic fallback.

Claude Context, an open-source semantic code search tool from Zilliz, indexes your codebase by splitting source files into searchable chunks. The choice of code splitter directly impacts search accuracy and recall. According to the zilliztech/claude-context source code, you have two production-ready implementations that share a common Splitter interface defined in packages/core/src/splitter/index.ts.

The Two Available Code Splitters

Claude Context ships with AST and LangChain splitters, selectable via the SplitterType enum (AST or LANGCHAIN). Both implement the same interface: split(code, language, filePath), setChunkSize(), and setChunkOverlap().

AST Splitter: Semantic-Aware Chunking

The AstCodeSplitter class in packages/core/src/splitter/ast-splitter.ts parses source code into an Abstract Syntax Tree using tree-sitter, then extracts logical units like functions, classes, and methods as discrete chunks.

Default configuration:

  • Chunk size: 2500 bytes
  • Chunk overlap: 300 bytes

Key implementation details from ast-splitter.ts:

  1. Language support map (lines 86-105): Maps language identifiers to tree-sitter parsers and SPLITTABLE_NODE_TYPES (functions, classes, etc.)
  2. Fallback logic (lines 44-51): If getLanguageConfig returns null, automatically delegates to LangChainCodeSplitter
  3. Chunk extraction (lines 109-138): Walks the AST, creating chunks for each splittable node type
  4. Refinement (lines 64-77): Oversized chunks are further split using character-based logic

When to use the AST splitter:

  • Supported languages with robust tree-sitter grammars: JavaScript, TypeScript, Python, Java, C++, Go, Rust, C#, Scala
  • When semantic boundaries matter: retrieving whole functions or class definitions rather than arbitrary text windows
  • Performance-critical indexing: parses once, avoids redundant character-splitting
  • When you want automatic fallback: built-in delegation to LangChain when parsing fails

LangChain Splitter: Flexible Language-Agnostic Chunking

The LangChainCodeSplitter class in packages/core/src/splitter/langchain-splitter.ts uses LangChain's RecursiveCharacterTextSplitter with language-specific optimizations and generic fallback.

Default configuration:

  • Chunk size: 1000 bytes
  • Chunk overlap: 200 bytes

Key implementation details from langchain-splitter.ts:

  1. Language mapping (lines 64-89): Translates language strings to LangChain-supported identifiers (e.g., 'javascript' → 'js')
  2. Primary splitting (lines 16-48): Uses RecursiveCharacterTextSplitter.fromLanguage() when mapped; otherwise generic RecursiveCharacterTextSplitter
  3. Line-number estimation (lines 15-31): When using the generic splitter, estimates start/end lines by locating the chunk in original source

When to use the LangChain splitter:

  • Unsupported or niche languages: DSLs, proprietary formats, or languages without tree-sitter grammars
  • Smaller chunk requirements: config files, microservices, or when higher recall on short snippets is needed
  • Language-agnostic preference: simpler mental model, no AST parsing overhead
  • Custom overlap tuning: when you need precise control over chunk boundaries for specific retrieval patterns

Selecting and Configuring Your Splitter

Runtime Selection via Environment Variable

As demonstrated in examples/basic-usage/index.ts (lines 23-49), choose the splitter at runtime:

import { AstCodeSplitter, LangChainCodeSplitter } from '@zilliz/claude-context-core';

const splitterType = process.env.SPLITTER_TYPE?.toLowerCase() || 'ast';

const splitter = splitterType === 'langchain'
  ? new LangChainCodeSplitter(1000, 200)
  : new AstCodeSplitter(2500, 300);

Pass this to your Context instance:

import { Context, MilvusVectorDatabase } from '@zilliz/claude-context-core';

const vectorDb = new MilvusVectorDatabase({ address: 'localhost:19530' });
const ctx = new Context({ 
  vectorDatabase: vectorDb, 
  codeSplitter: splitter 
});

await ctx.indexCodebase('/path/to/project');

VS Code Extension Configuration

The VS Code extension persists splitter settings via configManager.ts (lines 345-381):

  1. Open Settings → Extensions → Claude Context → Splitter
  2. Choose AST or LangChain
  3. Adjust Chunk Size and Chunk Overlap as needed

The extension reads these values via getSplitterConfig() and instantiates the appropriate splitter in extension.ts (lines 155-170).

Quick Reference: Choosing Your Splitter

Situation Recommended Splitter Rationale
JavaScript, TypeScript, Python, Java, C++, Go, Rust, C#, Scala AST splitter Native tree-sitter support, semantic boundaries
Need whole functions/classes in results AST splitter AST extraction preserves logical units
Performance-critical indexing AST splitter Single parse, minimal overhead
Automatic fallback on parse failure AST splitter Built-in delegation to LangChain
Niche languages, DSLs, or proprietary formats LangChain splitter Language-agnostic, no grammar required
Config files, microservices, short snippets LangChain splitter Smaller default chunk size (1000 bytes)
Precise overlap control for retrieval tuning LangChain splitter Easier to customize chunk boundaries

Summary

  • Two code splitters ship with Claude Context: AstCodeSplitter and LangChainCodeSplitter, both implementing the common Splitter interface from packages/core/src/splitter/index.ts

  • AST splitter (default: 2500 bytes, 300 overlap) uses tree-sitter for semantic, language-aware chunking with automatic fallback—ideal for supported languages requiring whole-function or class-level retrieval

  • LangChain splitter (default: 1000 bytes, 200 overlap) provides language-agnostic, character-based splitting with smaller chunks—optimal for niche languages, config files, or fine-grained recall tuning

  • Selection methods include environment variable (SPLITTER_TYPE), direct instantiation in code, or VS Code extension settings managed via configManager.ts

Frequently Asked Questions

What happens if the AST splitter encounters an unsupported language?

The AST splitter automatically falls back to the LangChain splitter. In packages/core/src/splitter/ast-splitter.ts (lines 44-51), if getLanguageConfig returns null for the given language, the code delegates splitting to a LangChainCodeSplitter instance, ensuring indexing continues without error.

Can I use different splitters for different files in the same project?

Currently, Claude Context uses a single splitter instance per Context. As shown in examples/basic-usage/index.ts (lines 44-49), you select one splitter at initialization time. To use different strategies for different file types, you would need to create separate Context instances or implement a custom Splitter that delegates based on file extension.

How do chunk size and overlap affect search quality?

Chunk size determines the maximum amount of code per searchable unit. Larger sizes (AST default: 2500 bytes) preserve more context and reduce embedding fragmentation, improving coherence for semantic search. Smaller sizes (LangChain default: 1000 bytes) increase the number of chunks, potentially improving recall for specific short queries but risking context loss.

Chunk overlap (AST: 300 bytes, LangChain: 200 bytes) ensures adjacent chunks share content, preventing relevant code from being split across chunk boundaries. Increase overlap when you notice search missing code that spans logical boundaries; decrease it to reduce storage overhead.

Does the AST splitter support all tree-sitter languages?

No. The AST splitter's language support is defined in packages/core/src/splitter/ast-splitter.ts (lines 86-105), which maps specific language identifiers to tree-sitter parsers and SPLITTABLE_NODE_TYPES. Supported languages include JavaScript, TypeScript, Python, Java, C++, Go, Rust, C#, and Scala. Languages outside this set trigger the fallback to LangChain splitting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →