Optimal Embedding Batch Size in Claude Context: How It Affects Indexing Performance
The optimal embedding batch size in Claude Context is 100 by default, configurable via EMBEDDING_BATCH_SIZE, with larger values (up to 512) reducing indexing time by minimizing API calls while balancing memory usage and rate limits.
Claude Context creates vector embeddings for code chunks in batches to optimize performance when indexing repositories. The batch size directly impacts how quickly your codebase can be processed, how much memory is consumed, and whether you hit embedding provider rate limits. Understanding how to tune this parameter is essential for production deployments of the zilliztech/claude-context system.
How Embedding Batch Size Is Configured
The batch size is controlled entirely through the EMBEDDING_BATCH_SIZE environment variable. The system reads this value at runtime and applies sensible bounds to prevent misconfiguration.
Default Value and Bounds
In packages/core/src/context.ts, the batch size is parsed and clamped:
const EMBEDDING_BATCH_SIZE = Math.max(
1,
parseInt(envManager.get('EMBEDDING_BATCH_SIZE') || '100', 10)
);
The implementation enforces a minimum of 1 to prevent empty batches, with a default of 100 when the variable is unset. The MCP documentation notes that you can raise this value—for example, to 512—to "optimize the performance of the MCP server" (source: packages/mcp/README.md line 164).
Setting the Environment Variable
Configure the batch size in your .env file:
# .env
EMBEDDING_BATCH_SIZE=256 # Increase for faster indexing if your model & resources permit
Or pass it directly when running:
EMBEDDING_BATCH_SIZE=512 npm run start
How Batch Size Influences Indexing Performance
The batch size creates a fundamental trade-off between throughput, latency, and resource consumption. Larger batches reduce API call overhead but increase memory pressure and risk of rate limit violations.
Performance Characteristics by Batch Size Range
| Batch Size | Effect on Indexing |
|---|---|
| Small (1–50) | More API calls → higher total latency. Minimal memory usage. Respects strict rate limits. Best for constrained environments or very large individual chunks. |
| Medium (100–200) | Balanced throughput. The default 100 provides good performance without overwhelming most embedding services. Moderate memory consumption. |
| Large (300–512) | Fewer API calls → significantly shorter overall indexing time. Parallel processing within each batch maximizes throughput. However, requires more RAM and may hit provider token/size limits, causing 429 errors or timeouts. |
The Batching Mechanism in Practice
The actual batching logic resides in Context.processFileBatch (packages/core/src/context.ts lines 733–736). The implementation accumulates chunks until the buffer reaches EMBEDDING_BATCH_SIZE, then flushes the entire batch:
// Accumulate chunks
chunkBuffer.push({ chunk, codebasePath });
totalChunks++;
// When the buffer reaches the configured size, send the batch
if (chunkBuffer.length >= EMBEDDING_BATCH_SIZE) {
await this.processChunkBuffer(chunkBuffer);
chunkBuffer = []; // clear buffer for next batch
}
After all files are processed, any remaining chunks are flushed as a final batch (lines 68–73). This ensures no data is lost regardless of whether the total chunk count divides evenly by the batch size.
Determining Your Optimal Batch Size
The truly optimal value depends on three interacting factors:
-
Embedding model throughput — High-throughput models like OpenAI
text-embedding-3-largecan sustain larger batches without latency degradation. -
Available memory — Each chunk's text is held in memory until batch flush. Large batches with large code files can exhaust RAM.
-
Provider rate limits — Exceeding per-request token caps or request rates triggers 429 errors. The SDK retries automatically, but this increases total indexing time.
Practical Tuning Strategy
Start with the default 100, then iterate:
# Test progression
EMBEDDING_BATCH_SIZE=200 npm run index # 2x baseline
EMBEDDING_BATCH_SIZE=512 npm run index # Maximum recommended
Monitor for:
- Memory usage — Watch for OOM kills or swapping
- Error logs — "rate limit exceeded" or timeout messages
- Total indexing time — Measure wall-clock improvement
If you encounter rate limits or memory pressure, scale back by 25–50%.
Summary
- Embedding batch size in Claude Context is controlled by
EMBEDDING_BATCH_SIZE, defaulting to 100 and clamped to minimum 1. - The implementation in
packages/core/src/context.tsbuffers chunks until the batch size is reached, then sends them to the embedding provider viaprocessChunkBuffer. - Larger batches reduce API call overhead and total indexing time, but increase memory consumption and risk of hitting provider rate limits.
- Optimal values typically range from 100–512, with 512 being the upper bound recommended in the MCP documentation. Tune based on your embedding model throughput, available RAM, and provider constraints.
Frequently Asked Questions
What happens if I set EMBEDDING_BATCH_SIZE to 1?
Setting the batch size to 1 forces Claude Context to make an API call for every single code chunk. This minimizes memory usage and never hits per-request size limits, but maximizes total indexing time due to API call overhead and network latency. Use this only in extremely memory-constrained environments or when processing individual chunks that exceed provider size limits.
How do I know if my batch size is too large?
Signs of an oversized batch include "rate limit exceeded" (429) errors, request timeout messages, or process termination due to out-of-memory (OOM) conditions. The Claude Context SDK automatically retries failed requests, but you will observe increased indexing time and error messages in logs. If you encounter these symptoms, reduce your batch size by 25–50% and retest.
Does the optimal batch size vary by embedding model?
Yes. High-throughput models like OpenAI's text-embedding-3-large or text-embedding-3-small can efficiently process larger batches without significant latency degradation. Models with lower throughput or stricter rate limits benefit from smaller batches to maintain steady request flow. Consult your provider's documentation for throughput characteristics and recommended request sizes, then tune EMBEDDING_BATCH_SIZE accordingly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →