How Claude Context Manages Collection Naming for Multiple Indexed Codebases
Claude Context generates deterministic, unique collection names by hashing the absolute path of each codebase, ensuring no collisions between projects while maintaining stable identifiers across re-indexing operations.
When working with multiple codebases in Claude Context, the system needs a reliable way to isolate vector embeddings in separate collections. The solution implemented in the zilliztech/claude-context repository uses path-based hashing to create unique, deterministic collection names that prevent data mixing between projects.
The Collection Naming Algorithm
The core logic resides in Context.getCollectionName() at packages/core/src/context.ts (lines 234-240). The method follows a three-step process to generate each collection name.
Step 1: Determine the Prefix Based on Hybrid Mode
// From context.ts lines 234-240
const hybridMode = process.env.HYBRID_MODE !== 'false'; // defaults to true
const prefix = hybridMode ? 'hybrid_code_chunks' : 'code_chunks';
The HYBRID_MODE environment variable controls whether the collection stores dense vectors only or a hybrid of dense plus sparse vectors. This prefix makes the search mode immediately visible in the collection name.
Step 2: Normalize and Hash the Codebase Path
// Path normalization and hashing
const resolvedPath = path.resolve(codebasePath);
const hash = crypto.createHash('md5').update(resolvedPath).digest('hex');
const shortHash = hash.substring(0, 8);
The path.resolve() call converts any relative path to an absolute, platform-independent form. This ensures that ./my-project and /home/user/my-project produce the same hash when resolved from the same working directory.
MD5 hashing provides a deterministic 32-character hex string, from which the first 8 characters are extracted. This produces 16^8 (over 4 billion) possible combinations—sufficient for collision resistance in typical usage while keeping names readable.
Step 3: Assemble the Final Collection Name
// Final name construction
return `${prefix}_${shortHash}`;
A complete collection name looks like hybrid_code_chunks_a1b2c3d4.
Why This Approach Prevents Collection Collisions
The path-based hashing strategy guarantees three critical properties for multi-codebase management:
-
Uniqueness across different projects: Two codebases at
/home/user/project-alphaand/home/user/project-betaproduce entirely different hashes, even though they share the same parent directory. -
Stability for the same project: Re-indexing
/home/user/project-alphaa week later generates the identical collection name, allowing the system to reuse existing vectors or perform incremental updates. -
Visibility of search configuration: The
hybrid_prefix immediately indicates whether sparse vectors are available for keyword-dense searches.
How Collection Names Are Used Throughout the System
The getCollectionName() method is the single source of truth for collection identification across all Claude Context operations.
During Indexing
In packages/vscode-extension/src/commands/indexCommand.ts (lines 84-92), the collection name is obtained before creating the synchronizer and invoking the index tool:
// From indexCommand.ts
const collectionName = context.getCollectionName(codebasePath);
await vectorDb.prepareCollection(collectionName, {
hybridMode: context.hybridMode
});
During Search and Status Operations
The name is retrieved consistently in packages/core/src/context.ts for:
semanticSearch()— lines 14-15hasIndex()— lines 23-25clearIndex()— various internal calls
During MCP Snapshot Handling
In packages/mcp/src/handlers.ts (lines 31-33, 73-75, 539-540, 719-720), the collection name validates snapshot consistency and manages cleanup operations.
Practical Example: Indexing Multiple Projects
import { Context } from '@zilliztech/claude-context';
import { MilvusVectorDatabase } from '@zilliztech/claude-context/vector-db';
const vectorDb = new MilvusVectorDatabase({
address: 'localhost:19530'
});
const ctx = new Context({ vectorDatabase: vectorDb });
async function indexProject(path: string) {
const collection = ctx.getCollectionName(path);
console.log(`Indexing ${path} → collection "${collection}"`);
await ctx.indexCodebase(path);
}
// Index two distinct repositories
await indexProject('/home/user/project-alpha');
await indexProject('/home/user/project-beta');
// Search within a specific project
const results = await ctx.semanticSearch(
'/home/user/project-alpha',
'how to handle file uploads'
);
The output demonstrates the deterministic naming:
Indexing /home/user/project-alpha → collection "hybrid_code_chunks_3f8a9c2e"
Indexing /home/user/project-beta → collection "hybrid_code_chunks_7b4d1f5a"
Each project receives a unique collection, enabling isolated vector storage and targeted semantic search.
Summary
-
Collection names are deterministic hashes of the absolute codebase path, preventing collisions between different projects.
-
The
getCollectionName()method inpackages/core/src/context.ts(lines 234-240) implements the hashing logic with MD5 and an 8-character truncation. -
Hybrid mode configuration is encoded in the collection prefix, making the search strategy visible in the name itself.
-
All system components retrieve collection names through the same method, ensuring consistency across indexing, search, status checks, and MCP operations.
Frequently Asked Questions
What happens if two projects have the same folder name but different parent directories?
The collection names will still be unique. The path.resolve() call produces the full absolute path before hashing, so /home/user/project and /opt/data/project generate completely different MD5 hashes and thus different collection names.
Can I change the hybrid mode for an existing collection without creating a new one?
No—changing HYBRID_MODE after initial indexing will generate a different collection name prefix, causing the system to create and use a separate collection. To re-index with a different mode, you must explicitly clear the existing collection or use the force-recreate option.
How long can the collection name become with very long paths?
The collection name is always bounded at 24 characters: the prefix (hybrid_code_chunks = 18 characters or code_chunks = 11 characters) plus underscore plus exactly 8 hex characters. The MD5 hash truncation ensures consistent length regardless of input path length.
Is MD5 secure enough for this use case?
Yes—MD5 is used here solely for deterministic name generation, not for cryptographic security. The 8-character truncation provides sufficient collision resistance for typical development workloads, and the absolute path inputs are not attacker-controlled in standard Claude Context usage scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →