# How Claude Context Manages Collection Naming for Multiple Indexed Codebases

> Discover how Claude Context ensures unique collection names when indexing multiple codebases. Learn about deterministic naming and collision prevention for stable identifiers.

- Repository: [Zilliz/claude-context](https://github.com/zilliztech/claude-context)
- Tags: how-to-guide
- Published: 2026-04-22

---

**Claude Context generates deterministic, unique collection names by hashing the absolute path of each codebase, ensuring no collisions between projects while maintaining stable identifiers across re-indexing operations.**

When working with multiple codebases in Claude Context, the system needs a reliable way to isolate vector embeddings in separate collections. The solution implemented in the `zilliztech/claude-context` repository uses path-based hashing to create unique, deterministic collection names that prevent data mixing between projects.

## The Collection Naming Algorithm

The core logic resides in `Context.getCollectionName()` at [`packages/core/src/context.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/context.ts) (lines 234-240). The method follows a three-step process to generate each collection name.

### Step 1: Determine the Prefix Based on Hybrid Mode

```typescript
// From context.ts lines 234-240
const hybridMode = process.env.HYBRID_MODE !== 'false'; // defaults to true
const prefix = hybridMode ? 'hybrid_code_chunks' : 'code_chunks';

```

The `HYBRID_MODE` environment variable controls whether the collection stores dense vectors only or a hybrid of dense plus sparse vectors. This prefix makes the search mode immediately visible in the collection name.

### Step 2: Normalize and Hash the Codebase Path

```typescript
// Path normalization and hashing
const resolvedPath = path.resolve(codebasePath);
const hash = crypto.createHash('md5').update(resolvedPath).digest('hex');
const shortHash = hash.substring(0, 8);

```

The `path.resolve()` call converts any relative path to an absolute, platform-independent form. This ensures that `./my-project` and `/home/user/my-project` produce the same hash when resolved from the same working directory.

MD5 hashing provides a deterministic 32-character hex string, from which the first 8 characters are extracted. This produces 16^8 (over 4 billion) possible combinations—sufficient for collision resistance in typical usage while keeping names readable.

### Step 3: Assemble the Final Collection Name

```typescript
// Final name construction
return `${prefix}_${shortHash}`;

```

A complete collection name looks like `hybrid_code_chunks_a1b2c3d4`.

## Why This Approach Prevents Collection Collisions

The path-based hashing strategy guarantees three critical properties for multi-codebase management:

- **Uniqueness across different projects**: Two codebases at `/home/user/project-alpha` and `/home/user/project-beta` produce entirely different hashes, even though they share the same parent directory.

- **Stability for the same project**: Re-indexing `/home/user/project-alpha` a week later generates the identical collection name, allowing the system to reuse existing vectors or perform incremental updates.

- **Visibility of search configuration**: The `hybrid_` prefix immediately indicates whether sparse vectors are available for keyword-dense searches.

## How Collection Names Are Used Throughout the System

The `getCollectionName()` method is the single source of truth for collection identification across all Claude Context operations.

### During Indexing

In [`packages/vscode-extension/src/commands/indexCommand.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/vscode-extension/src/commands/indexCommand.ts) (lines 84-92), the collection name is obtained before creating the synchronizer and invoking the index tool:

```typescript
// From indexCommand.ts
const collectionName = context.getCollectionName(codebasePath);
await vectorDb.prepareCollection(collectionName, {
  hybridMode: context.hybridMode
});

```

### During Search and Status Operations

The name is retrieved consistently in [`packages/core/src/context.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/context.ts) for:

- `semanticSearch()` — lines 14-15
- `hasIndex()` — lines 23-25
- `clearIndex()` — various internal calls

### During MCP Snapshot Handling

In [`packages/mcp/src/handlers.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/mcp/src/handlers.ts) (lines 31-33, 73-75, 539-540, 719-720), the collection name validates snapshot consistency and manages cleanup operations.

## Practical Example: Indexing Multiple Projects

```typescript
import { Context } from '@zilliztech/claude-context';
import { MilvusVectorDatabase } from '@zilliztech/claude-context/vector-db';

const vectorDb = new MilvusVectorDatabase({
  address: 'localhost:19530'
});

const ctx = new Context({ vectorDatabase: vectorDb });

async function indexProject(path: string) {
  const collection = ctx.getCollectionName(path);
  console.log(`Indexing ${path} → collection "${collection}"`);
  await ctx.indexCodebase(path);
}

// Index two distinct repositories
await indexProject('/home/user/project-alpha');
await indexProject('/home/user/project-beta');

// Search within a specific project
const results = await ctx.semanticSearch(
  '/home/user/project-alpha',
  'how to handle file uploads'
);

```

The output demonstrates the deterministic naming:

```

Indexing /home/user/project-alpha → collection "hybrid_code_chunks_3f8a9c2e"
Indexing /home/user/project-beta → collection "hybrid_code_chunks_7b4d1f5a"

```

Each project receives a unique collection, enabling isolated vector storage and targeted semantic search.

## Summary

- **Collection names are deterministic hashes** of the absolute codebase path, preventing collisions between different projects.

- **The `getCollectionName()` method** in [`packages/core/src/context.ts`](https://github.com/zilliztech/claude-context/blob/main/packages/core/src/context.ts) (lines 234-240) implements the hashing logic with MD5 and an 8-character truncation.

- **Hybrid mode configuration** is encoded in the collection prefix, making the search strategy visible in the name itself.

- **All system components** retrieve collection names through the same method, ensuring consistency across indexing, search, status checks, and MCP operations.

## Frequently Asked Questions

### What happens if two projects have the same folder name but different parent directories?

The collection names will still be unique. The `path.resolve()` call produces the full absolute path before hashing, so `/home/user/project` and `/opt/data/project` generate completely different MD5 hashes and thus different collection names.

### Can I change the hybrid mode for an existing collection without creating a new one?

No—changing `HYBRID_MODE` after initial indexing will generate a different collection name prefix, causing the system to create and use a separate collection. To re-index with a different mode, you must explicitly clear the existing collection or use the force-recreate option.

### How long can the collection name become with very long paths?

The collection name is always bounded at 24 characters: the prefix (`hybrid_code_chunks` = 18 characters or `code_chunks` = 11 characters) plus underscore plus exactly 8 hex characters. The MD5 hash truncation ensures consistent length regardless of input path length.

### Is MD5 secure enough for this use case?

Yes—MD5 is used here solely for deterministic name generation, not for cryptographic security. The 8-character truncation provides sufficient collision resistance for typical development workloads, and the absolute path inputs are not attacker-controlled in standard Claude Context usage scenarios.