Memory Overhead of the slot_to_id Map in IdMapIndex: Per-Vector Cost Breakdown
The slot_to_id map in IdMapIndex consumes 8 bytes per indexed vector, and when combined with the inverse id_to_slot hash map, the total bidirectional mapping overhead is approximately 32 bytes per vector.
The IdMapIndex struct in the turbovec crate manages stable external IDs over a dense quantized index. Understanding the memory overhead of the slot_to_id map in IdMapIndex is critical for sizing workloads, as this layer sits on top of the actual vector storage. According to the RyanCodrai/turbovec source code, the mapping is implemented in turbovec/src/id_map.rs using a Vec<u64> paired with a HashMap<u64, usize>.
How IdMapIndex Manages the Slot-to-ID Mapping
IdMapIndex maintains two complementary structures to translate between dense internal slots and stable external identifiers.
slot_to_id— aVec<u64>where each index corresponds to an internal slot and each value stores the external ID assigned to the vector in that slot.id_to_slot— aHashMap<u64, usize>providing the inverse lookup from external ID back to the internal slot index.
This design allows vectors to be deleted or moved without changing their public IDs.
Breaking Down the Memory Overhead of the slot_to_id Map
slot_to_id Vec Storage
Each entry in slot_to_id is a plain u64, which occupies 8 bytes per slot. The Vec itself incurs a fixed 24-byte header for its pointer, length, and capacity on 64-bit platforms, but this header does not scale with the number of vectors.
id_to_slot HashMap<u64, usize> Storage
The inverse map stores key-value pairs alongside hash metadata. On a 64-bit platform, each entry requires roughly:
- 8 bytes for the
u64key - 8 bytes for the
usizevalue - 8 bytes for the stored hash and bucket overhead
This yields approximately 24 bytes per entry in the HashMap.
Total Mapping Overhead per Vector
Adding the two structures together gives the marginal cost for every indexed vector:
slot_to_id: 8 bytesid_to_slot: ~24 bytes- Total: ~32 bytes per vector
This cost is independent of the quantized vector payload.
Where the Mapping Is Defined in the turbovec Source
The two fields are declared in turbovec/src/id_map.rs at lines 47–52:
slot_to_id: Vec<u64>id_to_slot: HashMap<u64, usize>
The IdMapIndex type is publicly re-exported from turbovec/src/lib.rs, and unit tests covering insertion, removal, and lookup behavior reside in turbovec/tests/id_map.rs.
Measuring Memory Usage in Practice
You can approximate the mapping overhead at runtime using the crate's internal structures or standard library size calculations.
use turbovec::IdMapIndex;
let mut idx = IdMapIndex::new(1536, 4).unwrap();
let vectors = vec![0.0_f32; 1536 * 3];
idx.add_with_ids(&vectors, &[1001, 1002, 1003]).unwrap();
// Approximate overhead: 3 vectors × 32 bytes ≈ 96 bytes
println!("slots: {}", idx.len());
println!("slot→id vec capacity: {}", idx.slot_to_id.capacity());
println!("id→slot hashmap buckets: {}", idx.id_to_slot.capacity());
For a raw breakdown of the Rust types involved:
use std::mem::size_of;
let per_slot_to_id = size_of::<u64>(); // 8 bytes
let per_id_to_slot = size_of::<(u64, usize)>(); // 16 bytes
let hash_overhead = size_of::<usize>(); // 8 bytes (hash stored per bucket)
let per_entry = per_slot_to_id + per_id_to_slot + hash_overhead; // ≈ 32 bytes
println!("≈ {} bytes per vector for the bidirectional map", per_entry);
Summary
- The
slot_to_idfield inIdMapIndexis aVec<u64>that costs 8 bytes per vector. - The companion
id_to_slothash map adds roughly 24 bytes per vector. - Together, the bidirectional mapping layer introduces approximately 32 bytes of memory overhead per indexed vector.
- These structures are defined in
turbovec/src/id_map.rsand are required for stable ID support on top of the dense quantized index.
Frequently Asked Questions
What is the memory overhead of the slot_to_id map in IdMapIndex?
The slot_to_id map itself—the Vec<u64> declared in turbovec/src/id_map.rs—contributes 8 bytes per vector. This flat array stores the external ID for each occupied internal slot in the underlying TurboQuantIndex.
How does the id_to_slot HashMap affect total memory usage?
The id_to_slot HashMap<u64, usize> provides the inverse lookup and costs approximately 24 bytes per entry on 64-bit systems. Combined with slot_to_id, the total mapping overhead is roughly 32 bytes per vector.
Are there fixed allocation costs even for an empty index?
Yes. The Vec<u64> holds a 24-byte header regardless of length, and the HashMap allocates a small control structure for its bucket array even when empty. These baseline costs are negligible once the index holds more than a few vectors.
Where can I inspect the source code for these structures?
The mapping fields are defined in turbovec/src/id_map.rs between lines 47 and 52. Public exports live in turbovec/src/lib.rs, and integration tests are available in turbovec/tests/id_map.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →