How the TurboVec `prepare()` Method Optimizes First Search Latency
The prepare() method eagerly materializes three internal caches—the rotation matrix, Lloyd-Max centroids, and SIMD-blocked layout—before any queries run, eliminating one-time allocation and computation costs that would otherwise inflate the latency of the first search.
When working with vector search in RyanCodrai/turbovec, the prepare() method serves as a critical optimization tool for production deployments requiring consistent query performance. By pre-computing essential data structures that the asymmetric distance computation relies on, this method moves expensive initialization work out of the critical search path. Understanding how TurboQuantIndex::prepare() functions allows you to eliminate unpredictable latency spikes during the first query after loading or modifying an index.
The Three Caches Built by prepare()
The optimization works by forcing eager initialization of three specific caches that search() would otherwise build lazily. In turbovec/src/lib.rs (lines 514-538), the implementation uses OnceLock::get_or_init to populate these structures.
Rotation Matrix Cache
TurboVec rotates vectors into a quantized sub-space before distance calculation. The prepare() method pre-computes this transformation matrix by calling:
self.rotation.get_or_init(|| rotation::make_rotation_matrix(dim))
This allocates and computes the rotation matrix once, rather than calculating it during the first query's execution path.
Lloyd-Max Centroids Cache
For asymmetric distance computation, the index requires a codebook of centroids. The method initializes these via:
self.centroids.get_or_init(|| {
let (_, c) = codebook::codebook(self.bit_width, dim);
c
})
This generates the Lloyd-Max centroids used to map vectors to quantization codes, avoiding the computational cost of codebook generation during active queries.
SIMD-Blocked Layout Cache
To enable vectorized processing, the index packs encoded vectors into blocks optimizes for SIMD instructions:
self.blocked.get_or_init(|| {
let (data, n_blocks) = pack::repack(&self.packed_codes, self.n_vectors, self.bit_width, dim);
BlockedCache { data, n_blocks }
})
This SIMD-blocked layout transformation involves repacking the entire dataset, which can be memory-intensive and slow when performed during a live search request.
Why First Searches Are Slower Without prepare()
Without explicit preparation, the search() method internally triggers the same lazy initializers shown above. When these caches are cold, the first query must sequentially:
- Allocate memory for the rotation matrix and compute its values
- Generate the Lloyd-Max codebook centroids
- Repack all stored vectors into the SIMD-optimized blocked layout
This one-time work adds significant overhead to the first search call. By invoking prepare() before any queries—such as immediately after TurboQuantIndex::load() or following a batch of add() operations—you absorb these costs during initialization, ensuring the first search() experiences only the pure lookup latency.
Key Implementation Characteristics
According to the source code in turbovec/src/id_map.rs (lines 270-272) and the core implementation, prepare() exhibits several important properties:
- Idempotent and thread-safe – Uses
OnceLockprimitives, so concurrent calls race to initialize but only the first succeeds; subsequent calls return instantly without recomputation. - Early-exit for empty indexes – If the index has never received an
add()call (indicated bydimbeingNone), the method returns immediately since no caches require building. - Optional but recommended – The index remains functional without calling
prepare();search()will lazily initialize caches on demand. The method simply provides explicit control over when initialization costs occur.
Usage Examples
Rust Implementation
use turbovec::TurboQuantIndex;
// Load or create an index
let idx = TurboQuantIndex::load("my_index.tqv").unwrap();
// Prime the caches before the first search
idx.prepare(); // Moves one-off cost out of query path
// Now the first search is fast
let results = idx.search(&query_vector, 10);
Python Binding
import turbovec
# Load a saved index
idx = turbovec.load("my_index.tqv")
# Prime the internal caches (optional)
idx.prepare() # Removes first-search overhead
# Perform a search
scores, ids = idx.search(query, k=10)
Performance Impact
In benchmarks, calling prepare() immediately after loading a populated index typically reduces first-search latency by approximately 30–50 milliseconds, depending on vector dimensionality and hardware specifications. This improvement stems from eliminating the heavy rotation matrix computation, centroid generation, and code repacking work from the active query path, as documented in the API reference at docs/api.md (line 47).
Summary
prepare()eliminates first-search latency spikes by eagerly building three critical caches: the rotation matrix, Lloyd-Max centroids, and SIMD-blocked layout.- The method is idempotent and thread-safe, using
OnceLockto ensure safe concurrent initialization without redundant computation. - Early-exit logic prevents unnecessary work on empty indexes that have not yet received vector additions.
- While optional, calling
prepare()after loading or batch-adding vectors ensures predictable, low-latency performance from the first query. - Typical latency reduction ranges from 30–50ms for the initial search call.
Frequently Asked Questions
Is calling prepare() required before searching?
No, calling prepare() is optional. The search() method will automatically trigger lazy initialization of the rotation matrix, centroids, and blocked layout if they haven't been built yet. However, without explicit preparation, the first query will include the one-time cost of building these structures, resulting in higher latency for that specific request.
Is the prepare() method thread-safe?
Yes, prepare() is fully thread-safe. The implementation uses Rust's OnceLock primitives for cache initialization, ensuring that concurrent calls race to initialize each cache but only the first caller succeeds in building the data structure. Subsequent calls return instantly without recomputation or blocking, making the method safe to call from multiple threads simultaneously.
What happens if I call prepare() on an empty index?
The method returns immediately without performing any work. If the index has never received an add() call (indicated by dim being None), there are no vectors to prepare and no caches to build. This early-exit behavior prevents unnecessary allocation and computation on uninitialized indexes.
How much latency does prepare() actually save?
Benchmarks indicate that calling prepare() typically reduces first-search latency by approximately 30–50 milliseconds, though the exact improvement depends on vector dimensionality, index size, and hardware capabilities. This saving represents the cost of allocating memory for the rotation matrix, generating Lloyd-Max centroids, and repacking codes into the SIMD-optimized layout—operations that would otherwise occur during the first query execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →