TurboVec prepare() for Search Cache Warmup: Eager Cache Materialization Explained

TurboVec's prepare() eagerly materializes lazy search caches—rotation matrix, codebook centroids, and SIMD-blocked layouts—so the first approximate-nearest-neighbor query avoids a one-time latency spike.

The Rust-based vector search library TurboVec, maintained in the RyanCodrai/turbovec repository, relies on heavyweight data structures to accelerate ANN search. To keep index construction cheap, these structures are stored in lazy caches that are populated on demand. Calling prepare() before serving traffic performs a search cache warmup that moves allocation and computation out of the critical query path.

How TurboVec Defers Search Data Structures

Inside TurboQuantIndex, the engine stores three expensive artifacts in OnceCell fields rather than building them during add or load:

  • rotation — the rotation matrix derived from the vector dimension.
  • centroids — the codebook centroids produced by codebook::codebook.
  • blocked — the SIMD-friendly repacking of packed codes produced by pack::repack.

This lazy strategy, implemented in turbovec/src/lib.rs, minimizes upfront cost when vectors are ingested. However, the first call to search() must allocate and fill each cache, which can introduce noticeable latency.

Inside TurboQuantIndex::prepare(): Eager Cache Initialization

The core warmup logic lives in turbovec/src/lib.rs at lines 494–509. The method first confirms the index is materialized by checking that self.dim is known, then calls get_or_init on each OnceCell field:

// turbovec/src/lib.rs
pub fn prepare(&self) {
    let Some(dim) = self.dim else { return };
    self.rotation.get_or_init(|| rotation::make_rotation_matrix(dim));
    self.centroids.get_or_init(|| {
        let (_, c) = codebook::codebook(self.bit_width, dim);
        c
    });
    self.blocked.get_or_init(|| {
        let (data, n_blocks) = pack::repack(&self.packed_codes, self.n_vectors, self.bit_width, dim);
        BlockedCache { data, n_blocks }
    });
}

Each closure runs exactly once. rotation::make_rotation_matrix(dim) builds the rotation matrix. codebook::codebook(self.bit_width, dim) returns the centroids. pack::repack reorganizes self.packed_codes into a BlockedCache suited for SIMD traversal. Because get_or_init is atomic, the operation is safe to invoke from any thread even if searches are already running.

Thread Safety and Idempotency

TurboVec uses OnceCell for each cache, so prepare() is idempotent and thread-safe. You can call it multiple times or concurrently from many threads; only the first successful initialization pays the cost. The test suite in turbovec/tests/concurrent_search.rs demonstrates this pattern under real concurrency.

Practical Usage Patterns for Search Cache Warmup

Typical production scenarios call prepare() after state changes that invalidate or create the underlying index data.

Warm Up After Loading a Persisted Index

When you restore an index from disk with TurboQuantIndex::load, the packed codes exist but the search caches do not. Call prepare() immediately so the first production query is fast:

use turbovec::TurboQuantIndex;

let idx = TurboQuantIndex::load("my_index.tvim").unwrap();
idx.prepare(); // warm-up rotation, centroids, blocked layout
let results = idx.search(&queries, k);

Rebuild Caches After Bulk Insertion

Adding vectors in bulk may leave the blocked layout stale. In TurboVec, invoking prepare() after the batch ensures the rotation, centroids, and blocked caches reflect the current dataset before traffic resumes:

idx.add(&new_vectors).unwrap(); // bulk add leaves caches stale
idx.prepare();                  // rebuild rotation, centroids, blocked layout
let results = idx.search(&queries, k);

Safe Multi-Threaded Warmup

In a multi-threaded server, spawn a background task or call prepare() at startup. The underlying OnceCell guarantees safe publication across threads without locks:

std::thread::spawn(move || {
    idx.prepare(); // safe even if other threads are already searching
    // ...
});

The ID-map wrapper exposes the same warmup path. In turbovec/src/id_map.rs at lines 68–72, TurboQuantIdMap::prepare() simply forwards to the inner index:

use turbovec::TurboQuantIdMap;

let map = TurboQuantIdMap::load("my_index.tvim").unwrap();
map.prepare(); // forwards to inner TurboQuantIndex::prepare()

Summary

  • TurboVec prepare() performs eager search cache warmup by populating three OnceCell fields: rotation matrix, codebook centroids, and SIMD-blocked codes.
  • The implementation lives in turbovec/src/lib.rs#L494-L509 and is safe to call from any thread thanks to OnceCell atomicity.
  • Primary use cases are after TurboQuantIndex::load and after bulk add operations, moving one-time cost out of the query critical path.
  • TurboQuantIdMap::prepare() in turbovec/src/id_map.rs#L68-L72 provides an identical warmup interface for ID-mapped indexes.

Frequently Asked Questions

What does prepare() do in TurboVec?

It eagerly materializes the lazy search caches inside TurboQuantIndex—specifically the rotation matrix, codebook centroids, and SIMD-blocked code layout—so that the first call to search() does not incur allocation and computation latency. The method is implemented in turbovec/src/lib.rs.

When should I call prepare()?

Call it immediately after loading a persisted index or after a bulk add of vectors. This ensures the caches are hot before production traffic arrives. Skipping it does not affect correctness; the caches are built lazily on the first query instead.

Is prepare() thread-safe?

Yes. The method relies on OnceCell::get_or_init, which guarantees that initialization happens exactly once even when many threads call prepare() concurrently. The turbovec/tests/concurrent_search.rs test validates this behavior.

Where is the ID-map version of prepare() implemented?

TurboQuantIdMap::prepare() is a thin wrapper in turbovec/src/id_map.rs#L68-L72 that forwards the call to the underlying TurboQuantIndex::prepare(). It exists so that both index variants share the same warmup semantics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →