Using prepare() to Warm Up Search Caches in TurboVec

Call TurboQuantIndex::prepare() immediately after constructing or loading an index to eagerly initialize the search caches and eliminate first-query latency penalties.

TurboVec is a high-performance vector search library that quantizes high-dimensional vectors for fast similarity search. When building latency-sensitive applications, you must explicitly warm up the internal caches that the search engine relies on. The prepare() method forces the lazy initialization of these heavy-weight data structures so that your first query is as fast as every subsequent one.

What prepare() Initializes

The prepare() method materializes three deterministic caches that would otherwise be built on the first call to search. According to the implementation in turbovec/src/lib.rs (lines 523–538), these include:

  • Rotation matrix – A deterministic dim × dim matrix used to rotate vectors before quantization.
  • Codebook centroids – The quantizer’s lookup tables computed for the chosen bit-width.
  • Blocked SIMD layout – A cache storing packed codes in a SIMD-friendly block arrangement for vectorized distance computation.

This work is performed only once per index and is guarded by OnceLock, making subsequent calls essentially free.

Why Call prepare() Before Querying

The search method takes &self and can be invoked from many threads without external locking. However, the first query pays the one-time cost of building the caches described above, which can add noticeable latency to the critical path. By calling prepare() right after constructing or loading an index—or after a batch of add operations—you shift that initialization cost to a non-critical initialization phase, leading to predictable, low-latency response times for all queries.

Thread Safety and Concurrency Guarantees

prepare() is safe to call from multiple threads simultaneously. The implementation uses OnceLock::get_or_init to ensure that exactly one thread performs the work while others block until completion. This makes the method idempotent; calling it repeatedly from concurrent workers has no negative side effects and will not re-initialize the caches once they are warmed up.

The test suite confirms these guarantees in turbovec/tests/concurrent_search.rs (lines 120–132) and verifies cache invalidation semantics in turbovec/tests/state_sequences.rs (lines 185–200).

Implementation Details

You can find the method definition and its documentation in turbovec/src/lib.rs between lines 514 and 538. The surrounding comments (lines 514–522) explicitly describe its purpose and concurrency guarantees.

The method short-circuits to a no-op if the index is still in lazy mode—meaning no vectors have been added and the dimensionality is unknown—because the caches depend on knowing dim.

Code Examples

Basic Usage After Adding Vectors

Warm up the caches after populating the index to ensure the first search is fast:

use turbovec::TurboQuantIndex;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // Construct an index (1536-dim, 4-bit encoding)
    let mut idx = TurboQuantIndex::new(1536, 4)?;

    // Add a batch of vectors (flattened f32 slice)
    let vectors = vec![0.0_f32; 1536 * 10];
    idx.add(&vectors);

    // Warm up the caches before serving queries
    idx.prepare();          // forces rotation, centroids, blocked cache

    // Perform a search – now the first query has no extra latency
    let queries = vec![0.0_f32; 1536 * 2];
    let results = idx.search(&queries, 5);
    println!("Top-k scores: {:?}", results.scores);
    Ok(())
}

Warming Up a Persisted Index

When loading an index from disk, call prepare() immediately to avoid cache initialization during the first query:

let idx = TurboQuantIndex::load("my_index.tv")?;
idx.prepare();               // useful when the index is freshly loaded
let results = idx.search(&my_queries, 10);

Concurrent Preparation

Multiple worker threads can safely invoke prepare() without coordination:

use std::thread;

let idx = TurboQuantIndex::load("my_index.tv")?;
let handles: Vec<_> = (0..4)
    .map(|_| {
        let idx_ref = &idx;
        thread::spawn(move || {
            idx_ref.prepare();            // safe, will run once
            // … now run many searches concurrently …
        })
    })
    .collect();

for h in handles { h.join().unwrap(); }

Summary

  • TurboQuantIndex::prepare() eagerly initializes the rotation matrix, codebook centroids, and blocked SIMD layout.
  • Call it after construction, loading, or batch additions to shift initialization cost out of the query path.
  • The method is thread-safe and idempotent, using OnceLock to ensure exactly one thread performs the work.
  • It is a no-op if the index dimensionality is unknown (no vectors added yet).
  • The implementation resides in turbovec/src/lib.rs at lines 523–538.

Frequently Asked Questions

When should I call prepare() in my application lifecycle?

Call prepare() immediately after constructing a new index, loading a persisted index from disk, or completing a batch of add operations. This ensures that the heavy cache initialization occurs before you serve user queries, keeping response times consistent and predictable.

Is prepare() safe to call from multiple threads simultaneously?

Yes. The method uses OnceLock::get_or_init internally, guaranteeing that exactly one thread builds the caches while others wait for the result. Once initialized, subsequent calls return immediately without repeating the work, making it safe to invoke from any thread without external synchronization.

Do I need to call prepare() after adding new vectors?

Yes, if you want to maintain optimal query performance. Adding vectors can invalidate the blocked cache (as shown in turbovec/tests/state_sequences.rs), so you should call prepare() again after your final batch of add operations and before high-volume searching begins.

What happens if I call prepare() on an empty index?

The method returns immediately without performing any work. If no vectors have been added and the dimensionality is unknown, prepare() cannot compute the required caches and exits as a no-op. It will only initialize the caches once vectors have been added and the index dimensions are established.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →