# How the TurboVec `prepare()` Method Optimizes First Search Latency

> Discover how TurboVec's prepare() method optimizes first search latency by pre-materializing internal caches, eliminating allocation and computation costs for faster queries.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: performance
- Published: 2026-07-18

---

**The `prepare()` method eagerly materializes three internal caches—the rotation matrix, Lloyd-Max centroids, and SIMD-blocked layout—before any queries run, eliminating one-time allocation and computation costs that would otherwise inflate the latency of the first search.**

When working with vector search in **RyanCodrai/turbovec**, the `prepare()` method serves as a critical optimization tool for production deployments requiring consistent query performance. By pre-computing essential data structures that the asymmetric distance computation relies on, this method moves expensive initialization work out of the critical search path. Understanding how `TurboQuantIndex::prepare()` functions allows you to eliminate unpredictable latency spikes during the first query after loading or modifying an index.

## The Three Caches Built by prepare()

The optimization works by forcing eager initialization of three specific caches that `search()` would otherwise build lazily. In [`turbovec/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs) (lines 514-538), the implementation uses `OnceLock::get_or_init` to populate these structures.

### Rotation Matrix Cache

**TurboVec** rotates vectors into a quantized sub-space before distance calculation. The `prepare()` method pre-computes this transformation matrix by calling:

```rust
self.rotation.get_or_init(|| rotation::make_rotation_matrix(dim))

```

This allocates and computes the rotation matrix once, rather than calculating it during the first query's execution path.

### Lloyd-Max Centroids Cache

For asymmetric distance computation, the index requires a codebook of centroids. The method initializes these via:

```rust
self.centroids.get_or_init(|| {
    let (_, c) = codebook::codebook(self.bit_width, dim);
    c
})

```

This generates the **Lloyd-Max centroids** used to map vectors to quantization codes, avoiding the computational cost of codebook generation during active queries.

### SIMD-Blocked Layout Cache

To enable vectorized processing, the index packs encoded vectors into blocks optimizes for SIMD instructions:

```rust
self.blocked.get_or_init(|| {
    let (data, n_blocks) = pack::repack(&self.packed_codes, self.n_vectors, self.bit_width, dim);
    BlockedCache { data, n_blocks }
})

```

This **SIMD-blocked layout** transformation involves repacking the entire dataset, which can be memory-intensive and slow when performed during a live search request.

## Why First Searches Are Slower Without prepare()

Without explicit preparation, the `search()` method internally triggers the same lazy initializers shown above. When these caches are cold, the first query must sequentially:

1. Allocate memory for the rotation matrix and compute its values
2. Generate the Lloyd-Max codebook centroids
3. Repack all stored vectors into the SIMD-optimized blocked layout

This one-time work adds significant overhead to the first search call. By invoking `prepare()` **before** any queries—such as immediately after `TurboQuantIndex::load()` or following a batch of `add()` operations—you absorb these costs during initialization, ensuring the first `search()` experiences only the pure lookup latency.

## Key Implementation Characteristics

According to the source code in [`turbovec/src/id_map.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/id_map.rs) (lines 270-272) and the core implementation, `prepare()` exhibits several important properties:

- **Idempotent and thread-safe** – Uses `OnceLock` primitives, so concurrent calls race to initialize but only the first succeeds; subsequent calls return instantly without recomputation.
- **Early-exit for empty indexes** – If the index has never received an `add()` call (indicated by `dim` being `None`), the method returns immediately since no caches require building.
- **Optional but recommended** – The index remains functional without calling `prepare()`; `search()` will lazily initialize caches on demand. The method simply provides explicit control over when initialization costs occur.

## Usage Examples

### Rust Implementation

```rust
use turbovec::TurboQuantIndex;

// Load or create an index
let idx = TurboQuantIndex::load("my_index.tqv").unwrap();

// Prime the caches before the first search
idx.prepare();  // Moves one-off cost out of query path

// Now the first search is fast
let results = idx.search(&query_vector, 10);

```

### Python Binding

```python
import turbovec

# Load a saved index

idx = turbovec.load("my_index.tqv")

# Prime the internal caches (optional)

idx.prepare()  # Removes first-search overhead

# Perform a search

scores, ids = idx.search(query, k=10)

```

## Performance Impact

In benchmarks, calling `prepare()` immediately after loading a populated index typically reduces first-search latency by approximately **30–50 milliseconds**, depending on vector dimensionality and hardware specifications. This improvement stems from eliminating the heavy rotation matrix computation, centroid generation, and code repacking work from the active query path, as documented in the API reference at [`docs/api.md`](https://github.com/RyanCodrai/turbovec/blob/main/docs/api.md) (line 47).

## Summary

- **`prepare()`** eliminates first-search latency spikes by eagerly building three critical caches: the rotation matrix, Lloyd-Max centroids, and SIMD-blocked layout.
- The method is **idempotent and thread-safe**, using `OnceLock` to ensure safe concurrent initialization without redundant computation.
- **Early-exit logic** prevents unnecessary work on empty indexes that have not yet received vector additions.
- While **optional**, calling `prepare()` after loading or batch-adding vectors ensures predictable, low-latency performance from the first query.
- Typical latency reduction ranges from **30–50ms** for the initial search call.

## Frequently Asked Questions

### Is calling `prepare()` required before searching?

No, calling `prepare()` is optional. The `search()` method will automatically trigger lazy initialization of the rotation matrix, centroids, and blocked layout if they haven't been built yet. However, without explicit preparation, the first query will include the one-time cost of building these structures, resulting in higher latency for that specific request.

### Is the `prepare()` method thread-safe?

Yes, `prepare()` is fully thread-safe. The implementation uses Rust's `OnceLock` primitives for cache initialization, ensuring that concurrent calls race to initialize each cache but only the first caller succeeds in building the data structure. Subsequent calls return instantly without recomputation or blocking, making the method safe to call from multiple threads simultaneously.

### What happens if I call `prepare()` on an empty index?

The method returns immediately without performing any work. If the index has never received an `add()` call (indicated by `dim` being `None`), there are no vectors to prepare and no caches to build. This early-exit behavior prevents unnecessary allocation and computation on uninitialized indexes.

### How much latency does `prepare()` actually save?

Benchmarks indicate that calling `prepare()` typically reduces first-search latency by approximately **30–50 milliseconds**, though the exact improvement depends on vector dimensionality, index size, and hardware capabilities. This saving represents the cost of allocating memory for the rotation matrix, generating Lloyd-Max centroids, and repacking codes into the SIMD-optimized layout—operations that would otherwise occur during the first query execution.