# How Turbovec's Search Kernel Works with Rotated Queries: SIMD Lookup Table Architecture

> Discover how Turbovec's search kernel handles rotated queries using SIMD lookup tables. Learn about its efficient pre-computation and zero runtime rotation overhead for faster vector search.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: internals
- Published: 2026-08-22

---

**Turbovec applies a deterministic rotation matrix to both database and query vectors before indexing, enabling SIMD search kernels to compute distances using pre-computed lookup tables without runtime rotation overhead.**

The `RyanCodrai/turbovec` library implements a high-performance vector search engine that relies on coordinate rotation to optimize quantization and SIMD utilization. Understanding how the **turbovec search kernel rotated queries** pipeline functions is essential for optimizing retrieval performance across ARM and x86 architectures. The system treats both stored vectors and incoming queries in a unified rotated coordinate space, ensuring deterministic results across platforms.

## The Rotation Pipeline

Before any vector enters the index, it passes through a deterministic transformation that scrambles coordinates while preserving inner-product geometry.

### Matrix Generation and Batch Processing

In [`src/rotation.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/rotation.rs), the `Rotation::new(dim, seed)` constructor builds a rotation matrix using bit-wise permutations and XOR-shifts, operating in `O(dim·log B)` time where `B` represents the block size. The `rotate_batch_into(vectors, n, dim, rotation, rotated_scratch)` function in [`src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/encode.rs) handles the actual transformation, processing `n` rows in-place using only fixed-order additions and bit-rotations to guarantee bit-identical results across all platforms.

### Calibration and Quantization

After rotation, vectors undergo TQ+ calibration using per-coordinate shift/scale pairs applied to the rotated coordinates. The [`src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/encode.rs) module then quantizes these rotated-and-calibrated values for storage in a blocked layout (`BLOCK = 32`), matching the cache-friendly structure required by the SIMD kernels.

## Query Handling in the Search Kernel

When `index.search(query, k)` is invoked, the query vector flows through the same rotation matrix via the internal `turbovec::search::rotated` wrapper. The resulting rotated query generates lookup tables (LUTs) that power the architecture-specific SIMD kernels without requiring additional rotation during the hot path.

### ARM NEON Implementation

The `score_4bit_block_neon` kernel in [`src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/search.rs) processes queries by splitting the rotated coordinates into 4-bit nibbles. Each nibble indexes a pre-computed LUT of quantized scores, enabling parallel distance calculations on ARM architectures.

### x86 AVX2 Implementation

In [`src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/search.rs), the `search_multi_query_avx2` function constructs four 8-bit LUTs from the rotated query. The kernel utilizes `_mm256_shuffle_epi8` to shuffle these LUTs against the nibbles of the database codes, performing vectorized accumulation across the rotated space.

### AVX-512 Optimizations

For AVX-512BW, the `search_multi_query_avx512bw` kernel extends the AVX2 approach to process two blocks in parallel using 512-bit shuffles. The AVX-512VBMI + VNNI variant (`search_multi_query_vnni_dispatch`) concatenates the rotated query into a 64-byte LUT and employs `vpermb` with `vpdpbusd` instructions to execute true dot-product calculations over the rotated coordinates.

## Why Rotate Queries?

Rotation serves three critical functions in the turbovec architecture.

**Uniform Distribution**: Random rotation spreads vector energy evenly across dimensions, maximizing the effectiveness of per-coordinate scalar quantization.

**Cache-Friendly Layout**: The rotated representation enables a blocked storage format that aligns with SIMD kernel inner loops, supporting high-throughput vectorized lookups.

**Hardware Agnosticism**: Because rotation occurs deterministically during indexing and query preparation, the same rotated representation functions identically across ARM NEON, AVX2, and AVX-512 without runtime conversion, as validated in [`src/kernel_tests.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/kernel_tests.rs).

## Implementation Examples

### Rust Usage

```rust
use turbovec::{Index, Metric};

fn main() -> anyhow::Result<()> {
    // Create an index for 1536-dim vectors with 4-bit quantization
    let mut idx = Index::builder()
        .dim(1536)
        .bit_width(4)
        .metric(Metric::Cosine)
        .build()?;

    // Add vectors - automatically rotated via rotate_batch_into
    let vectors: Vec<Vec<f32>> = (0..10_000)
        .map(|_| (0..1536).map(|_| rand::random::<f32>()).collect())
        .collect();
    idx.add(vectors)?;

    // Query is rotated internally before SIMD kernel execution
    let query: Vec<f32> = (0..1536).map(|_| rand::random::<f32>()).collect();
    let top_k = idx.search(&query, 10)?;
    println!("Top-10 IDs: {:?}", top_k);
    Ok(())
}

```

### Python Usage

```python
import turbovec

# Initialize index with rotation-enabled parameters

idx = turbovec.Index(dim=1536, bit_width=4, metric="cosine")

# Insert vectors - rotation applied automatically in src/encode.rs

idx.add(vectors)  # NumPy array of shape (N, 1536)

# Search rotates query on-the-fly before invoking the kernel

scores, ids = idx.search(query, k=10)  # query is (1536,) array

print("Top-10 IDs:", ids)

```

## Summary

- Turbovec applies a deterministic rotation matrix from [`src/rotation.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/rotation.rs) to all vectors before quantization, ensuring geometric preservation while scrambling coordinates.
- The `rotate_batch_into` function in [`src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/encode.rs) processes batches using bit-wise operations that guarantee cross-platform bit-identical results.
- Search kernels in [`src/search.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/search.rs) consume pre-rotated queries to build LUTs, with specialized implementations for NEON, AVX2, AVX-512BW, and AVX-512VBMI+VNNI.
- Rotation enables uniform energy distribution and cache-efficient blocked layouts (`BLOCK = 32`), eliminating architecture-specific conversion during search.

## Frequently Asked Questions

### Does Turbovec rotate queries at search time?

Yes. When you call `index.search()`, the query vector passes through the same rotation matrix used during indexing via the internal `turbovec::search::rotated` wrapper. This happens before the SIMD kernels execute, ensuring the query exists in the same coordinate space as the indexed vectors.

### Why does Turbovec use rotation instead of raw vector coordinates?

Rotation spreads vector energy uniformly across dimensions, improving scalar quantization accuracy. It also enables a deterministic, platform-independent representation that supports bit-identical results across ARM and x86 architectures, as verified in [`src/kernel_tests.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/kernel_tests.rs).

### How does the AVX-512 kernel use rotated queries differently than AVX2?

While both kernels build lookup tables from rotated queries, the AVX2 implementation (`search_multi_query_avx2`) uses 256-bit shuffles (`_mm256_shuffle_epi8`), whereas the AVX-512VBMI+VNNI variant (`search_multi_query_vnni_dispatch`) employs 64-byte concatenated LUTs with `vpermb` and `vpdpbusd` for true dot-product calculation over the rotated space.

### Is the rotation deterministic across different machines?

Yes. The `Rotation::new` constructor in [`src/rotation.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/rotation.rs) uses a fixed seed and composition of bit-wise permutations and XOR-shifts that produce identical results on every platform. The `rotate_batch_into` function uses reduction-free, fixed-order operations to guarantee bit-identical outputs regardless of hardware.