How Turbovec's Search Kernel Works with Rotated Queries: SIMD Lookup Table Architecture

Turbovec applies a deterministic rotation matrix to both database and query vectors before indexing, enabling SIMD search kernels to compute distances using pre-computed lookup tables without runtime rotation overhead.

The RyanCodrai/turbovec library implements a high-performance vector search engine that relies on coordinate rotation to optimize quantization and SIMD utilization. Understanding how the turbovec search kernel rotated queries pipeline functions is essential for optimizing retrieval performance across ARM and x86 architectures. The system treats both stored vectors and incoming queries in a unified rotated coordinate space, ensuring deterministic results across platforms.

The Rotation Pipeline

Before any vector enters the index, it passes through a deterministic transformation that scrambles coordinates while preserving inner-product geometry.

Matrix Generation and Batch Processing

In src/rotation.rs, the Rotation::new(dim, seed) constructor builds a rotation matrix using bit-wise permutations and XOR-shifts, operating in O(dim·log B) time where B represents the block size. The rotate_batch_into(vectors, n, dim, rotation, rotated_scratch) function in src/encode.rs handles the actual transformation, processing n rows in-place using only fixed-order additions and bit-rotations to guarantee bit-identical results across all platforms.

Calibration and Quantization

After rotation, vectors undergo TQ+ calibration using per-coordinate shift/scale pairs applied to the rotated coordinates. The src/encode.rs module then quantizes these rotated-and-calibrated values for storage in a blocked layout (BLOCK = 32), matching the cache-friendly structure required by the SIMD kernels.

Query Handling in the Search Kernel

When index.search(query, k) is invoked, the query vector flows through the same rotation matrix via the internal turbovec::search::rotated wrapper. The resulting rotated query generates lookup tables (LUTs) that power the architecture-specific SIMD kernels without requiring additional rotation during the hot path.

ARM NEON Implementation

The score_4bit_block_neon kernel in src/search.rs processes queries by splitting the rotated coordinates into 4-bit nibbles. Each nibble indexes a pre-computed LUT of quantized scores, enabling parallel distance calculations on ARM architectures.

x86 AVX2 Implementation

In src/search.rs, the search_multi_query_avx2 function constructs four 8-bit LUTs from the rotated query. The kernel utilizes _mm256_shuffle_epi8 to shuffle these LUTs against the nibbles of the database codes, performing vectorized accumulation across the rotated space.

AVX-512 Optimizations

For AVX-512BW, the search_multi_query_avx512bw kernel extends the AVX2 approach to process two blocks in parallel using 512-bit shuffles. The AVX-512VBMI + VNNI variant (search_multi_query_vnni_dispatch) concatenates the rotated query into a 64-byte LUT and employs vpermb with vpdpbusd instructions to execute true dot-product calculations over the rotated coordinates.

Why Rotate Queries?

Rotation serves three critical functions in the turbovec architecture.

Uniform Distribution: Random rotation spreads vector energy evenly across dimensions, maximizing the effectiveness of per-coordinate scalar quantization.

Cache-Friendly Layout: The rotated representation enables a blocked storage format that aligns with SIMD kernel inner loops, supporting high-throughput vectorized lookups.

Hardware Agnosticism: Because rotation occurs deterministically during indexing and query preparation, the same rotated representation functions identically across ARM NEON, AVX2, and AVX-512 without runtime conversion, as validated in src/kernel_tests.rs.

Implementation Examples

Rust Usage

use turbovec::{Index, Metric};

fn main() -> anyhow::Result<()> {
    // Create an index for 1536-dim vectors with 4-bit quantization
    let mut idx = Index::builder()
        .dim(1536)
        .bit_width(4)
        .metric(Metric::Cosine)
        .build()?;

    // Add vectors - automatically rotated via rotate_batch_into
    let vectors: Vec<Vec<f32>> = (0..10_000)
        .map(|_| (0..1536).map(|_| rand::random::<f32>()).collect())
        .collect();
    idx.add(vectors)?;

    // Query is rotated internally before SIMD kernel execution
    let query: Vec<f32> = (0..1536).map(|_| rand::random::<f32>()).collect();
    let top_k = idx.search(&query, 10)?;
    println!("Top-10 IDs: {:?}", top_k);
    Ok(())
}

Python Usage

import turbovec

# Initialize index with rotation-enabled parameters

idx = turbovec.Index(dim=1536, bit_width=4, metric="cosine")

# Insert vectors - rotation applied automatically in src/encode.rs

idx.add(vectors)  # NumPy array of shape (N, 1536)

# Search rotates query on-the-fly before invoking the kernel

scores, ids = idx.search(query, k=10)  # query is (1536,) array

print("Top-10 IDs:", ids)

Summary

  • Turbovec applies a deterministic rotation matrix from src/rotation.rs to all vectors before quantization, ensuring geometric preservation while scrambling coordinates.
  • The rotate_batch_into function in src/encode.rs processes batches using bit-wise operations that guarantee cross-platform bit-identical results.
  • Search kernels in src/search.rs consume pre-rotated queries to build LUTs, with specialized implementations for NEON, AVX2, AVX-512BW, and AVX-512VBMI+VNNI.
  • Rotation enables uniform energy distribution and cache-efficient blocked layouts (BLOCK = 32), eliminating architecture-specific conversion during search.

Frequently Asked Questions

Does Turbovec rotate queries at search time?

Yes. When you call index.search(), the query vector passes through the same rotation matrix used during indexing via the internal turbovec::search::rotated wrapper. This happens before the SIMD kernels execute, ensuring the query exists in the same coordinate space as the indexed vectors.

Why does Turbovec use rotation instead of raw vector coordinates?

Rotation spreads vector energy uniformly across dimensions, improving scalar quantization accuracy. It also enables a deterministic, platform-independent representation that supports bit-identical results across ARM and x86 architectures, as verified in src/kernel_tests.rs.

How does the AVX-512 kernel use rotated queries differently than AVX2?

While both kernels build lookup tables from rotated queries, the AVX2 implementation (search_multi_query_avx2) uses 256-bit shuffles (_mm256_shuffle_epi8), whereas the AVX-512VBMI+VNNI variant (search_multi_query_vnni_dispatch) employs 64-byte concatenated LUTs with vpermb and vpdpbusd for true dot-product calculation over the rotated space.

Is the rotation deterministic across different machines?

Yes. The Rotation::new constructor in src/rotation.rs uses a fixed seed and composition of bit-wise permutations and XOR-shifts that produce identical results on every platform. The rotate_batch_into function uses reduction-free, fixed-order operations to guarantee bit-identical outputs regardless of hardware.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →