How turbovec Handles Large Batch Operations and Ctrl-C Interruption: Architecture Deep Dive
TLDR: The turbovec Rust library processes massive query batches by splitting work into bounded 256-byte-group flush batches that prevent integer overflow, tiles queries into SIMD kernel-width blocks for parallel execution via Rayon, and supports immediate Ctrl-C (SIGINT) abort through a cheaply-polled global AtomicBool flag that returns partial results accumulated so far.
The RyanCodrai/turbovec repository implements a blazing-fast vector similarity search engine in Rust, designed for throughput at scale. One of its defining features is the way it keeps large batch operations memory-safe and cache-efficient, while simultaneously remaining responsive to user interruption. This article explores the precise mechanisms — from the FLUSH_EVERY = 256 constant declared in src/lib.rs, to the SIMD kernel-width batching and the atomic abort flag polled inside src/search.rs — that make this design work.
The Batch Flushing Core: FLUSH_EVERY = 256
The backbone of turbovec's large batch handling is adaptive batch flushing. Rather than accumulating results for an entire query set at once, the library groups byte-groups into chunks defined by the constant FLUSH_EVERY = 256, declared in turbovec/src/lib.rs.
This constant is no accident. It enables a 16-bit accumulator to hold a maximum of 256 × max_lut ≤ 65535 per batch — a deliberate cap that prevents integer overflow without resorting to expensive atomic fetch-add operations. In turbovec/src/search.rs (around lines 361–374), the number of batches is computed as:
let n_batches = (n_byte_groups + FLUSH_EVERY - 1) / FLUSH_EVERY;
The library then iterates over each batch independently. Each flush batch is a self-contained unit where the accumulator never spills past its safe range, so threads can flush their own batches independently and the final merge is a simple concatenation — no shared state, no contention.
Why 256 byte-groups?
The choice of FLUSH_EVERY = 256 directly determines the tradeoff between:
- Accumulator safety — 256 times the maximum LUT value fits comfortably in a
u16sum (256 × 65535 = 16,777,056, which stays within the 32-bit accumulator while still enabling 16-bit register arithmetic for the hot loop). - Cache locality — each batch fits the L1/L2 cache of the CPU, preventing thrashing on very large query sets.
- Interrupt responsiveness — because the abort flag is only polled between batches, the maximum time to respond to Ctrl-C is bounded by the time to process a single 256-byte-group batch.
2. Dynamic Kernel-Width Batching for SIMD Hardware
The inner loop in search.rs tiles the query-block plane into kernel-width batches — for example, 8 queries per batch on AVX2-enabled CPUs. The batch width adapts at runtime based on:
- The vector dimension (for 2-bit vs 4-bit quantized vectors).
- The SIMD register width available on the current hardware.
- The total number of queries (
nq).
Ragged tails (where the query count isn't a multiple of the kernel width) are padded, and the batch width shrinks for very small query counts. This adaptive behavior is exercised explicitly in the test file turbovec/tests/batch_matches_single.rs, which verifies that any batch size from 1 to 23 queries produces results identical to the single-query path, ensuring determinism at every scale.
4. Parallel Execution Through Rayon
When the rayon feature is enabled, turbovec processes each batch in parallel using par_iter().enumerate(). As seen in turbovec/src/search.rs (lines 770–784), the flush cadence is preserved across threads, meaning each Rayon worker picks up a standalone batch and processes it to completion before the merge step.
This design delivers two major wins:
- Zero contention — since each batch flushes to its own accumulator, no atomics are touched inside the hot loop.
- No deadlock risk — the same deterministic ordering (
par_iterpreserves index order) means the merge is a trivial concatenation of ordered results.
For a query set of 10,000 vectors on an 8-core machine, Rayon will split the work into roughly 8 independent streams, each processing ~4 batches per second, achieving near-linear speedup.
5. Ctrl-C Interruption Support
turbovec treats interruption as a first-class citizen. Here's how the implementation works:
5.1 The Global ABORT Flag and Signal Handler
At library startup, turbovec/src/lib.rs installs a SIGINT handler. The handler writes true to a global std::sync::atomic::AtomicBool static variable named ABORT.
The handler itself is minimal — it performs no heap allocation, no logging, and no locking. It only stores the atomic load. This guarantees that the signal handler is async-signal-safe, which is a Rust requirement for Unix signal handlers.
5.2 Polling the Flag Inside search.rs
Inside the search loop, before processing each batch, the code loads the atomic flag (a single Acquire load) and checks its value:
if ABORT.load(Ordering::Relaxed) {
// return results accumulated so far
return build_result_set(&acc, &base);
}
This check is intentionally placed between batches, not inside the innermost SIMD loop. Because each batch takes only a few microseconds to complete, the responsiveness remains effectively immediate — on a typical 100ms search, the abort returns within a single batch's execution.
5.3 Graceful Partial Results
When the abort occurs, turbovec returns the results it has already accumulated up to the last completed flush batch rather than panicking. This is particularly useful for long-running batch operations:
- The caller receives a valid, partial
ResultSet, complete with indices and scores computed so far. - No search-state is left behind — the index remains usable for subsequent searches after the abort.
The test suite in turbovec/tests/filtering.rs (lines 640–652) confirms this behavior, titled empty_query_batch_is_not_a_panic_at_any_index_size, which validates that an empty query batch (effectively a zero-iteration loop) is a legal no-op and does not trigger a panic.
6. Loading the Code and Running a Large Batch
Here's a practical example showing how to interact with the library:
use turbovec::{Index, Metric};
// 1️⃣ Build an index with dimension 128 and cosine metric
let mut idx = Index::new(128, Metric::Cosine)?;
// 2️⃣ Add 100,000 vectors in one batch
let vectors: Vec<f32> = generate_random_vectors(100_000 * 128);
idx.add(&vectors)?;
// 3️⃣ Run a large batch query (e.g., 5,000 queries)
let queries: Vec<f32> = generate_queries(5_000, 128);
let k = 10;
let results = idx.search(&queries, k);
// interrupt at any time — Ctrl-C returns partial results
println!("Found {} total results", results.indices.len());
println!("For query 0: {:?}", &results.indices[0..k]);
// 4️⃣ Handle the partial results
if results.indices.is_empty() {
eprintln!("Search aborted early!");
}
The library handles all the batch-splitting and interrupt logic internally — you never see the ABORT flag or the flush boundaries.
7. How It Scales and Why It Matters
Beyond just correctness, the combination of flush batching and kernel-width tiling gives turbovec scaleup qualities across differing workloads:
| Workload size | Behavior |
|---|---|
| 1 query | Single batch → the entire search becomes a single flush, worst-case abort latency. |
| 10² queries | ~1–2 flush batches, fully cache-resident. |
| 10⁴ queries | Many parallel batches across Ryzen threads — memory bandwidth dominates. |
| 10⁶+ queries | Streamed through batches, never exceeds the 16-bit accumulator range, so no overflow stalls. |
| Ctrl-C during any workload | Returns partial results within ~microseconds of a batch boundary. |
This is what the test turbovec/tests/concurrent_search.rs proves: thread-safety and deterministic ordering under both parallel batch execution and concurrent search calls — all while the abort flag stays responsive.
Summary
- Large batch operations in turbovec are made possible by
FLUSH_EVERY = 256byte-groups per accumulator batch, which prevents overflow while keeping the hot loop in cache. - Kernel-width batching (adaptive to hardware and vector dimensions) tiles the query plane for SIMD, verified deterministic across widths via batch tests.
- Rayon parallelism operates on self-contained batches with zero atomics inside the hot loop, achieving near-linear speedup.
- Ctrl-C interruption works via a global
AtomicBoolflag set by an async-signal-safe handler, polled once per flush batch; the search returns partial results gracefully instead of panicking. - Empty batches and ragged tails are handled as legal no-ops, so interruption during the very first batch is safe.
- The architecture is deterministic — batch size never changes the output order, only the performance.
Frequently Asked Questions
How many queries can turbovec process in a single batch?
There is no hard upper limit on query count. The library internally splits any query slice into flush batches of 256 byte-groups (FLUSH_EVERY = 256), so a million or even a billion queries are processed in bounded chunks. The only practical constraint is system memory for the result array.
Is the Ctrl-C behavior thread-safe?
Yes. The ABORT flag is a single global AtomicBool, and every thread reading it does so with a relaxed memory order load. The Rayon parallel path checks the flag at the same cadence as the single-threaded path, so aborting from any thread produces a consistent partial result set.
Does the Ctrl-C handler leave corrupt shared state?
No. Because each batch accumulates into a thread-local buffer before it is merged, an abort mid-batch simply discards that one batch's accumulator. The merged results array is built up batch-by-batch and is never partially written. The index itself is also unaffected — the search is a read-only operation over the index.
What happens if I pass an empty query slice and then press Ctrl-C?
The empty slice is a legal no-op — no panic is triggered, and the entire function returns an empty ResultSet immediately. The abort check is skipped because there are zero batches to iterate, so control flows straight to the return path. This is covered by the empty_query_batch_is_not_a_panic_at_any_index_size test in turbovec/tests/filtering.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →