Handling NaN/Inf Values in Vectors During Add and Search in TurboVec
TurboVec strictly rejects any vector containing NaN, Infinity, or values with magnitude ≥1e16 before they reach the SIMD encoding kernels, preventing silent index corruption through the first_invalid_coord validation helper.
TurboVec is a high-performance approximate nearest neighbor search library that packs vectors into quantized codes with per-vector scales. Because non-finite values poison the SIMD kernels used for encoding and scoring, the library implements defensive input validation that intercepts invalid data during both indexing and query operations in the RyanCodrai/turbovec repository.
Why NaN and Inf Values Destroy Vector Index Integrity
The Poisoning Mechanism
TurboVec stores vectors as packed codes alongside a per-vector scale factor in vec_scales[slot]. When a vector contains NaN, the expression 0 * NaN = NaN propagates to the scale storage, creating unreachable slots that appear in len() but return corrupted results during search. Similarly, Infinity values result in 1 / Inf = 0 scales that effectively zero out vector magnitudes, rendering them invisible to similarity search.
The Overflow Risk
Values with extremely large magnitude (≥ 1e16) cause the sum-of-squares norm calculation to overflow to +Inf during encoding. According to the TurboVec source code in turbovec/src/lib.rs, this produces an infinite scale factor that dominates every top-k query, causing the corrupted vector to win every similarity search regardless of actual content.
Input Validation Architecture in TurboVec
The first_invalid_coord Helper
All validation flows through a single helper function defined in turbovec/src/lib.rs at lines 94-101. This function scans input slices and returns the first offending coordinate:
pub fn first_invalid_coord(values: &[f32], dim: usize) -> Option<(usize, usize, f32)> {
for (i, x) in values.iter().enumerate() {
if !x.is_finite() || x.abs() >= MAX_INPUT_MAGNITUDE {
let vector_index = if dim == 0 { 0 } else { i / dim };
let coord_index = if dim == 0 { i } else { i % dim };
return Some((vector_index, coord_index, *x));
}
}
None
}
The 1e16 Magnitude Threshold
The constant MAX_INPUT_MAGNITUDE is defined as 1e16 in turbovec/src/lib.rs (lines 75-80). This threshold sits far above realistic embedding magnitudes but safely below the f32 overflow threshold, preventing the norm calculation from producing infinity during SIMD operations.
Validation in Add and Search Operations
Add Path Validation
Both TurboQuantIndex::add and TurboQuantIndex::add_2d invoke first_invalid_coord before encoding. If validation fails in the standard add method (lines 65-70), the code panics with a descriptive message:
if let Some((vi, ci, v)) = first_invalid_coord(vectors, dim) {
panic!(
"invalid input value at vector {vi}, coord {ci}: {v} \
(must be finite and |value| < 1e16 to avoid f32 norm overflow)",
);
}
The add_2d variant (lines 70-77) returns a typed AddError::InvalidInputValue instead of panicking, supporting error propagation in production bindings.
Search Path Validation
The search and search_with_mask methods (lines 99-105, 442-450) apply identical validation to query vectors. Invalid queries trigger an immediate panic identifying the specific query index and coordinate:
if let Some((vi, ci, v)) = first_invalid_coord(queries, dim) {
panic!(
"invalid query value at query {vi}, coord {ci}: {v} \
(must be finite and |value| < 1e16 to avoid f32 overflow)",
);
}
Python Wrapper Error Handling
The Python binding in turbovec-python/src/lib.rs translates Rust panics into Python exceptions. When using the Python API, AddError::InvalidInputValue surfaces as a ValueError with detailed coordinate information, allowing graceful handling of malformed input data.
Practical Code Examples
Handling Validation Errors in Rust
When adding vectors to a TurboVec index, non-finite values trigger immediate panics:
use turbovec::TurboQuantIndex;
fn main() -> Result<(), turbovec::ConstructError> {
let mut idx = TurboQuantIndex::new(1536, 4)?;
// Valid vectors
let vectors = vec![0.1_f32; 1536 * 10];
idx.add(&vectors); // Success
// Invalid vector with NaN
let mut bad = vectors.clone();
bad[123] = f32::NAN;
idx.add(&bad); // Panics: "invalid input value at vector 0, coord 123: NaN..."
Ok(())
}
Graceful Error Handling in Python
The Python wrapper converts validation failures into catchable exceptions:
import numpy as np
import turbovec
idx = turbovec.TurboQuantIndex.new_lazy(bit_width=4)
vectors = np.random.rand(10, 1536).astype(np.float32)
# Valid addition
idx.add(vectors)
# Invalid value insertion
vectors[0, 0] = np.inf
try:
idx.add(vectors)
except turbovec.AddError as e:
print(e) # "InvalidInputValue: vector 0, coord 0, value inf"
Query Validation
Search queries undergo identical validation:
let queries = vec![0.0_f32; 1536 * 2];
let results = idx.search(&queries, 5); // Success
let mut bad_queries = queries.clone();
bad_queries[10] = f32::INFINITY;
idx.search(&bad_queries, 5); // Panics with coordinate details
Summary
- TurboVec validates every input vector and query through
first_invalid_coordinturbovec/src/lib.rsbefore SIMD processing. - NaN and Inf values poison the per-vector scale storage, creating unreachable or dominant index entries.
- The 1e16 magnitude limit prevents
f32overflow during norm calculations that would corrupt similarity scores. - The Rust API panics on invalid input with detailed coordinates, while
add_2dreturnsAddError::InvalidInputValuefor typed error handling. - Python bindings translate validation failures into
ValueErrorexceptions with specific vector and coordinate indices.
Frequently Asked Questions
What happens if I accidentally add a NaN value to a TurboVec index?
TurboVec immediately panics with a descriptive message indicating the exact vector and coordinate index containing the NaN. In the Python bindings, this manifests as a ValueError. The vector is not added to the index, preventing the NaN from poisoning the scale storage and corrupting future search results.
Why does TurboVec reject values with magnitude larger than 1e16?
Values ≥ 1e16 cause the sum-of-squares norm calculation to overflow to +Inf during the encoding process. This results in an infinite scale factor that dominates every top-k query, causing the corrupted vector to incorrectly win every similarity search. The 1e16 threshold provides a safety margin well above typical embedding magnitudes while preventing overflow.
Does TurboVec return an error code or panic on invalid input?
The standard add and search methods panic on invalid input to surface programmer errors immediately. However, the add_2d method returns a typed AddError::InvalidInputValue result, allowing Rust applications to handle validation failures gracefully without unwinding the stack. The Python wrapper exclusively uses exceptions derived from this error type.
How does the Python wrapper handle invalid vector values?
The Python binding layer catches Rust validation errors and converts them into Python ValueError exceptions (or specific turbovec.AddError subclasses) that include the offending vector index, coordinate index, and value. This allows Python applications to sanitize input data and retry operations without crashing the interpreter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →