# Security Considerations for Loading Untrusted .tv Files in Turbovec

> Learn how Turbovec secures untrusted .tv files by validating every field before memory allocation, preventing DoS and memory corruption attacks. Protect your system now.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: security-best-practices
- Published: 2026-06-16

---

**Turbovec validates every field in binary `.tv` and `.tvim` files—from magic bytes to dimensionality checks—before allocating memory, preventing denial-of-service and memory corruption attacks.**

When handling vector indexes from untrusted sources—whether user uploads, network streams, or third-party pipelines—RyanCodrai/turbovec implements a defense-in-depth strategy that scrutinizes every byte before any computation begins. The loader in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) performs rigorous header validation, overflow-checked arithmetic, and incremental streaming to ensure malicious inputs cannot trigger out-of-memory conditions or undefined behavior. These protections are continuously exercised by the security regression suite in [`turbovec-python/tests/test_security.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/tests/test_security.py).

## Validation Layers in the Turbovec Loader

The core validation logic resides in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs), where the `load` and `load_id_map` functions implement a sequence of defensive checks. Each validation step targets specific attack vectors while maintaining compatibility with legitimate version 2 and version 3 file formats.

### Magic-Byte and Version Verification

Before interpreting any header fields, the loader verifies that the file begins with the **magic bytes** `"TVPI"` (for positional indexes) or `"TVIM"` (for ID-mapped indexes). It then checks that the version byte is either `2` or `3`. This prevents accidental loading of arbitrary binary blobs and rejects legacy version-1 files that lack proper magic headers.

As implemented in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) (lines 70–92), this check executes immediately after the initial read, ensuring that malformed or truncated files fail fast without allocating resources.

### Bit-Width and Dimensionality Constraints

The loader enforces strict constraints on quantization parameters to prevent divide-by-zero and out-of-bounds memory access:

- **`bit_width`** must be `2`, `3`, or `4`. Any other value triggers an immediate error before the packing layer (`pack::repack`) is invoked.
- **`dim`** (dimensionality) must be a multiple of 8, must be non-zero unless `n_vectors` is also zero (indicating a lazy-index sentinel), and cannot exceed `MAX_DIM` (65,536).

These checks in `read_header_codes_scales` (lines 84–103) block headers that would otherwise allocate gigantic rotation matrices or cause panics due to zero-column layouts.

### Overflow-Safe Size Calculations

To prevent integer overflow attacks that could allocate gigabytes on 32-bit systems, the loader uses **checked arithmetic** (`checked_mul`) for all size calculations:

- The expected size of the packed-code block is computed using `checked_mul` before allocation.
- The `slot_to_id` table size (`n_vectors * 8`) is validated to fit within `usize` limits.

These safeguards in `read_header_codes_scales` (lines 113–120) and `load_id_map` (lines 64–69) defend against malicious headers that claim billions of vectors to induce out-of-memory conditions.

### Incremental Reading with Bounded Streams

Rather than pre-allocating based on header claims, the loader uses **`read_exact_vec`** with `take(n)` to stream exactly `n` bytes. As implemented in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) (lines 126–135), this approach rejects truncated files immediately, ensuring that a tiny file claiming massive vector counts fails quickly rather than OOM-ing during pre-allocation.

### TQ+ Calibration Sanity Checks

For version 3 files supporting TQ+ (TurboQuant Plus), the loader validates the trailer field **`n_calib`**. This value must be either `0` or exactly equal to `dim`. The `read_core_v3` function (lines 55–60) enforces this constraint to prevent malformed calibration data from corrupting distance calculations.

### ID-Map Specific Protections (.tvim Files)

The `.tvim` format extends `.tv` with auxiliary side-tables mapping slots to IDs. In `load_id_map` (lines 62–74), the loader mirrors all `.tv` safety guarantees while adding specific validation for the `slot_to_id` table size. This ensures that the ID-map payload cannot be exploited to perform out-of-bounds reads on the auxiliary metadata.

## Security Regression Testing

The Python bindings expose these same hardening guarantees. The test suite in [`turbovec-python/tests/test_security.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/tests/test_security.py) deliberately crafts malformed files to verify that `TurboQuantIndex.load` and `IdMapIndex.load` raise `ValueError` or `OSError` instead of panicking or silently misbehaving.

The following example demonstrates how to construct malicious files for testing:

```python
import struct
import pathlib
import pytest
from turbovec import TurboQuantIndex

def craft_bad_tv(path, bit_width, dim, n_vectors, n_scales=None, codes=b"", n_calib=0):
    """Write a v3 .tv file with attacker-controlled header fields."""
    if n_scales is None:
        n_scales = n_vectors
    with open(path, "wb") as f:
        f.write(b"TVPI")                     # magic

        f.write(bytes([3]))                  # version 3

        f.write(bytes([bit_width & 0xFF]))   # bit_width

        f.write(struct.pack("<I", dim))      # dim

        f.write(struct.pack("<I", n_vectors))# n_vectors

        f.write(codes)                       # packed codes

        f.write(struct.pack("<f", 1.0) * n_scales)  # per-vector scales

        f.write(struct.pack("<I", n_calib))  # TQ+ calibration size

# Example 1 – reject an illegal bit_width

tmp = pathlib.Path("/tmp/bad_bitwidth.tv")
craft_bad_tv(tmp, bit_width=0, dim=8, n_vectors=1, codes=b"\x00"*8)
with pytest.raises((ValueError, OSError)):
    TurboQuantIndex.load(str(tmp))

# Example 2 – reject a file that declares billions of vectors

tmp = pathlib.Path("/tmp/huge.tv")
craft_bad_tv(tmp, bit_width=2, dim=8, n_vectors=0xFFFFFFFF, n_scales=0)
with pytest.raises((ValueError, OSError)):
    TurboQuantIndex.load(str(tmp))

```

These tests validate that the Rust core correctly propagates errors across the FFI boundary, maintaining security invariants for Python consumers.

## Summary

- **Magic-byte validation** in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) rejects non-Turbovec files before parsing.
- **Bit-width constraints** (2, 3, or 4) and **dimensionality checks** (multiple of 8, max 65,536) prevent memory corruption.
- **Overflow-safe arithmetic** using `checked_mul` blocks integer overflow attacks on 32-bit and 64-bit platforms.
- **Incremental streaming** via `read_exact_vec` prevents OOM attacks from truncated files claiming massive vector counts.
- **TQ+ calibration validation** ensures `n_calib` matches `dim` or zero.
- **ID-map hardening** in `load_id_map` protects `.tvim` side-tables from out-of-bounds access.
- **Continuous validation** via [`test_security.py`](https://github.com/RyanCodrai/turbovec/blob/main/test_security.py) ensures these defenses remain effective across releases.

## Frequently Asked Questions

### What file extensions does turbovec use for its indexes?

Turbovec uses **`.tv`** for positional indexes and **`.tvim`** for indexes that include an ID-map side-table. These binary formats store quantized vector codes and optional metadata, with `.tvim` files containing additional slot-to-ID mappings for document retrieval.

### How does turbovec prevent integer overflow when loading large files?

The loader uses **checked arithmetic** (`checked_mul`) in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) to compute buffer sizes before allocation. For example, when calculating the packed-code block size or the `slot_to_id` table dimensions, any overflow results in an immediate error rather than wrapping to a small value that could cause heap corruption.

### Can turbovec safely load user-uploaded vector indexes?

Yes, provided you use the standard `TurboQuantIndex.load()` or `IdMapIndex.load()` APIs. These methods implement defense-in-depth validation that rejects malformed headers, invalid quantization parameters, and impossible dimensionality claims before allocating memory. However, as with any file-processing system, you should apply additional sandboxing and resource limits appropriate for your threat model.

### What error types are raised when loading a corrupted .tv file?

The Python bindings raise either **`ValueError`** (for validation failures like invalid bit-width or dimensionality mismatches) or **`OSError`** (for I/O failures such as truncated files or read errors). These exceptions originate from the Rust core in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) and propagate through the Python FFI without crashing the interpreter.