# Turbovec .tv and .tvim File Format Structure: Quantized Vector Storage Explained

> Explore the .tv and .tvim file format structure for quantized vector storage. Understand how Turbovec stores packed quantization codes, scales, and calibration data efficiently.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: internals
- Published: 2026-07-27

---

**The .tv and .tvim formats are versioned binary containers—defined in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs)—that store packed quantization codes, per-vector scales, optional TQ+ calibration data, and an ID-map side table in the case of .tvim, with all fields validated before allocation.**

Turbovec, as implemented in `RyanCodrai/turbovec`, is a Rust library for fast quantized vector search that persists indexes to disk using the `.tv` and `.tvim` file formats. The `.tv` file stores the core positional index, while the `.tvim` file extends it with a `slot_to_id` mapping for vector identification. Both formats share a common binary layout for the quantized payload and are governed by the serialization logic in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).

## Magic Number and Version Byte

Every file begins with a 4-byte magic identifier followed by a single version byte. The `.tv` format uses the magic sequence `TVPI`, and the `.tvim` format uses the magic sequence `TVIM`, as declared in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs). Both formats currently emit version 4 when writing, while the loader explicitly accepts versions 2, 3, and 4 to preserve backward-compatible reads.

## Core Header Layout in Version 4

After the magic and version, a **v4** header stores the metadata required to reconstruct the index. The exact layout is controlled by the constant `V4_HEADER_SIZE` in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).

- **1 byte:** `bit_width` — the quantization width, which must be 2, 3, or 4.
- **4 bytes:** `dim` — the vector dimensionality, stored as a little-endian `u32`.
- **8 bytes:** `n_vectors` — the total number of vectors, stored as a little-endian `u64`.
- **8 bytes:** rotation fingerprint hash — a `u64` hash that protects against rotation drift on load.
- **`N_PROBES` × 4 bytes:** fingerprint probes — stored as `f32` values.

This header upgrades prior versions by expanding `n_vectors` from `u32` to `u64` and adding the rotation fingerprint field for integrity verification.

## Shared Core Payload

Following the header, both `.tv` and `.tvim` files write the same core payload in three sequential blocks.

### Packed Quantization Codes

The first block contains the quantized vector data. It occupies exactly `bit_width × (dim / 8) × n_vectors` bytes, tightly packing the low-bit codes for every dimension of every vector. This byte stream is read directly into the engine’s distance-computation buffers.

### Per-Vector Scales

The second block stores one 32-bit float per vector, consuming `n_vectors × 4` bytes. These scales are applied during dequantization to rescale the packed integer codes back to their original approximate magnitudes.

### TQ+ Calibration Trailer

The final shared block is the optional TQ+ calibration trailer. It begins with a `n_calib` count stored as a 4-byte `u32`, which is either `0` or exactly equal to `dim`. When present, the trailer continues with:

- `tqplus_shift`: `n_calib × 4` bytes of `f32` shift values.
- `tqplus_scale`: `n_calib × 4` bytes of `f32` scale values.

The helper function `assert_tqplus_calibration` validates these arrays, while `read_tqplus_trailer` handles their deserialization in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).

## .tvim-Specific ID-Map Extension

The `.tvim` format appends an extra table after the core payload. This block is a **slot-to-ID map** consisting of `n_vectors` entries, each an 8-byte little-endian `u64`. The writer enforces that the length of this table exactly matches `n_vectors`, ensuring a one-to-one correspondence between quantized slots and external identifiers. When loading, `IdMapIndex::load_id_map` returns this vector alongside the core quantized data.

## Version Differences and Backward Compatibility

Turbovec maintains a small but meaningful version lineage inside [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).

1. **Version 2:** Uses a `u32` for `n_vectors` and omits the TQ+ calibration trailer entirely.
2. **Version 3:** Retains the v2 header but adds the TQ+ calibration trailer.
3. **Version 4:** Upgrades `n_vectors` to `u64`, introduces the rotation fingerprint hash, and keeps the TQ+ trailer.

The loader inspects the version byte and dispatches to the appropriate deserialization path, allowing modern Turbovec builds to read indexes produced by earlier releases.

## Validation and Safety Guards

Before any memory is allocated, Turbovec performs strict validation on every field in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).

- Header sanity checks ensure that `bit_width`, `dim`, and `n_vectors` fall within plausible ranges.
- Buffer sizes are computed using overflow-checked arithmetic to prevent out-of-memory crashes on malformed headers.
- Per-vector scales and every TQ+ calibration value must be finite; non-finite floats trigger an error and abort loading.

These checks prevent corrupted or malicious files from causing silent data corruption or memory exhaustion.

## Reading and Writing .tv and .tvim Files

The public API in [`turbovec/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs) exposes safe methods for persisting and restoring indexes.

```rust
use turbovec::TurboQuantIndex;
use std::path::Path;

// Write a positional-only .tv index (version 4)
let idx: TurboQuantIndex = /* ... */;
idx.write(Path::new("my_index.tv"))?;

```

To load a positional index later:

```rust
use turbovec::TurboQuantIndex;

let idx = TurboQuantIndex::load("my_index.tv")?;

```

For indexes that require an ID mapping, use the `.tvim` helpers:

```rust
use turbovec::IdMapIndex;
use std::path::Path;

// Write the index plus slot_to_id table
let id_map_idx: IdMapIndex = /* ... */;
id_map_idx.write_id_map(Path::new("my_index.tvim"))?;

```

Loading returns all runtime components explicitly:

```rust
use turbovec::IdMapIndex;

let (bit_width, dim, n_vectors, packed, scales, tq_shift, tq_scale, slot_to_id) =
    IdMapIndex::load_id_map("my_index.tvim")?;

```

## Summary

- The `.tv` and `.tvim` formats are little-endian binary containers defined in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).
- Both start with a 4-byte magic (`TVPI` or `TVIM`), a version byte, and a v4 header containing `bit_width`, `dim`, `n_vectors`, a rotation fingerprint, and probe values.
- After the header, the core payload stores packed codes, per-vector `f32` scales, and an optional TQ+ calibration trailer.
- The `.tvim` file appends a `slot_to_id` table of `u64` values to map internal slots to external IDs.
- Versions 2–4 are supported for reading, with v4 expanding `n_vectors` to 64 bits and adding drift detection.
- All inputs are validated for plausibility, arithmetic overflow, and finite floating-point values before allocation.

## Frequently Asked Questions

### What is the difference between .tv and .tvim?

The `.tv` file stores only the positional quantized index, including header, packed codes, scales, and optional TQ+ calibration. The `.tvim` file contains the identical core payload but appends a `slot_to_id` mapping table, making it suitable when each vector must be associated with an external identifier.

### Which Turbovec versions can read older file formats?

The loader in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) accepts versions 2, 3, and 4. Version 4 is emitted by default during writes, ensuring that new files benefit from 64-bit vector counts and rotation fingerprinting while legacy files remain readable.

### What does the rotation fingerprint in v4 prevent?

The rotation fingerprint is an 8-byte hash stored in the v4 header. When a file is loaded, Turbovec verifies this fingerprint against the current transform state to detect **rotation drift**, which would otherwise produce incorrect distance calculations if the quantization rotation matrix had changed.

### How does Turbovec validate a file before loading?

All fields are checked for semantic plausibility, buffer sizes are computed with overflow-checked arithmetic, and every floating-point scale or calibration parameter is verified to be finite. These guards run in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) before any heap allocation occurs, eliminating common vectors for malformed-file exploits.