# Turbovec .tv and .tvim File Format: Binary Layout Specification

> Explore the Turbovec .tv and .tvim file format binary layout. Understand how Turbovec stores quantized vector indexes with packed codes, TQ+ calibration, and slot-to-id lookup.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: api-reference
- Published: 2026-06-10

---

**Turbovec stores quantized vector indexes in two versioned binary formats—`.tv` files contain packed codes and TQ+ calibration data under the `TVPI` magic header, while `.tvim` files append a slot-to-id lookup table under the `TVIM` magic header—both currently at version 3.**

Turbovec is a high-performance Rust library for approximate nearest neighbor search developed by RyanCodrai. The library persists indexes to disk using proprietary `.tv` and `.tvim` extensions defined in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs). Understanding the binary structure of these Turbovec file formats is essential for debugging storage issues, implementing cross-language loaders, and managing version compatibility.

## Binary Format Overview

Both `.tv` and `.tvim` files share an identical **core payload** but differ in their magic headers and the presence of an ID mapping table.

| Extension | Rust Type | Magic Bytes | Version | Payload Difference |
|-----------|-----------|-------------|---------|-------------------|
| `.tv` | `TurboQuantIndex` | `b"TVPI"` | 3 | Core payload only |
| `.tvim` | `IdMapIndex` | `b"TVIM"` | 3 | Core payload + slot-to-id table |

The format begins with a **4-byte magic string** followed by a **1-byte version number**. According to the source code in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs), the current specification is **version 3**, which includes TQ+ per-coordinate calibration data.

## Core Payload Structure

After the 5-byte header (magic + version), the core payload is written sequentially by the `write` function:

1. **Header** – 1 byte `bit_width`, 4 bytes `dim`, 4 bytes `n_vectors`
2. **Packed codes** – `(dim / 8) * bit_width * n_vectors` bytes containing quantized vector data
3. **Scales** – `n_vectors` × 4-byte `f32` values (one per vector)
4. **TQ+ calibration** – Optional per-coordinate rescaling:
   - 4 bytes `n_calib` (`0` indicates identity calibration, otherwise equals `dim`)
   - `n_calib` × 4-byte `f32` shift values
   - `n_calib` × 4-byte `f32` scale values

This layout is implemented in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) within the `write` function (lines 42–53) and consumed by `load` (lines 64–96).

## The .tvim Extension: ID Mapping

**`.tvim` files** extend the core payload with a **slot-to-id table** that maps each physical vector slot to its original identifier. After the core payload, the file appends:

5. **Slot-to-id table** – `n_vectors` × 8-byte `u64` values (little-endian)

This trailing `Vec<u64>` allows the `IdMapIndex` to reconstruct the relationship between internal storage indices and external IDs. The `write_id_map` function (lines 98–110) handles this serialization, while `load_id_map` (lines 33–45) validates the `TVIM` magic before parsing the table.

## Reading and Writing .tv Files

### Writing a .tv Index

Use the `io::write` function to persist a `TurboQuantIndex` without ID mapping:

```rust
use turbovec::io;

// Prepare quantized data
let bit_width = 4;
let dim = 128;
let n_vectors = 10_000;
let packed_codes: Vec<u8> = /* quantized vector data */;
let scales: Vec<f32> = /* per-vector scales */;
let tqplus_shift: Vec<f32> = /* dim-length for v3, empty for v2 */;
let tqplus_scale: Vec<f32> = /* same length as shift */;

io::write(
    "index.tv",
    bit_width,
    dim,
    n_vectors,
    &packed_codes,
    &scales,
    &tqplus_shift,
    &tqplus_scale,
)?;

```

This writes the `TVPI` magic, version `3`, and the core payload to [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).

### Loading a .tv Index

The `load` function automatically detects version 2 versus 3 and returns the decomposition:

```rust
use turbovec::io;

let (bit_width, dim, n_vectors, packed_codes, scales, tqplus_shift, tqplus_scale) =
    io::load("index.tv")?;

```

## Reading and Writing .tvim Files

### Writing a .tvim Index

For indexes requiring ID preservation, use `write_id_map`:

```rust
use turbovec::io;

let slot_to_id: Vec<u64> = vec![100, 101, 102, /* ... */];

io::write_id_map(
    "index.tvim",
    bit_width,
    dim,
    n_vectors,
    &packed_codes,
    &scales,
    &tqplus_shift,
    &tqplus_scale,
    &slot_to_id,
)?;

```

### Loading a .tvim Index

The `load_id_map` function validates the `TVIM` magic, reads the core payload, then parses the trailing ID table:

```rust
use turbovec::io;

let (
    bit_width,
    dim,
    n_vectors,
    packed_codes,
    scales,
    tqplus_shift,
    tqplus_scale,
    slot_to_id,
) = io::load_id_map("index.tvim")?;

```

## Version Compatibility and Migration

The Turbovec file format supports backward-compatible reading but not forward compatibility:

- **Version 3** (current) – Includes TQ+ calibration data; written by default
- **Version 2** – Lacks TQ+ data; readable by current versions (calibration treated as identity)
- **Version 1** – Pre-0.4.4 files without magic headers; **incompatible** and will return an explicit error requiring index rebuild (handled in `load` at lines 70–88)

When loading, the presence of the magic header is validated first. If missing, the loader returns an error indicating the file must be rebuilt.

## Summary

- **`.tv`** files use the `TVPI` magic header and store only the core quantized payload (bit-width, dimensions, packed codes, scales, and optional TQ+ calibration).
- **`.tvim`** files use the `TVIM` magic header and append a `Vec<u64>` slot-to-id table to the core payload.
- Both formats are currently at **version 3**, defined in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs).
- **Version 2** files are readable but lack calibration; **version 1** files are incompatible and lack magic headers.
- Use `io::write` and `io::load` for `.tv` files; use `io::write_id_map` and `io::load_id_map` for `.tvim` files.

## Frequently Asked Questions

### What is the difference between .tv and .tvim files?

**`.tv` files** store the minimal data needed for similarity search: quantized codes, per-vector scales, and calibration parameters. **`.tvim` files** contain the same data plus a trailing table that maps each internal slot index to its original `u64` identifier, enabling the `IdMapIndex` to return original IDs during queries.

### Can Turbovec read older versions of .tv files?

Yes. The `load` and `load_id_map` functions in [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) transparently read **version 2** files by detecting the version byte and treating missing TQ+ calibration as identity transforms. However, version 2 files cannot store per-coordinate calibration data.

### What happens if I try to load a v1 .tv file?

Loading a version 1 file results in an error. These pre-0.4.4 files lack the 4-byte magic header (`TVPI` or `TVIM`), causing the loader to reject them with a message indicating the index must be rebuilt from the source vectors.

### How is the slot-to-id table used in .tvim files?

The slot-to-id table is a `Vec<u64>` written after the core payload in little-endian format. When `IdMapIndex` loads the file, it uses this table to translate internal vector positions (0 to n_vectors-1) back to their original identifiers, which is critical when the index must reference external database keys or document IDs.