Turbovec .tv and .tvim File Format: Binary Layout Specification

Turbovec stores quantized vector indexes in two versioned binary formats—.tv files contain packed codes and TQ+ calibration data under the TVPI magic header, while .tvim files append a slot-to-id lookup table under the TVIM magic header—both currently at version 3.

Turbovec is a high-performance Rust library for approximate nearest neighbor search developed by RyanCodrai. The library persists indexes to disk using proprietary .tv and .tvim extensions defined in turbovec/src/io.rs. Understanding the binary structure of these Turbovec file formats is essential for debugging storage issues, implementing cross-language loaders, and managing version compatibility.

Binary Format Overview

Both .tv and .tvim files share an identical core payload but differ in their magic headers and the presence of an ID mapping table.

Extension Rust Type Magic Bytes Version Payload Difference
.tv TurboQuantIndex b"TVPI" 3 Core payload only
.tvim IdMapIndex b"TVIM" 3 Core payload + slot-to-id table

The format begins with a 4-byte magic string followed by a 1-byte version number. According to the source code in turbovec/src/io.rs, the current specification is version 3, which includes TQ+ per-coordinate calibration data.

Core Payload Structure

After the 5-byte header (magic + version), the core payload is written sequentially by the write function:

  1. Header – 1 byte bit_width, 4 bytes dim, 4 bytes n_vectors
  2. Packed codes – (dim / 8) * bit_width * n_vectors bytes containing quantized vector data
  3. Scales – n_vectors × 4-byte f32 values (one per vector)
  4. TQ+ calibration – Optional per-coordinate rescaling:
    • 4 bytes n_calib (0 indicates identity calibration, otherwise equals dim)
    • n_calib × 4-byte f32 shift values
    • n_calib × 4-byte f32 scale values

This layout is implemented in turbovec/src/io.rs within the write function (lines 42–53) and consumed by load (lines 64–96).

The .tvim Extension: ID Mapping

.tvim files extend the core payload with a slot-to-id table that maps each physical vector slot to its original identifier. After the core payload, the file appends:

  1. Slot-to-id table – n_vectors × 8-byte u64 values (little-endian)

This trailing Vec<u64> allows the IdMapIndex to reconstruct the relationship between internal storage indices and external IDs. The write_id_map function (lines 98–110) handles this serialization, while load_id_map (lines 33–45) validates the TVIM magic before parsing the table.

Reading and Writing .tv Files

Writing a .tv Index

Use the io::write function to persist a TurboQuantIndex without ID mapping:

use turbovec::io;

// Prepare quantized data
let bit_width = 4;
let dim = 128;
let n_vectors = 10_000;
let packed_codes: Vec<u8> = /* quantized vector data */;
let scales: Vec<f32> = /* per-vector scales */;
let tqplus_shift: Vec<f32> = /* dim-length for v3, empty for v2 */;
let tqplus_scale: Vec<f32> = /* same length as shift */;

io::write(
    "index.tv",
    bit_width,
    dim,
    n_vectors,
    &packed_codes,
    &scales,
    &tqplus_shift,
    &tqplus_scale,
)?;

This writes the TVPI magic, version 3, and the core payload to turbovec/src/io.rs.

Loading a .tv Index

The load function automatically detects version 2 versus 3 and returns the decomposition:

use turbovec::io;

let (bit_width, dim, n_vectors, packed_codes, scales, tqplus_shift, tqplus_scale) =
    io::load("index.tv")?;

Reading and Writing .tvim Files

Writing a .tvim Index

For indexes requiring ID preservation, use write_id_map:

use turbovec::io;

let slot_to_id: Vec<u64> = vec![100, 101, 102, /* ... */];

io::write_id_map(
    "index.tvim",
    bit_width,
    dim,
    n_vectors,
    &packed_codes,
    &scales,
    &tqplus_shift,
    &tqplus_scale,
    &slot_to_id,
)?;

Loading a .tvim Index

The load_id_map function validates the TVIM magic, reads the core payload, then parses the trailing ID table:

use turbovec::io;

let (
    bit_width,
    dim,
    n_vectors,
    packed_codes,
    scales,
    tqplus_shift,
    tqplus_scale,
    slot_to_id,
) = io::load_id_map("index.tvim")?;

Version Compatibility and Migration

The Turbovec file format supports backward-compatible reading but not forward compatibility:

  • Version 3 (current) – Includes TQ+ calibration data; written by default
  • Version 2 – Lacks TQ+ data; readable by current versions (calibration treated as identity)
  • Version 1 – Pre-0.4.4 files without magic headers; incompatible and will return an explicit error requiring index rebuild (handled in load at lines 70–88)

When loading, the presence of the magic header is validated first. If missing, the loader returns an error indicating the file must be rebuilt.

Summary

  • .tv files use the TVPI magic header and store only the core quantized payload (bit-width, dimensions, packed codes, scales, and optional TQ+ calibration).
  • .tvim files use the TVIM magic header and append a Vec<u64> slot-to-id table to the core payload.
  • Both formats are currently at version 3, defined in turbovec/src/io.rs.
  • Version 2 files are readable but lack calibration; version 1 files are incompatible and lack magic headers.
  • Use io::write and io::load for .tv files; use io::write_id_map and io::load_id_map for .tvim files.

Frequently Asked Questions

What is the difference between .tv and .tvim files?

.tv files store the minimal data needed for similarity search: quantized codes, per-vector scales, and calibration parameters. .tvim files contain the same data plus a trailing table that maps each internal slot index to its original u64 identifier, enabling the IdMapIndex to return original IDs during queries.

Can Turbovec read older versions of .tv files?

Yes. The load and load_id_map functions in turbovec/src/io.rs transparently read version 2 files by detecting the version byte and treating missing TQ+ calibration as identity transforms. However, version 2 files cannot store per-coordinate calibration data.

What happens if I try to load a v1 .tv file?

Loading a version 1 file results in an error. These pre-0.4.4 files lack the 4-byte magic header (TVPI or TVIM), causing the loader to reject them with a message indicating the index must be rebuilt from the source vectors.

How is the slot-to-id table used in .tvim files?

The slot-to-id table is a Vec<u64> written after the core payload in little-endian format. When IdMapIndex loads the file, it uses this table to translate internal vector positions (0 to n_vectors-1) back to their original identifiers, which is critical when the index must reference external database keys or document IDs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →