Structure of the .tv File Format: Binary Layout and Data Organization

The .tv file format stores quantized vector indices with a 14-byte header containing the magic string "TVPI" and a 9-byte core metadata block, followed by tightly packed quantization codes, per-vector scale factors, and an optional TQ+ calibration trailer, using little-endian byte order for all multi-byte fields except the big-endian magic.

The .tv (TurboVect Index) file is the on-disk representation of a TurboQuantIndex, the positional vector index used in the RyanCodrai/turbovec repository. According to the source code in turbovec/src/io.rs, this binary format encodes high-dimensional vectors using product quantization with configurable bit-widths, storing both the compressed codes and the reconstruction scales needed for approximate nearest neighbor search.

Binary Layout Overview

The .tv structure follows a strict byte sequence designed for efficient memory mapping and deserialization. The layout consists of a fixed 14-byte prefix (magic + version + core header), followed by variable-length data sections.

Magic String and Version Byte

The file begins with a 4-byte identifier that distinguishes TurboVect files from other binary formats:

  • Bytes 0–3: Magic string "TVPI" (big-endian ASCII)
  • Byte 4: Format version (u8), currently 3

Versions 1 and 2 remain supported for backward compatibility, though version 1 files lack the magic header entirely and trigger a rebuild error when loaded (as implemented in io.rs lines 70–78).

Core Header (9 Bytes)

Immediately following the version byte, the core header stores the index metadata as defined by the constant CORE_HEADER_SIZE = 9 in io.rs (line 76):

Byte Offset Field Type Description
0 bit_width u8 Quantization width per sub-dimension (2, 3, or 4 bits)
1–4 dim u32 (LE) Original vector dimensionality
5–8 n_vectors u32 (LE) Number of vectors stored in the index

All multi-byte fields in the core header use little-endian byte order.

Data Sections

After the 14-byte header, the file contains three distinct data regions whose sizes depend on the dim, n_vectors, and bit_width values specified in the header.

Packed Quantization Codes

The bulk of the file consists of tightly packed quantization codes. The total size in bytes is calculated as:

(dim / 8) * bit_width * n_vectors

As implemented in read_header_codes_scales (lines 71–74 of io.rs), these bytes contain the actual quantized vector data where each sub-dimension is encoded using bit_width bits. The packing is dense, with no padding between vectors.

Per-Vector Scales

Following the packed codes, the file stores one 32-bit IEEE 754 float (f32) per vector in little-endian format. This section occupies exactly 4 * n_vectors bytes. According to the source code in io.rs (lines 75–76), these scales represent the norms used during inner-product search to reconstruct approximate distances from the quantized codes.

TQ+ Calibration Trailer (Version 3)

Introduced in format version 3, the calibration trailer enables per-coordinate rescaling through TQ+ (TurboQuant Plus) post-processing. This section appears at the end of the file and is controlled by a count field:

  • n_calib (u32, 4 bytes): Specifies the calibration mode
    • 0: Identity calibration (no additional data, behaves like version 2)
    • dim: Full per-coordinate calibration follows

When n_calib > 0 (specifically when equal to dim as shown in lines 200–207 of io.rs), the file appends two consecutive arrays:

  1. Shift values: n_calib × 4 bytes of f32 (little-endian)
  2. Scale values: n_calib × 4 bytes of f32 (little-endian)

These arrays store the per-coordinate offsets and gains that refine quantization accuracy, as read in lines 208–212 of io.rs.

Version History and Compatibility

The Turbovec library maintains backward compatibility while rejecting obsolete formats:

  • Version 1: No magic string; files start directly with the 9-byte core header. The current loader explicitly rejects these files with an error message suggesting a rebuild (lines 70–78).
  • Version 2: Contains the "TVPI" magic and version byte, but lacks the TQ+ calibration trailer. The loader treats these as having empty calibration arrays (identity transform) via read_core_v2 (lines 37–40).
  • Version 3: Current format supporting the full calibration trailer, read via read_core_v3 (lines 44–60).

Working with .tv Files in Rust and Python

Rust: Serializing and Deserializing

The turbovec::io module provides the write and load functions for direct binary manipulation:

use turbovec::io::{write, load};

// Parameters
let bit_width = 4;
let dim = 1536;
let n_vectors = 1000;

// Prepare data
let packed_codes = vec![0u8; (dim / 8) * bit_width * n_vectors];
let scales = vec![1.0f32; n_vectors];
let tqplus_shift = vec![0.0f32; dim];  // Identity calibration
let tqplus_scale = vec![1.0f32; dim];

// Write version 3 format
write(
    "index.tv",
    bit_width,
    dim,
    n_vectors,
    &packed_codes,
    &scales,
    &tqplus_shift,
    &tqplus_scale,
).unwrap();

// Load round-trip
let (bw, d, n, codes, sc, shift, scale) = load("index.tv").unwrap();

Source: turbovec/src/io.rs, lines 42–61 (write) and 64–95 (load).

Python: High-Level Index Operations

The Python bindings expose the same format through the TurboQuantIndex class:

from turbovec import TurboQuantIndex
import numpy as np

# Create and populate index

idx = TurboQuantIndex(dim=1536, bit_width=4)
vectors = np.random.randn(1000, 1536).astype(np.float32)
idx.add(vectors)

# Persist to .tv format (version 3 with optional calibration)

idx.write("vectors.tv")

# Load preserving all metadata and calibration

loaded = TurboQuantIndex.load("vectors.tv")
assert loaded.dim == 1536

Source: Python bindings in turbovec-python/src/lib.rs and API documentation in docs/api.md.

Summary

  • The .tv file format begins with the big-endian magic "TVPI" followed by a version byte and a 9-byte little-endian core header specifying bit_width, dim, and n_vectors.
  • Quantized vector data occupies (dim / 8) * bit_width * n_vectors bytes immediately after the header.
  • Each vector has an associated 4-byte f32 scale factor stored consecutively after the packed codes.
  • Version 3 files may include a TQ+ calibration trailer with per-coordinate shift and scale arrays when n_calib equals the vector dimension.
  • The implementation in turbovec/src/io.rs handles versions 1–3 with explicit compatibility logic in read_core_v2 and read_core_v3.

Frequently Asked Questions

What byte order does the .tv file format use?

The .tv format uses big-endian byte order only for the 4-byte magic string "TVPI" at the start of the file. All other multi-byte fields—including the core header dimensions, vector counts, scale values, and calibration data—use little-endian byte order, as implemented in the serialization logic of turbovec/src/io.rs.

How is the size of the packed codes section calculated?

The packed codes section size equals (dim / 8) * bit_width * n_vectors bytes. This formula accounts for the product quantization scheme where each vector dimension is sub-quantized using bit_width bits (typically 2, 3, or 4), with eight sub-dimensions packed into each byte.

Can I load a version 1 .tv file in the current Turbovec release?

No. Version 1 files lack the "TVPI" magic header and the current loader in io.rs explicitly rejects them with an error directing users to rebuild the index. Versions 2 and 3 remain fully supported, with version 2 automatically treated as having identity (empty) calibration.

What happens when n_calib is set to zero in a version 3 file?

When the n_calib field (the 4-byte integer preceding the calibration arrays) equals zero, the loader interprets this as identity calibration, meaning no shift or scale arrays follow in the file. The index behaves identically to a version 2 file, using no per-coordinate rescaling during query operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →