Turbovec Index File Formats Explained: `.tv` and `.tvim` Binary Structure

TLDR: Turbovec stores vector indexes in two binary formats — .tv (version 6) for compact, ID-less snapshots and .tvim (version 7) for persistent indexes with external ID maps and optional TQ+ calibration data — both sharing a common header layout defined in the turbovec/src/io.rs source file.

The Turbovec repository (RyanCodrai/turbovec) is a high-performance vector similarity search library written in Rust with Python bindings. Its index persistence layer uses two distinct on-disk formats that differ by version, magic bytes, and metadata sections. Understanding the turbovec index file format is essential if you plan to read, write, or memory-map indexes directly, or if you need to debug interop issues between the v6 and v7 code paths. This guide breaks down the byte-level structure, version differences, and practical code examples for both .tv and .tvim files.

Format Overview: v6 vs v7

Turbovec writes indexes in one of two binary layouts depending on which data structure you're persisting:

Format Extension Version Primary Use
TurboQuant v6 .tv 6 Compact, read-only snapshots of a TurboQuantIndex (no external IDs or calibration data).
TurboQuant v7 .tvim 7 Persistent indexes that also contain an formula ID map and optional TQ+ calibration information (used by IdMapIndex).

The exposed write and load methods in the library auto-detect which version to produce or consume. The top-level serialization logic resides in [src/io.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/io.rs), while the v7 extensions are isolated in [src/io_v7.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/io_v7.rs).

Common Header Structure for .tv and .tvim

Both file types begin with a binary header followed by a packed payload. The first 32 bytes are shared between the two versions, containing the magic number, version identifier, vector dimensionality, bit width, and total vector count.

Here is the exact block layout:

Offset (bytes) Size Field Meaning
0 4 MAGIC ASCII string "TVIM" for v7, "TVV\0" (or "TV" followed by zero) for v6.
4 2 VERSION 0x0006 for v6, 0x0007 for v7 (little-endian).
6 2 RESERVED Currently unused; set to zero.
8 4 DIM Vector dimensionality (u32). Must be a positive multiple of 8 and ≤ MAX_DIM (16,384).
12 1 BIT_WIDTH Bits per coordinate (valid valuses: 2, 3, or 4).
13 6 RESERVED2 Padding to 16-byte alignment (zeroed).
16 8 N_VECTORS Total number of stored vectors (u64).
24 8 HEADER_SIZE Total size of this header block in bytes (u64). Fixed at 32 for v6; v7 grows this because of optional metadata fields.

For version 7, the header continues immediately after HEADER_SIZE:

Offset Size Field Meaning
32 1 HAS_ID_MAP 0 = no ID map, 1 = ID map present.
33 1 HAS_CALIBRATION 0 = uncalibrated, 1 = calibrated (TQ+).
34 6 RESERVED3 Padding (zero).
40 8 × N_VECTORS ID_MAP (optional) Array of u64 external IDs, one per vector, only if HAS_ID_MAP == 1.
variable – Calibration data (optional) If HAS_CALIBRATION == 1, two contiguous blocks of f32 values: SHIFT array (dim × f32) and SCALE array (dim × f32).

The header is always aligned to 16 bytes so the subsequent compressed payload can be efficiently formatted.

Payload Structure After the Header

Both .tv and .tvim store the actual quantized vectors in a blocked, bit-packed format. This is the region that the SIMD search kernels consume directly.

The payload consists of three main components:

  • Blocked cache (a sequence of blocks, each covering 32 rows). For every block, vector coordinates are packed into a bit-plane representation. The number of blocks is calculated as n_blocks = ceiling(N_VECTORS / 32). The exact geometry is derived by pack::blocked_geometry in [src/pack.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/pack.rs).
  • Scales array. One f32 per vector representing the length-renormalization factor. These are stored consecutively immediately after the blocked cache.
  • Additional v7-only sections (if present): the ID map is written before the blocked cache so it can be memory-mapped without loading the entire file. The calibration blocks (shift and scale) follow the ID map, also before the cache when HAS_CALIBRATION is true.

Key Differences Between .tv (v6) and .tvim (v7)

Feature .tv (v6) .tvim (v7)
Magic bytes "TVV\0" "TVIM"
Header size Fixed 32 bytes (no extra fields) Variable; includes optional ID map and calibration
ID map Not stored – external IDs are absent Optional u64 array per vector
Calibration (TQ+) Not represented – index is always uncalibrated Optional SHIFT/SCALE arrays (presence is marked by HAS_CALIBRATION)
Compatibility Only TurboQuantIndex Used by IdMapIndex (which wraps a TurboQuantIndex)

When you call TurboQuantIndex::write, the library always emits the v6 (.tv) snapshot. When IdMapIndex::write is called (from either Rust or Python), the v7 (.tvim) layout is used, preserving the external ID map and any calibration applied before saving.

Practical Code Examples

Writing a .tv file in Rust

use turbovec::TurboQuantIndex;

// Create a 1536-dim index, 4 bits per coordinate.
let mut idx = TurboQuantIndex::new(1536, 4).unwrap();
idx.add(&vectors);

// Serialize to the v6 format.
idx.write("my_index.tv").unwrap(); // <-- generates `.tv`

The write method on TurboQuantIndex delegates to io::write_v6, as defined in [src/lib.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/lib.rs).

Writing an .tvim file in Python

from turbovec import IdMapIndex
import numpy as np

# Build an index that tracks external IDs.

idx = IdMapIndex(dim=1536, bit_width=4)
idx.add_with_ids(vectors, np.arange(len(vectors), dtype=np.uint64))

# Persist to v7 (includes IDs + calibration).

idx.write("my_index.tvim")

The Python binding forwards to IdMapIndex::write, which eventually calls io_v7::write_v7 loaded from src/io_v7.rs.

Loading a .tvim file again (Rust)

use turbovec::IdMapIndex;

// Automatically restores IDs and any calibration.
let idx = IdMapIndex::load("my_index.tvim").unwrap();

load reads the first 32 bytes, checks the version field, and then constructs either a plain TurboQuantIndex or an IdMapIndex, depending on the magic number and HAS_ID_MAP/HAS_CALIBRATION flags (see src/io.rs).

Quick Reference for write / load internals

Constant / function Location Purpose
MAGIC_V6 / MAGIC_V7 io.rs Detection of file version.
write_v6 io.rs Serialises a TurboQuantIndex to the .tv layout.
write_v7 io_v7.rs Serialises an IdMapIndex (plus optional ID map and calibration) to the .tvim layout.
read_header io.rs Parses the first 32 bytes, validates magic, determines version.
read_v6 / read_v7 io.rs / io_v7.rs Parse the rest of the file according to the version-specific schema.
blocked_geometry pack.rs Computes block count, byte-group count, and total payload size from n_vectors, bit_width, and dim.
repack / repack_block_range pack.rs Converts packed-row layout to the blocked cache format (needed when loading a v6 file and then mutating data).

Summary

  • .tv files are version 6 binary snapshots containing only vector codes, block traps, and a fixed header.
  • .tvim files are version 7 and add the magic "TVIM", flags for an ID map and calibration, and optional u64 ID arrays plus SHIFT/SCALE blocks.
  • Both formats share a common 32-byte header with DIM, BIT_WIDTH, N_VECTORS, and HEADER_SIZE fields.
  • The payload after the header uses a blocked, bit-packed cache format together with per-vector scales.
  • Writing from the low-level TurboQuantIndex produces .tv; writing from IdMapIndex produces .tvim and restores the important additional metadata.

Frequently Asked Questions

How do I know whether a file on disk is .tv or .tvim?

Check the first four bytes. If they are "TVIM" the file is a v7 .tvim; if they begin with "TV" followed by a zero byte (or are simply "tv" in ASCII), it's a v6 .tv. The read_header function in io.rs performs exactly this check when loading.

Can I convert a .tv file into a .tvim file?

Yes, but you need to construct an IdMapIndex from the data first. Load the .tv using TurboQuantIndex::load, then call IdMapIndex::from_index providing your external IDs, and finally write out to the .tvim destination.

Why do my .tv files have no ID tracks?

Version 6 explicitly stores no external separator IDs. If your application requires a stable mapping between vectors and external identifiers (e.g., document or graph node IDs), you should use v7 (.tvim) from the start or convert to v7 during the write phase.

Where in the source can I find the full binary layout structs?

The primary struct definitions are in src/io.rs (the GenericIndexHeader equivalent and v6 writer) and src/io_v7.rs (the v7 writer and reader). The pack.rs module documents the exact byte-offset calculations via blocked_geometry.


The details above are derived from the Turbovec source files io.rs, io_v7.rs, and pack.rs. For the complete latest layout, always check the official repository documentation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →