Turbovec .tv and .tvim File Format: Binary Layout Specification
Turbovec stores quantized vector indexes in two versioned binary formats—.tv files contain packed codes and TQ+ calibration data under the TVPI magic header, while .tvim files append a slot-to-id lookup table under the TVIM magic header—both currently at version 3.
Turbovec is a high-performance Rust library for approximate nearest neighbor search developed by RyanCodrai. The library persists indexes to disk using proprietary .tv and .tvim extensions defined in turbovec/src/io.rs. Understanding the binary structure of these Turbovec file formats is essential for debugging storage issues, implementing cross-language loaders, and managing version compatibility.
Binary Format Overview
Both .tv and .tvim files share an identical core payload but differ in their magic headers and the presence of an ID mapping table.
| Extension | Rust Type | Magic Bytes | Version | Payload Difference |
|---|---|---|---|---|
.tv |
TurboQuantIndex |
b"TVPI" |
3 | Core payload only |
.tvim |
IdMapIndex |
b"TVIM" |
3 | Core payload + slot-to-id table |
The format begins with a 4-byte magic string followed by a 1-byte version number. According to the source code in turbovec/src/io.rs, the current specification is version 3, which includes TQ+ per-coordinate calibration data.
Core Payload Structure
After the 5-byte header (magic + version), the core payload is written sequentially by the write function:
- Header – 1 byte
bit_width, 4 bytesdim, 4 bytesn_vectors - Packed codes –
(dim / 8) * bit_width * n_vectorsbytes containing quantized vector data - Scales –
n_vectors× 4-bytef32values (one per vector) - TQ+ calibration – Optional per-coordinate rescaling:
- 4 bytes
n_calib(0indicates identity calibration, otherwise equalsdim) n_calib× 4-bytef32shift valuesn_calib× 4-bytef32scale values
- 4 bytes
This layout is implemented in turbovec/src/io.rs within the write function (lines 42–53) and consumed by load (lines 64–96).
The .tvim Extension: ID Mapping
.tvim files extend the core payload with a slot-to-id table that maps each physical vector slot to its original identifier. After the core payload, the file appends:
- Slot-to-id table –
n_vectors× 8-byteu64values (little-endian)
This trailing Vec<u64> allows the IdMapIndex to reconstruct the relationship between internal storage indices and external IDs. The write_id_map function (lines 98–110) handles this serialization, while load_id_map (lines 33–45) validates the TVIM magic before parsing the table.
Reading and Writing .tv Files
Writing a .tv Index
Use the io::write function to persist a TurboQuantIndex without ID mapping:
use turbovec::io;
// Prepare quantized data
let bit_width = 4;
let dim = 128;
let n_vectors = 10_000;
let packed_codes: Vec<u8> = /* quantized vector data */;
let scales: Vec<f32> = /* per-vector scales */;
let tqplus_shift: Vec<f32> = /* dim-length for v3, empty for v2 */;
let tqplus_scale: Vec<f32> = /* same length as shift */;
io::write(
"index.tv",
bit_width,
dim,
n_vectors,
&packed_codes,
&scales,
&tqplus_shift,
&tqplus_scale,
)?;
This writes the TVPI magic, version 3, and the core payload to turbovec/src/io.rs.
Loading a .tv Index
The load function automatically detects version 2 versus 3 and returns the decomposition:
use turbovec::io;
let (bit_width, dim, n_vectors, packed_codes, scales, tqplus_shift, tqplus_scale) =
io::load("index.tv")?;
Reading and Writing .tvim Files
Writing a .tvim Index
For indexes requiring ID preservation, use write_id_map:
use turbovec::io;
let slot_to_id: Vec<u64> = vec![100, 101, 102, /* ... */];
io::write_id_map(
"index.tvim",
bit_width,
dim,
n_vectors,
&packed_codes,
&scales,
&tqplus_shift,
&tqplus_scale,
&slot_to_id,
)?;
Loading a .tvim Index
The load_id_map function validates the TVIM magic, reads the core payload, then parses the trailing ID table:
use turbovec::io;
let (
bit_width,
dim,
n_vectors,
packed_codes,
scales,
tqplus_shift,
tqplus_scale,
slot_to_id,
) = io::load_id_map("index.tvim")?;
Version Compatibility and Migration
The Turbovec file format supports backward-compatible reading but not forward compatibility:
- Version 3 (current) – Includes TQ+ calibration data; written by default
- Version 2 – Lacks TQ+ data; readable by current versions (calibration treated as identity)
- Version 1 – Pre-0.4.4 files without magic headers; incompatible and will return an explicit error requiring index rebuild (handled in
loadat lines 70–88)
When loading, the presence of the magic header is validated first. If missing, the loader returns an error indicating the file must be rebuilt.
Summary
.tvfiles use theTVPImagic header and store only the core quantized payload (bit-width, dimensions, packed codes, scales, and optional TQ+ calibration)..tvimfiles use theTVIMmagic header and append aVec<u64>slot-to-id table to the core payload.- Both formats are currently at version 3, defined in
turbovec/src/io.rs. - Version 2 files are readable but lack calibration; version 1 files are incompatible and lack magic headers.
- Use
io::writeandio::loadfor.tvfiles; useio::write_id_mapandio::load_id_mapfor.tvimfiles.
Frequently Asked Questions
What is the difference between .tv and .tvim files?
.tv files store the minimal data needed for similarity search: quantized codes, per-vector scales, and calibration parameters. .tvim files contain the same data plus a trailing table that maps each internal slot index to its original u64 identifier, enabling the IdMapIndex to return original IDs during queries.
Can Turbovec read older versions of .tv files?
Yes. The load and load_id_map functions in turbovec/src/io.rs transparently read version 2 files by detecting the version byte and treating missing TQ+ calibration as identity transforms. However, version 2 files cannot store per-coordinate calibration data.
What happens if I try to load a v1 .tv file?
Loading a version 1 file results in an error. These pre-0.4.4 files lack the 4-byte magic header (TVPI or TVIM), causing the loader to reject them with a message indicating the index must be rebuilt from the source vectors.
How is the slot-to-id table used in .tvim files?
The slot-to-id table is a Vec<u64> written after the core payload in little-endian format. When IdMapIndex loads the file, it uses this table to translate internal vector positions (0 to n_vectors-1) back to their original identifiers, which is critical when the index must reference external database keys or document IDs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →