Turbovec Index File Formats Explained: `.tv` and `.tvim` Binary Structure
TLDR: Turbovec stores vector indexes in two binary formats — .tv (version 6) for compact, ID-less snapshots and .tvim (version 7) for persistent indexes with external ID maps and optional TQ+ calibration data — both sharing a common header layout defined in the turbovec/src/io.rs source file.
The Turbovec repository (RyanCodrai/turbovec) is a high-performance vector similarity search library written in Rust with Python bindings. Its index persistence layer uses two distinct on-disk formats that differ by version, magic bytes, and metadata sections. Understanding the turbovec index file format is essential if you plan to read, write, or memory-map indexes directly, or if you need to debug interop issues between the v6 and v7 code paths. This guide breaks down the byte-level structure, version differences, and practical code examples for both .tv and .tvim files.
Format Overview: v6 vs v7
Turbovec writes indexes in one of two binary layouts depending on which data structure you're persisting:
| Format | Extension | Version | Primary Use |
|---|---|---|---|
| TurboQuant v6 | .tv |
6 | Compact, read-only snapshots of a TurboQuantIndex (no external IDs or calibration data). |
| TurboQuant v7 | .tvim |
7 | Persistent indexes that also contain an formula ID map and optional TQ+ calibration information (used by IdMapIndex). |
The exposed write and load methods in the library auto-detect which version to produce or consume. The top-level serialization logic resides in [src/io.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/io.rs), while the v7 extensions are isolated in [src/io_v7.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/io_v7.rs).
Common Header Structure for .tv and .tvim
Both file types begin with a binary header followed by a packed payload. The first 32 bytes are shared between the two versions, containing the magic number, version identifier, vector dimensionality, bit width, and total vector count.
Here is the exact block layout:
| Offset (bytes) | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | MAGIC |
ASCII string "TVIM" for v7, "TVV\0" (or "TV" followed by zero) for v6. |
| 4 | 2 | VERSION |
0x0006 for v6, 0x0007 for v7 (little-endian). |
| 6 | 2 | RESERVED |
Currently unused; set to zero. |
| 8 | 4 | DIM |
Vector dimensionality (u32). Must be a positive multiple of 8 and ≤ MAX_DIM (16,384). |
| 12 | 1 | BIT_WIDTH |
Bits per coordinate (valid valuses: 2, 3, or 4). |
| 13 | 6 | RESERVED2 |
Padding to 16-byte alignment (zeroed). |
| 16 | 8 | N_VECTORS |
Total number of stored vectors (u64). |
| 24 | 8 | HEADER_SIZE |
Total size of this header block in bytes (u64). Fixed at 32 for v6; v7 grows this because of optional metadata fields. |
For version 7, the header continues immediately after HEADER_SIZE:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 32 | 1 | HAS_ID_MAP |
0 = no ID map, 1 = ID map present. |
| 33 | 1 | HAS_CALIBRATION |
0 = uncalibrated, 1 = calibrated (TQ+). |
| 34 | 6 | RESERVED3 |
Padding (zero). |
| 40 | 8 × N_VECTORS |
ID_MAP (optional) |
Array of u64 external IDs, one per vector, only if HAS_ID_MAP == 1. |
| variable | – | Calibration data (optional) | If HAS_CALIBRATION == 1, two contiguous blocks of f32 values: SHIFT array (dim × f32) and SCALE array (dim × f32). |
The header is always aligned to 16 bytes so the subsequent compressed payload can be efficiently formatted.
Payload Structure After the Header
Both .tv and .tvim store the actual quantized vectors in a blocked, bit-packed format. This is the region that the SIMD search kernels consume directly.
The payload consists of three main components:
- Blocked cache (a sequence of blocks, each covering
32rows). For every block, vector coordinates are packed into a bit-plane representation. The number of blocks is calculated asn_blocks = ceiling(N_VECTORS / 32). The exact geometry is derived bypack::blocked_geometryin [src/pack.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/pack.rs). - Scales array. One
f32per vector representing the length-renormalization factor. These are stored consecutively immediately after the blocked cache. - Additional v7-only sections (if present): the ID map is written before the blocked cache so it can be memory-mapped without loading the entire file. The calibration blocks (
shiftandscale) follow the ID map, also before the cache whenHAS_CALIBRATIONis true.
Key Differences Between .tv (v6) and .tvim (v7)
| Feature | .tv (v6) |
.tvim (v7) |
|---|---|---|
| Magic bytes | "TVV\0" |
"TVIM" |
| Header size | Fixed 32 bytes (no extra fields) | Variable; includes optional ID map and calibration |
| ID map | Not stored – external IDs are absent | Optional u64 array per vector |
| Calibration (TQ+) | Not represented – index is always uncalibrated | Optional SHIFT/SCALE arrays (presence is marked by HAS_CALIBRATION) |
| Compatibility | Only TurboQuantIndex |
Used by IdMapIndex (which wraps a TurboQuantIndex) |
When you call TurboQuantIndex::write, the library always emits the v6 (.tv) snapshot. When IdMapIndex::write is called (from either Rust or Python), the v7 (.tvim) layout is used, preserving the external ID map and any calibration applied before saving.
Practical Code Examples
Writing a .tv file in Rust
use turbovec::TurboQuantIndex;
// Create a 1536-dim index, 4 bits per coordinate.
let mut idx = TurboQuantIndex::new(1536, 4).unwrap();
idx.add(&vectors);
// Serialize to the v6 format.
idx.write("my_index.tv").unwrap(); // <-- generates `.tv`
The write method on TurboQuantIndex delegates to io::write_v6, as defined in [src/lib.rs](https://github.com/RyanCodrai/turbovec/blob/main/src/lib.rs).
Writing an .tvim file in Python
from turbovec import IdMapIndex
import numpy as np
# Build an index that tracks external IDs.
idx = IdMapIndex(dim=1536, bit_width=4)
idx.add_with_ids(vectors, np.arange(len(vectors), dtype=np.uint64))
# Persist to v7 (includes IDs + calibration).
idx.write("my_index.tvim")
The Python binding forwards to IdMapIndex::write, which eventually calls io_v7::write_v7 loaded from src/io_v7.rs.
Loading a .tvim file again (Rust)
use turbovec::IdMapIndex;
// Automatically restores IDs and any calibration.
let idx = IdMapIndex::load("my_index.tvim").unwrap();
load reads the first 32 bytes, checks the version field, and then constructs either a plain TurboQuantIndex or an IdMapIndex, depending on the magic number and HAS_ID_MAP/HAS_CALIBRATION flags (see src/io.rs).
Quick Reference for write / load internals
| Constant / function | Location | Purpose |
|---|---|---|
MAGIC_V6 / MAGIC_V7 |
io.rs |
Detection of file version. |
write_v6 |
io.rs |
Serialises a TurboQuantIndex to the .tv layout. |
write_v7 |
io_v7.rs |
Serialises an IdMapIndex (plus optional ID map and calibration) to the .tvim layout. |
read_header |
io.rs |
Parses the first 32 bytes, validates magic, determines version. |
read_v6 / read_v7 |
io.rs / io_v7.rs |
Parse the rest of the file according to the version-specific schema. |
blocked_geometry |
pack.rs |
Computes block count, byte-group count, and total payload size from n_vectors, bit_width, and dim. |
repack / repack_block_range |
pack.rs |
Converts packed-row layout to the blocked cache format (needed when loading a v6 file and then mutating data). |
Summary
.tvfiles are version 6 binary snapshots containing only vector codes, block traps, and a fixed header..tvimfiles are version 7 and add the magic"TVIM", flags for an ID map and calibration, and optionalu64ID arrays plusSHIFT/SCALEblocks.- Both formats share a common 32-byte header with
DIM,BIT_WIDTH,N_VECTORS, andHEADER_SIZEfields. - The payload after the header uses a blocked, bit-packed cache format together with per-vector scales.
- Writing from the low-level
TurboQuantIndexproduces.tv; writing fromIdMapIndexproduces.tvimand restores the important additional metadata.
Frequently Asked Questions
How do I know whether a file on disk is .tv or .tvim?
Check the first four bytes. If they are "TVIM" the file is a v7 .tvim; if they begin with "TV" followed by a zero byte (or are simply "tv" in ASCII), it's a v6 .tv. The read_header function in io.rs performs exactly this check when loading.
Can I convert a .tv file into a .tvim file?
Yes, but you need to construct an IdMapIndex from the data first. Load the .tv using TurboQuantIndex::load, then call IdMapIndex::from_index providing your external IDs, and finally write out to the .tvim destination.
Why do my .tv files have no ID tracks?
Version 6 explicitly stores no external separator IDs. If your application requires a stable mapping between vectors and external identifiers (e.g., document or graph node IDs), you should use v7 (.tvim) from the start or convert to v7 during the write phase.
Where in the source can I find the full binary layout structs?
The primary struct definitions are in src/io.rs (the GenericIndexHeader equivalent and v6 writer) and src/io_v7.rs (the v7 writer and reader). The pack.rs module documents the exact byte-offset calculations via blocked_geometry.
The details above are derived from the Turbovec source files io.rs, io_v7.rs, and pack.rs. For the complete latest layout, always check the official repository documentation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →