# Turbovec Index File Formats Explained: `.tv` and `.tvim` Binary Structure

> Explore the Turbovec index file formats .tv and .tvim. Understand the binary structure for compact snapshots and persistent indexes. Learn how Turbovec stores vector data.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: internals
- Published: 2026-08-22

---

**TLDR:** Turbovec stores vector indexes in two binary formats — `.tv` (version 6) for compact, ID-less snapshots and `.tvim` (version 7) for persistent indexes with external ID maps and optional TQ+ calibration data — both sharing a common header layout defined in the [`turbovec/src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/io.rs) source file.

The **Turbovec** repository (RyanCodrai/turbovec) is a high-performance vector similarity search library written in Rust with Python bindings. Its index persistence layer uses two distinct on-disk formats that differ by version, magic bytes, and metadata sections. Understanding the **turbovec index file format** is essential if you plan to read, write, or memory-map indexes directly, or if you need to debug interop issues between the v6 and v7 code paths. This guide breaks down the byte-level structure, version differences, and practical code examples for both `.tv` and `.tvim` files.

## Format Overview: v6 vs v7

Turbovec writes indexes in one of two binary layouts depending on which data structure you're persisting:

| Format | Extension | Version | Primary Use |
|--------|-----------|---------|-------------|
| **TurboQuant v6** | `.tv` | 6 | Compact, read-only snapshots of a `TurboQuantIndex` (no external IDs or calibration data). |
| **TurboQuant v7** | `.tvim` | 7 | Persistent indexes that also contain an formula ID map and optional TQ+ calibration information (used by `IdMapIndex`). |

The exposed `write` and `load` methods in the library auto-detect which version to produce or consume. The top-level serialization logic resides in [[`src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/io.rs)](https://github.com/RyanCodrai/turbovec/blob/main/src/io.rs), while the v7 extensions are isolated in [[`src/io_v7.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/io_v7.rs)](https://github.com/RyanCodrai/turbovec/blob/main/src/io_v7.rs).

## Common Header Structure for `.tv` and `.tvim`

Both file types begin with a binary header followed by a packed payload. The first 32 bytes are shared between the two versions, containing the magic number, version identifier, vector dimensionality, bit width, and total vector count.

Here is the exact block layout:

| Offset (bytes) | Size | Field | Meaning |
|----------------|------|-------|---------|
| 0 | 4 | `MAGIC` | ASCII string `"TVIM"` for v7, `"TVV\0"` (or `"TV"` followed by zero) for v6. |
| 4 | 2 | `VERSION` | `0x0006` for v6, `0x0007` for v7 (little-endian). |
| 6 | 2 | `RESERVED` | Currently unused; set to zero. |
| 8 | 4 | `DIM` | Vector dimensionality (`u32`). Must be a positive multiple of 8 and ≤ `MAX_DIM` (16,384). |
| 12 | 1 | `BIT_WIDTH` | Bits per coordinate (valid valuses: `2`, `3`, or `4`). |
| 13 | 6 | `RESERVED2` | Padding to 16-byte alignment (zeroed). |
| 16 | 8 | `N_VECTORS` | Total number of stored vectors (`u64`). |
| 24 | 8 | `HEADER_SIZE` | Total size of this header block in bytes (`u64`). Fixed at 32 for v6; v7 grows this because of optional metadata fields. |

For version 7, the header continues immediately after `HEADER_SIZE`:

| Offset | Size | Field | Meaning |
|--------|------|-------|---------|
| 32 | 1 | `HAS_ID_MAP` | `0` = no ID map, `1` = ID map present. |
| 33 | 1 | `HAS_CALIBRATION` | `0` = uncalibrated, `1` = calibrated (TQ+). |
| 34 | 6 | `RESERVED3` | Padding (zero). |
| 40 | 8 × `N_VECTORS` | `ID_MAP` (optional) | Array of `u64` external IDs, one per vector, only if `HAS_ID_MAP == 1`. |
| variable | – | **Calibration data** (optional) | If `HAS_CALIBRATION == 1`, two contiguous blocks of `f32` values: `SHIFT` array (`dim × f32`) and `SCALE` array (`dim × f32`). |

The header is always aligned to 16 bytes so the subsequent compressed payload can be efficiently formatted.

## Payload Structure After the Header

Both `.tv` and `.tvim` store the actual quantized vectors in a blocked, bit-packed format. This is the region that the SIMD search kernels consume directly.

The payload consists of three main components:

- **Blocked cache** (a sequence of blocks, each covering `32` rows). For every block, vector coordinates are packed into a bit-plane representation. The number of blocks is calculated as `n_blocks = ceiling(N_VECTORS / 32)`. The exact geometry is derived by `pack::blocked_geometry` in [[`src/pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/pack.rs)](https://github.com/RyanCodrai/turbovec/blob/main/src/pack.rs).
- **Scales** array. One `f32` per vector representing the length-renormalization factor. These are stored consecutively immediately after the blocked cache.
- Additional v7-only sections (if present): the ID map is written *before* the blocked cache so it can be memory-mapped without loading the entire file. The calibration blocks (`shift` and `scale`) follow the ID map, also before the cache when `HAS_CALIBRATION` is true.

## Key Differences Between `.tv` (v6) and `.tvim` (v7)

| Feature | `.tv` (v6) | `.tvim` (v7) |
|---------|------------|--------------|
| Magic bytes | `"TVV\0"` | `"TVIM"` |
| Header size | Fixed 32 bytes (no extra fields) | Variable; includes optional ID map and calibration |
| ID map | **Not** stored – external IDs are absent | Optional `u64` array per vector |
| Calibration (TQ+) | Not represented – index is always uncalibrated | Optional `SHIFT`/`SCALE` arrays (presence is marked by `HAS_CALIBRATION`) |
| Compatibility | Only `TurboQuantIndex` | Used by `IdMapIndex` (which wraps a `TurboQuantIndex`) |

When you call `TurboQuantIndex::write`, the library always emits the v6 (`.tv`) snapshot. When `IdMapIndex::write` is called (from either Rust or Python), the v7 (`.tvim`) layout is used, preserving the external ID map and any calibration applied before saving.

## Practical Code Examples

### Writing a `.tv` file in Rust


```rust
use turbovec::TurboQuantIndex;

// Create a 1536-dim index, 4 bits per coordinate.
let mut idx = TurboQuantIndex::new(1536, 4).unwrap();
idx.add(&vectors);

// Serialize to the v6 format.
idx.write("my_index.tv").unwrap(); // <-- generates `.tv`

```


The `write` method on `TurboQuantIndex` delegates to `io::write_v6`, as defined in [[`src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/lib.rs)](https://github.com/RyanCodrai/turbovec/blob/main/src/lib.rs).

### Writing an `.tvim` file in Python


```python
from turbovec import IdMapIndex
import numpy as np

# Build an index that tracks external IDs.

idx = IdMapIndex(dim=1536, bit_width=4)
idx.add_with_ids(vectors, np.arange(len(vectors), dtype=np.uint64))

# Persist to v7 (includes IDs + calibration).

idx.write("my_index.tvim")

```


The Python binding forwards to `IdMapIndex::write`, which eventually calls `io_v7::write_v7` loaded from [`src/io_v7.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/io_v7.rs).

### Loading a `.tvim` file again (Rust)


```rust
use turbovec::IdMapIndex;

// Automatically restores IDs and any calibration.
let idx = IdMapIndex::load("my_index.tvim").unwrap();

```


`load` reads the first 32 bytes, checks the version field, and then constructs either a plain `TurboQuantIndex` or an `IdMapIndex`, depending on the magic number and `HAS_ID_MAP`/`HAS_CALIBRATION` flags (see [`src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/io.rs)).

## Quick Reference for `write` / `load` internals

| Constant / function | Location | Purpose |
|----------------------|----------|---------|
| `MAGIC_V6` / `MAGIC_V7` | [`io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io.rs) | Detection of file version. |
| `write_v6` | [`io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io.rs) | Serialises a `TurboQuantIndex` to the `.tv` layout. |
| `write_v7` | [`io_v7.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io_v7.rs) | Serialises an `IdMapIndex` (plus optional ID map and calibration) to the `.tvim` layout. |
| `read_header` | [`io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io.rs) | Parses the first 32 bytes, validates magic, determines version. |
| `read_v6` / `read_v7` | [`io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io.rs) / [`io_v7.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io_v7.rs) | Parse the rest of the file according to the version-specific schema. |
| `blocked_geometry` | [`pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/pack.rs) | Computes block count, byte-group count, and total payload size from `n_vectors`, `bit_width`, and `dim`. |
| `repack` / `repack_block_range` | [`pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/pack.rs) | Converts packed-row layout to the blocked cache format (needed when loading a v6 file and then mutating data). |

## Summary

- **`.tv` files** are version 6 binary snapshots containing only vector codes, block traps, and a fixed header.
- **`.tvim` files** are version 7 and add the magic `"TVIM"`, flags for an ID map and calibration, and optional `u64` ID arrays plus `SHIFT`/`SCALE` blocks.
- Both formats share a common 32-byte header with `DIM`, `BIT_WIDTH`, `N_VECTORS`, and `HEADER_SIZE` fields.
- The payload after the header uses a blocked, bit-packed cache format together with per-vector scales.
- Writing from the low-level `TurboQuantIndex` produces `.tv`; writing from `IdMapIndex` produces `.tvim` and restores the important additional metadata.

## Frequently Asked Questions

### How do I know whether a file on disk is `.tv` or `.tvim`?

Check the first four bytes. If they are `"TVIM"` the file is a v7 `.tvim`; if they begin with `"TV"` followed by a zero byte (or are simply `"tv"` in ASCII), it's a v6 `.tv`. The `read_header` function in [`io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io.rs) performs exactly this check when loading.

### Can I convert a `.tv` file into a `.tvim` file?

Yes, but you need to construct an `IdMapIndex` from the data first. Load the `.tv` using `TurboQuantIndex::load`, then call `IdMapIndex::from_index` providing your external IDs, and finally write out to the `.tvim` destination.

### Why do my `.tv` files have no ID tracks?

Version 6 explicitly stores no external separator IDs. If your application requires a stable mapping between vectors and external identifiers (e.g., document or graph node IDs), you should use v7 (`.tvim`) from the start or convert to v7 during the write phase.

### Where in the source can I find the full binary layout structs?

The primary struct definitions are in [`src/io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/io.rs) (the `GenericIndexHeader` equivalent and v6 writer) and [`src/io_v7.rs`](https://github.com/RyanCodrai/turbovec/blob/main/src/io_v7.rs) (the v7 writer and reader). The [`pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/pack.rs) module documents the exact byte-offset calculations via `blocked_geometry`.

---

The details above are derived from the Turbovec source files [`io.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io.rs), [`io_v7.rs`](https://github.com/RyanCodrai/turbovec/blob/main/io_v7.rs), and [`pack.rs`](https://github.com/RyanCodrai/turbovec/blob/main/pack.rs). For the complete latest layout, always check the official repository documentation.