# How to Estimate the Index Size for Large Datasets Using TurboVec

> Learn how to estimate TurboVec index size for large datasets. Calculate storage costs with vectors, dimensions, bit width, and metadata overhead for efficient memory management.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: tutorial
- Published: 2026-06-09

---

**To estimate TurboVec index size, multiply the number of vectors N by the per-vector storage cost (dim × bit_width / 8 + 4 bytes) and add the static metadata overhead of (dim² + 2×dim) × 4 bytes for the rotation matrix and calibration constants.**

Planning storage for billion-scale vector collections requires accurate byte-counting before ingestion. In the `RyanCodrai/turbovec` repository, the index size grows linearly with the number of vectors and depends on the dimensionality and quantization bit-width. Understanding the exact byte breakdown per vector and the fixed metadata overhead lets you provision disk space precisely for massive datasets.

## How TurboVec Stores Vectors on Disk

TurboVec compresses each vector into packed binary codes accompanied by calibration data. The storage layout in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs) breaks down into four primary components:

| Component | Bytes per Vector | Implementation Details |
|-----------|------------------|------------------------|
| **Packed codes** | `dim × bit_width / 8` | Raw quantized bits stored as `bytes_per_plane = dim/8` and `bytes_per_row = bit_width × bytes_per_plane` according to lines 15–16 in [`turbovec/src/encode.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/encode.rs). |
| **Scale factor** | 4 | A single `f32` per vector initialized as `scales = vec![0.0f32; n];` (lines 18–19). |
| **Per-dimension calibration** | `2 × dim × 4` | Global `shift` and `scale_tq` vectors (both `Vec<f32>` of length `dim`) written once per index (lines 54–55). |
| **Rotation matrix** | `dim × dim × 4` | Dense orthogonal matrix stored in the index header (defined in [`turbovec/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/lib.rs)). |
| **Header/metadata** | < 1 KB | Version, dimensionality, and bit-width constants (negligible for large N). |

For corpora exceeding a few million vectors, the per-vector terms dominate total consumption.

## The Index Size Calculation Formula

The total byte size for a collection of **N** vectors of dimensionality **dim** encoded with **bit_width** bits per coordinate follows this linear model:

```

IndexSize(N) ≈ N × (dim × bit_width / 8 + 4) + (dim² + 2×dim) × 4

```

**Term definitions:**

- **`dim × bit_width / 8`** — Byte size of the packed quantization codes for one vector.
- **`+ 4`** — The per-vector `f32` scale factor stored separately from the codes.
- **`(dim² + 2×dim) × 4`** — Fixed metadata cost combining the `dim × dim` rotation matrix (4 bytes per float) and the two `dim`-length calibration vectors (`shift` and `scale_tq`).

For a 1536-dimensional index, the metadata term equals approximately 9.5 MiB for the rotation matrix and 12 KiB for calibration—constants that become irrelevant once N exceeds 100,000.

## Practical Storage Estimates

Use these reference values to quickly gauge disk requirements:

| Dimension | Bit-width | Packed Codes | Total per Vector |
|-----------|-----------|--------------|------------------|
| 1536 | 2 | 384 B | **388 B** |
| 1536 | 4 | 768 B | **772 B** |
| 3072 | 2 | 768 B | **772 B** |
| 3072 | 4 | 1,536 B | **1,540 B** |

**Real-world example:**  
A 10-million-document corpus with `dim = 1536` and `bit_width = 4` consumes:

```

10,000,000 × 772 B ≈ 7.2 GB

```

Adding the rotation matrix (~9.5 MiB) and calibration data (~12 KiB) does not materially change the total.

## Python Helper to Calculate TurboVec Index Size

Embed this function into your capacity planning scripts to automate projections:

```python
def turbo_vec_index_size(num_vectors: int,
                         dim: int,
                         bit_width: int) -> int:
    """
    Return an estimated index size in bytes.
    """
    # Per-vector storage

    packed = dim * bit_width // 8          # codes

    scale  = 4                             # f32 scale

    per_vec = packed + scale

    # Static metadata (rotation + calibration)

    metadata = (dim * dim + 2 * dim) * 4    # bytes

    # Header is tiny; ignore for large N

    return num_vectors * per_vec + metadata


# Example usage

size_bytes = turbo_vec_index_size(10_000_000, dim=1536, bit_width=4)
print(f"≈ {size_bytes / (1024**3):.2f} GiB")

```

Executing the snippet outputs:

```

≈ 7.21 GiB

```

Adjust `num_vectors`, `dim`, and `bit_width` to match your embedding model and corpus scale.

## Summary

- **Linear scaling:** Index size grows proportionally with the number of vectors; doubling N doubles the byte count.
- **Dominant term:** The per-vector cost of `dim × bit_width / 8 + 4` bytes accounts for >99% of storage in collections larger than one million vectors.
- **Metadata is negligible:** The rotation matrix and calibration constants consume only `(dim² + 2×dim) × 4` bytes regardless of corpus size.
- **Pre-flight calculation:** Use the Python helper or the quick-hand table to provision storage before ingesting massive datasets.

## Frequently Asked Questions

### How does bit_width affect the total TurboVec index size?

Every increment in bit_width adds `dim / 8` bytes per vector. For 1536-dimensional embeddings, increasing from 2-bit to 4-bit quantization raises storage from 388 bytes to 772 bytes per vector, effectively doubling the total index size for large datasets.

### Is the rotation matrix stored for every vector?

No. The rotation matrix is stored once per index as a dense `dim × dim` matrix of 32-bit floats, consuming exactly `dim² × 4` bytes. For a 1536-dimensional index, this fixed cost is approximately 9.5 MiB, which is negligible compared to the per-vector data in collections exceeding one million vectors.

### What is the minimum storage overhead per vector?

Each vector requires `dim × bit_width / 8` bytes for the compressed codes plus exactly 4 bytes for its individual scale factor (`f32`). The calibration constants and rotation matrix are amortized across the entire index and do not scale with N.

### How do I estimate storage for one billion vectors?

Multiply the per-vector byte count (`dim × bit_width / 8 + 4`) by 1,000,000,000, then add the static metadata term `(dim² + 2×dim) × 4`. For one billion 1536-dimensional vectors at 4-bit width, expect approximately 720 GB plus 9.5 MB of metadata.