How to Estimate the Index Size for Large Datasets Using TurboVec

To estimate TurboVec index size, multiply the number of vectors N by the per-vector storage cost (dim × bit_width / 8 + 4 bytes) and add the static metadata overhead of (dim² + 2×dim) × 4 bytes for the rotation matrix and calibration constants.

Planning storage for billion-scale vector collections requires accurate byte-counting before ingestion. In the RyanCodrai/turbovec repository, the index size grows linearly with the number of vectors and depends on the dimensionality and quantization bit-width. Understanding the exact byte breakdown per vector and the fixed metadata overhead lets you provision disk space precisely for massive datasets.

How TurboVec Stores Vectors on Disk

TurboVec compresses each vector into packed binary codes accompanied by calibration data. The storage layout in turbovec/src/encode.rs breaks down into four primary components:

Component Bytes per Vector Implementation Details
Packed codes dim × bit_width / 8 Raw quantized bits stored as bytes_per_plane = dim/8 and bytes_per_row = bit_width × bytes_per_plane according to lines 15–16 in turbovec/src/encode.rs.
Scale factor 4 A single f32 per vector initialized as scales = vec![0.0f32; n]; (lines 18–19).
Per-dimension calibration 2 × dim × 4 Global shift and scale_tq vectors (both Vec<f32> of length dim) written once per index (lines 54–55).
Rotation matrix dim × dim × 4 Dense orthogonal matrix stored in the index header (defined in turbovec/src/lib.rs).
Header/metadata < 1 KB Version, dimensionality, and bit-width constants (negligible for large N).

For corpora exceeding a few million vectors, the per-vector terms dominate total consumption.

The Index Size Calculation Formula

The total byte size for a collection of N vectors of dimensionality dim encoded with bit_width bits per coordinate follows this linear model:


IndexSize(N) ≈ N × (dim × bit_width / 8 + 4) + (dim² + 2×dim) × 4

Term definitions:

  • dim × bit_width / 8 — Byte size of the packed quantization codes for one vector.
  • + 4 — The per-vector f32 scale factor stored separately from the codes.
  • (dim² + 2×dim) × 4 — Fixed metadata cost combining the dim × dim rotation matrix (4 bytes per float) and the two dim-length calibration vectors (shift and scale_tq).

For a 1536-dimensional index, the metadata term equals approximately 9.5 MiB for the rotation matrix and 12 KiB for calibration—constants that become irrelevant once N exceeds 100,000.

Practical Storage Estimates

Use these reference values to quickly gauge disk requirements:

Dimension Bit-width Packed Codes Total per Vector
1536 2 384 B 388 B
1536 4 768 B 772 B
3072 2 768 B 772 B
3072 4 1,536 B 1,540 B

Real-world example:
A 10-million-document corpus with dim = 1536 and bit_width = 4 consumes:


10,000,000 × 772 B ≈ 7.2 GB

Adding the rotation matrix (~9.5 MiB) and calibration data (~12 KiB) does not materially change the total.

Python Helper to Calculate TurboVec Index Size

Embed this function into your capacity planning scripts to automate projections:

def turbo_vec_index_size(num_vectors: int,
                         dim: int,
                         bit_width: int) -> int:
    """
    Return an estimated index size in bytes.
    """
    # Per-vector storage

    packed = dim * bit_width // 8          # codes

    scale  = 4                             # f32 scale

    per_vec = packed + scale

    # Static metadata (rotation + calibration)

    metadata = (dim * dim + 2 * dim) * 4    # bytes

    # Header is tiny; ignore for large N

    return num_vectors * per_vec + metadata


# Example usage

size_bytes = turbo_vec_index_size(10_000_000, dim=1536, bit_width=4)
print(f"≈ {size_bytes / (1024**3):.2f} GiB")

Executing the snippet outputs:


≈ 7.21 GiB

Adjust num_vectors, dim, and bit_width to match your embedding model and corpus scale.

Summary

  • Linear scaling: Index size grows proportionally with the number of vectors; doubling N doubles the byte count.
  • Dominant term: The per-vector cost of dim × bit_width / 8 + 4 bytes accounts for >99% of storage in collections larger than one million vectors.
  • Metadata is negligible: The rotation matrix and calibration constants consume only (dim² + 2×dim) × 4 bytes regardless of corpus size.
  • Pre-flight calculation: Use the Python helper or the quick-hand table to provision storage before ingesting massive datasets.

Frequently Asked Questions

How does bit_width affect the total TurboVec index size?

Every increment in bit_width adds dim / 8 bytes per vector. For 1536-dimensional embeddings, increasing from 2-bit to 4-bit quantization raises storage from 388 bytes to 772 bytes per vector, effectively doubling the total index size for large datasets.

Is the rotation matrix stored for every vector?

No. The rotation matrix is stored once per index as a dense dim × dim matrix of 32-bit floats, consuming exactly dim² × 4 bytes. For a 1536-dimensional index, this fixed cost is approximately 9.5 MiB, which is negligible compared to the per-vector data in collections exceeding one million vectors.

What is the minimum storage overhead per vector?

Each vector requires dim × bit_width / 8 bytes for the compressed codes plus exactly 4 bytes for its individual scale factor (f32). The calibration constants and rotation matrix are amortized across the entire index and do not scale with N.

How do I estimate storage for one billion vectors?

Multiply the per-vector byte count (dim × bit_width / 8 + 4) by 1,000,000,000, then add the static metadata term (dim² + 2×dim) × 4. For one billion 1536-dimensional vectors at 4-bit width, expect approximately 720 GB plus 9.5 MB of metadata.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →