How to Serialize and Deserialize Turbovec Indices Using `tobytes` and `frombytes`

Use index.tobytes() to serialize an in-memory Turbovec index into a bytes object, and Index.frombytes(binary_data) to reconstruct a fully searchable index from that blob.

The turbovec library stores its searchable vector data in a binary .tvim format managed by the Rust-based IdMapIndex engine. The Python wrapper exposes tobytes and frombytes methods that let you serialize and deserialize turbovec indices without touching the filesystem. These methods are ideal for caching indexes in memory, transmitting them across a network, or snapshotting state for testing.

How tobytes and frombytes Work

Turbovec delegates heavy vector operations to the Rust core, where IdMapIndex::to_bytes and IdMapIndex::from_bytes handle the actual binary conversion. The Python API surfaces this functionality through two convenience methods on every Index instance:

  • index.tobytes() — Serializes the in-memory index into a contiguous bytes object. The payload contains the exact binary representation that would be written to a .tvim file.
  • Index.frombytes(data) — Accepts a bytes object and instantiates a fresh Index backed by a restored IdMapIndex.

Because the Python wrapper forwards calls directly to the Rust binding, any updates to the core serialization logic in turbovec-rust/src/lib.rs are immediately available in Python.

Serializing a Turbovec Index to Bytes

After you build an index, call tobytes() to capture its full binary state. The returned object is a raw byte string suitable for file storage, in-memory caches, or network transmission.

from turbovec import Index, Document

docs = [
    Document(id="doc-1", text="The quick brown fox jumps over the lazy dog."),
    Document(id="doc-2", text="Lorem ipsum dolor sit amet, consectetur adipiscing elit."),
]

idx = Index()
for doc in docs:
    idx.add_document(doc)

idx.build()

binary_blob = idx.tobytes()

with open("my_index.tvim", "wb") as f:
    f.write(binary_blob)

Deserializing a Turbovec Index from Bytes

Reconstructing an index is a single call to Index.frombytes(). The restored index is immediately searchable and behaves exactly like the original.

import turbovec

# From an in-memory bytes object

restored_idx = turbovec.Index.frombytes(binary_blob)

# Or load from a file on disk

with open("my_index.tvim", "rb") as f:
    file_blob = f.read()

restored_idx = turbovec.Index.frombytes(file_blob)

results = restored_idx.search("quick fox", k=5)
print([doc.id for doc in results])

Source Files and Low-Level Implementation

The binary serialization pipeline spans the Rust core and the Python shim. Key files in the RyanCodrai/turbovec repository include:

If you need a complete on-disk snapshot that includes both the binary index and the JSON side-car mapping handles to documents, use index.save(). Internally, it calls tobytes(), writes the .tvim file, and serializes the side-car via the routines in _persist.py.

Important Considerations for Binary Persistence

Keep the following in mind when using tobytes and frombytes:

  • Side-car data is excluded. tobytes() serializes only the vector index. If your application relies on the original document payloads during search, maintain the JSON side-car separately or use index.save() and index.load().
  • Format stability. The binary format is stable across Turbovec releases, but deserializing with a future major version may raise a ValueError if the on-disk format changes. Keep library versions in sync with persisted bytes.
  • Payload size. Because the blob is a direct dump of the Rust IdMapIndex, it can reach several hundred megabytes for million-vector indexes. When transmitting over a network, compress the payload with gzip or similar before transmission and decompress before calling frombytes().

Summary

  • index.tobytes() converts a Turbovec index into a raw bytes object containing the binary .tvim representation.
  • Index.frombytes(data) reconstructs a searchable index from that bytes object via the Rust IdMapIndex::from_bytes implementation.
  • The Python API in turbovec-python/python/turbovec/__init__.py forwards these calls directly to the Rust core in turbovec-rust/src/lib.rs.
  • Binary serialization does not include the JSON side-car; use index.save() and index.load() when you need both the index and document mappings.
  • Large byte payloads should be compressed for network transfer to reduce bandwidth.

Frequently Asked Questions

Does tobytes also save the document metadata?

No. The tobytes() method serializes only the binary vector index managed by IdMapIndex. Document metadata lives in a separate JSON side-car. If you need to preserve both, use the higher-level index.save() helper in turbovec-python/python/turbovec/_persist.py.

Can I deserialize an index created with a different version of Turbovec?

Deserializing across minor versions is generally safe, but a future major version may introduce breaking changes to the .tvim format. Always match the Turbovec library version to the version used to produce the serialized bytes to avoid ValueError exceptions.

What is the performance cost of calling tobytes versus index.save()?

tobytes() is faster because it skips the JSON side-car serialization and the atomic filesystem writes handled by _persist.py. If you only need the binary index object in memory or want to manage your own storage, prefer tobytes() and frombytes().

How large is the bytes object returned by tobytes()?

The payload is a direct memory dump of the Rust index, so its size scales with the number of vectors and the chosen algorithm parameters. For million-vector indexes, expect hundreds of megabytes. Compress the blob with standard tools before network transmission to reduce overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →