# How to Serialize and Deserialize Turbovec Indices Using `tobytes` and `frombytes`

> Serialize and deserialize Turbovec indices with tobytes and frombytes. Save and load your in-memory index efficiently using this method.

- Repository: [Ryan Codrai/turbovec](https://github.com/RyanCodrai/turbovec)
- Tags: how-to-guide
- Published: 2026-07-27

---

**Use `index.tobytes()` to serialize an in-memory Turbovec index into a `bytes` object, and `Index.frombytes(binary_data)` to reconstruct a fully searchable index from that blob.**

The `turbovec` library stores its searchable vector data in a binary `.tvim` format managed by the Rust-based `IdMapIndex` engine. The Python wrapper exposes `tobytes` and `frombytes` methods that let you serialize and deserialize turbovec indices without touching the filesystem. These methods are ideal for caching indexes in memory, transmitting them across a network, or snapshotting state for testing.

## How `tobytes` and `frombytes` Work

Turbovec delegates heavy vector operations to the Rust core, where `IdMapIndex::to_bytes` and `IdMapIndex::from_bytes` handle the actual binary conversion. The Python API surfaces this functionality through two convenience methods on every `Index` instance:

- **`index.tobytes()`** — Serializes the in-memory index into a contiguous `bytes` object. The payload contains the exact binary representation that would be written to a `.tvim` file.
- **`Index.frombytes(data)`** — Accepts a `bytes` object and instantiates a fresh `Index` backed by a restored `IdMapIndex`.

Because the Python wrapper forwards calls directly to the Rust binding, any updates to the core serialization logic in [`turbovec-rust/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-rust/src/lib.rs) are immediately available in Python.

## Serializing a Turbovec Index to Bytes

After you build an index, call `tobytes()` to capture its full binary state. The returned object is a raw byte string suitable for file storage, in-memory caches, or network transmission.

```python
from turbovec import Index, Document

docs = [
    Document(id="doc-1", text="The quick brown fox jumps over the lazy dog."),
    Document(id="doc-2", text="Lorem ipsum dolor sit amet, consectetur adipiscing elit."),
]

idx = Index()
for doc in docs:
    idx.add_document(doc)

idx.build()

binary_blob = idx.tobytes()

with open("my_index.tvim", "wb") as f:
    f.write(binary_blob)

```

## Deserializing a Turbovec Index from Bytes

Reconstructing an index is a single call to `Index.frombytes()`. The restored index is immediately searchable and behaves exactly like the original.

```python
import turbovec

# From an in-memory bytes object

restored_idx = turbovec.Index.frombytes(binary_blob)

# Or load from a file on disk

with open("my_index.tvim", "rb") as f:
    file_blob = f.read()

restored_idx = turbovec.Index.frombytes(file_blob)

results = restored_idx.search("quick fox", k=5)
print([doc.id for doc in results])

```

## Source Files and Low-Level Implementation

The binary serialization pipeline spans the Rust core and the Python shim. Key files in the `RyanCodrai/turbovec` repository include:

- **[`turbovec-python/python/turbovec/__init__.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/__init__.py)** — Exposes the public `Index` class and its `tobytes` and `frombytes` methods.
- **[`turbovec-rust/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-rust/src/lib.rs)** — Implements `IdMapIndex::to_bytes` and `IdMapIndex::from_bytes`, which perform the actual memory-to-binary conversion.
- **[`turbovec-python/python/turbovec/_persist.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/_persist.py)** — Provides the `atomic_save` helper that writes both the `.tvim` binary and the JSON side-car when using higher-level persistence.
- **[`turbovec-python/tests/test_persist.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/tests/test_persist.py)** — Contains unit tests that verify the round-trip correctness of `tobytes` and `frombytes`.

If you need a complete on-disk snapshot that includes both the binary index and the JSON side-car mapping handles to documents, use `index.save()`. Internally, it calls `tobytes()`, writes the `.tvim` file, and serializes the side-car via the routines in [`_persist.py`](https://github.com/RyanCodrai/turbovec/blob/main/_persist.py).

## Important Considerations for Binary Persistence

Keep the following in mind when using `tobytes` and `frombytes`:

- **Side-car data is excluded.** `tobytes()` serializes only the vector index. If your application relies on the original document payloads during search, maintain the JSON side-car separately or use `index.save()` and `index.load()`.
- **Format stability.** The binary format is stable across Turbovec releases, but deserializing with a future major version may raise a `ValueError` if the on-disk format changes. Keep library versions in sync with persisted bytes.
- **Payload size.** Because the blob is a direct dump of the Rust `IdMapIndex`, it can reach several hundred megabytes for million-vector indexes. When transmitting over a network, compress the payload with `gzip` or similar before transmission and decompress before calling `frombytes()`.

## Summary

- **`index.tobytes()`** converts a Turbovec index into a raw `bytes` object containing the binary `.tvim` representation.
- **`Index.frombytes(data)`** reconstructs a searchable index from that `bytes` object via the Rust `IdMapIndex::from_bytes` implementation.
- The Python API in [`turbovec-python/python/turbovec/__init__.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/__init__.py) forwards these calls directly to the Rust core in [`turbovec-rust/src/lib.rs`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-rust/src/lib.rs).
- Binary serialization does **not** include the JSON side-car; use `index.save()` and `index.load()` when you need both the index and document mappings.
- Large byte payloads should be compressed for network transfer to reduce bandwidth.

## Frequently Asked Questions

### Does `tobytes` also save the document metadata?

No. The `tobytes()` method serializes only the binary vector index managed by `IdMapIndex`. Document metadata lives in a separate JSON side-car. If you need to preserve both, use the higher-level `index.save()` helper in [`turbovec-python/python/turbovec/_persist.py`](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/python/turbovec/_persist.py).

### Can I deserialize an index created with a different version of Turbovec?

Deserializing across minor versions is generally safe, but a future major version may introduce breaking changes to the `.tvim` format. Always match the Turbovec library version to the version used to produce the serialized bytes to avoid `ValueError` exceptions.

### What is the performance cost of calling `tobytes` versus `index.save()`?

`tobytes()` is faster because it skips the JSON side-car serialization and the atomic filesystem writes handled by [`_persist.py`](https://github.com/RyanCodrai/turbovec/blob/main/_persist.py). If you only need the binary index object in memory or want to manage your own storage, prefer `tobytes()` and `frombytes()`.

### How large is the `bytes` object returned by `tobytes()`?

The payload is a direct memory dump of the Rust index, so its size scales with the number of vectors and the chosen algorithm parameters. For million-vector indexes, expect hundreds of megabytes. Compress the blob with standard tools before network transmission to reduce overhead.