How to Perform Incremental Saves with Turbovec Using `sync()`

Turbovec's sync() method enables incremental saves by detecting dirty slots and writing only the changed portions of the index to disk, rather than rewriting the entire file on every call.

Turbovec is a high-performance vector index library that stores vectors in a v7 container format with built-in support for delta persistence. As implemented in the RyanCodrai/turbovec repository, the sync() method tracks which slots have been modified since the last save, making repeated persistence calls extremely efficient. This article explains exactly how to perform incremental saves with turbovec using sync(), backed by the actual source code.

How sync() Achieves Incremental Saves

The core intelligence behind Turbovec's efficient persistence lies in its slot-level change tracking. When you mutate an index—by adding vectors, updating metadata, or performing deletions—Turbovec internally flags the affected slots as dirty. The next call to sync() consults these flags and writes only the changed data.

In the actual implementation, the Index::sync method (located in src/index.rs) performs three steps:

  1. Detects dirty slots using an internal generation counter that tracks which slots have been modified.
  2. Writes a temporary file containing only those dirty slots, along with an updated header that records the new generation.
  3. Atomically swaps the temporary file over the original, guaranteeing crash-safety even if the process is interrupted mid-write.

This design means that the first call to sync() necessarily writes the full container (since every slot is considered new), but every subsequent call only touches the parts that actually changed.

The Generation Counter Mechanism

Turbovec tracks each slot's generation number. When you mutate slot 5, its generation counter increments. During sync(), the method compares each in-memory slot's generation against the stored header from the last save. Any slot whose generation is newer than the recorded one is considered dirty and gets written. This is what enables true incremental saves without requiring a full file scan.

How to Perform Incremental Saves with Turbovec Using sync()

The usage pattern is straightforward. You create or load an index, mutate it using the available operations, then call sync() to persist changes. Thanks to the dirty-slot detection, every sync() call after the initial one becomes an incremental save.

Python Example

The Python wrapper (turbovec-python) exposes the same sync() method:

import turbovec

# Create an index with dimension 1536

idx = turbovec.Index(dim=1536)

# First sync - writes the full container (initial creation)

idx.sync("/tmp/my_index")      # Full write, generation 0

# Later, after adding new vectors...

idx.add(new_vectors)           # Marks only the new slots as dirty

idx.sync("/tmp/my_index")      # Incremental save - only changed slots

Note that idx.add() automatically marks the affected slots as dirty. You never need to manually indicate what changed.

Rust Example

In Rust, the flow is identical:

use turbovec::Index;
use std::path::Path;

fn main() -> Result<(), turbovec::Error> {
    // Create an empty index
    let mut idx = Index::new(1536);

    // First sync writes the whole container
    idx.sync(&Path::new("my_index.tvc"))?;

    // Add vectors later — new slots become dirty
    idx.add(&additional_vectors)?;
    idx.sync(&Path::new("my_index.tvc"))?; // Incremental save

    Ok(())
}

Each sync() call picks up only the slots that were modified since the last save, keeping file I/O to a minimum.

When Incremental Saves Are Not Possible

There are a few situations where sync() must perform a full rewrite instead of an incremental one. According to the test suite in [turbovec/tests/sync_v7.rs](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/sync_v7.rs), a full write happens when:

  • The index was loaded in memory and the on-disk generation does not match
  • A large batch of deletions causes slot reclamation, requiring a reshape of the file
  • The internal header is out of sync due to external modification of the file

In these cases, sync() detects that a delta write would be unsafe and performs a full container rewrite. This ensures data integrity even though it costs more I/O.

Key Source Files for sync()

To dig deeper into the implementation, the relevant source files are:

File Purpose
[turbovec/src/index.rs](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/src/index.rs) Contains the Index::sync method implementation, dirty-slot tracking, and atomic file swap logic.
[turbovec/tests/v7_only.rs](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/v7_only.rs) Demonstrates the full-write-then-incremental pattern with tests.
[turbovec/tests/sync_v7.rs](https://github.com/RyanCodrai/turbovec/blob/main/turbovec/tests/sync_v7.rs) Extensive tests covering incremental saves, error handling, and full-write fallback scenarios.
[turbovec-python/src/lib.rs](https://github.com/RyanCodrai/turbovec/blob/main/turbovec-python/src/lib.rs) The Python binding that exposes sync() to Python users.

Summary

  • Incremental saves in turbovec are performed automatically by sync(), which detects dirty slots using an internal generation counter.
  • The first sync() call always writes the entire container; subsequent calls only flush changed slots.
  • Turbovec uses an atomic swap of a temporary file, providing crash safety.
  • Forcing a full rewrite is handled internally when the generation mismatch requires it, so you never need to manage the file format yourself.
  • Both the Python and Rust APIs expose the same sync() method with identical semantics.

Frequently Asked Questions

How do I force a full save instead of an incremental save?

Turbovec does not expose a public flag to force a full rewrite. Instead, if you need a full write, you can create a new empty index and re-insert all vectors, or rely on the automatic fallback that happens when the generation mismatch is detected during the sync.

Is sync() thread-safe?

Yes. The sync method acquires an internal lock on the index, preventing concurrent mutations during the file write operation. This ensures that that no partial state is persisted.

What happens if the process crashes during a sync?

The atomic swap guarantees that either the old file or the new file is present on disk, never a partially written one. If a crash occurs mid-write, the temporary file is simply discarded and the original remains intact and valid.

Does sync() consume much memory?

No. The method streams only the dirty slots from memory to a temporary file. It does not load the entire database into RAM, which makes incremental saves both CPU-light and memory-efficient.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →