How Finestore Manages Long-Term Artifact Storage and Lazy Artifact Retrieval in Marin
Finestore implements durable artifact persistence through immutable Parquet archives with append-only transactional writes, while enabling efficient lazy retrieval via snapshot-isolated ReadViews and an in-process PersistentKvCache that defers remote I/O.
Finestore is the storage engine within the marin-community/marin repository designed for machine learning workflows requiring both durable versioning and high-performance access. The system separates write-heavy ingestion from read-heavy retrieval by combining an append-only transactional writer with snapshot-based readers and a tiered caching layer, ensuring artifacts remain immutable and compact over the long term without sacrificing query performance.
Long-Term Artifact Storage Architecture
Finestore persists artifacts—including binary blobs, tables, and schema—within a FineStore archive, a directory containing an immutable manifest and a collection of Parquet shards. This design prioritizes durability and immutability while allowing efficient compaction of historical data.
Archive Initialization and Structure
When a writer initializes storage via DataStore.open(root, …), the constructor builds a FineStoreLayout for the specified root path and calls initialize_archive(self._layout) to prepare the archive directory. According to the implementation in lib/finestore/src/finestore/store.py, this establishes the foundation for all subsequent write operations by creating the necessary metadata structures and shard directories.
Transactional Writes and Buffering
Artifacts are written through DataStore.write_object(name, data, metadata), which creates a blob descriptor and optional chunked parts, appending them to the internal blobs table. Writes are buffered in-memory per-table to optimize throughput. When the buffer exceeds max_buffer_bytes or an explicit flush() is invoked, the buffered rows convert into a ShardWriter and commit as a level-zero shard.
The commit sequence operates through the CommitCoordinator, recording metadata in the manifest to ensure the archive remains readable and consistent. This append-only approach prevents in-place modifications, preserving the history of artifact versions within lib/finestore/src/finestore/store.py.
Background Maintenance and Compaction
Long-term storage efficiency relies on a background maintenance thread (self._thread) that regularly executes self.maintain(). This process compacts over-full shards and optionally seals the archive, transforming multiple small shards into fewer, larger files to improve read performance. The maintenance routine defined in lib/finestore/src/finestore/store.py ensures that storage remains compact and immutable over extended periods without blocking active write operations.
Lazy Artifact Retrieval Mechanisms
Retrieval operations in Finestore avoid eager materialization of entire archives. Instead, the system uses snapshot isolation and streaming iterators to access only the necessary data segments.
Snapshot-Based ReadViews
Marin reads artifacts through the read-only ReadView class, which pins a snapshot of the manifest at construction time. According to lib/finestore/src/finestore/reader.py, this snapshot isolation guarantees consistent reads even as concurrent writes append new shards to the archive.
The ReadView.point(table, **keys) and ReadView.iter_rows(...) methods resolve the latest committed rows without loading the complete archive. These iterators stream rows from relevant shards while applying push-down filters and deduplication on the fly, minimizing memory overhead for large datasets.
Streaming Blob Access
For named binary objects, ReadView.read_blob(name) opens a forward-only _BlobReader stream. This implementation in lib/finestore/src/finestore/reader.py yields inline data or ordered chunked parts only when the consumer reads them, enabling lazy loading of large model checkpoints or datasets without holding the entire artifact in memory.
PersistentKvCache Memory Tier
The PersistentKvCache provides an in-process memory tier that bridges hot in-memory access with cold archive storage. As implemented in lib/finestore/src/finestore/cache.py, the cache operates with distinct read and write paths:
Lazy Reads: The load(key) method first checks the local dictionary; on a cache miss, it pins a ReadView and reads the blob from the archive, subsequently storing the result in memory for future fast hits.
Asynchronous Writes: The store(key, value) method writes immediately to the local memory dict. When configured for remote persistence (via marin_temp_bucket), the operation queues the write via _queue_remote_write. A daemon thread (_drain_remote_writes) drains this queue in the background, batching writes into an unbounded transaction and committing them to the FineStore archive in a single flush. This ensures the archive remains eventually consistent while the application experiences instant availability from the memory cache.
Implementation Examples
The following examples demonstrate the interaction between long-term storage and lazy retrieval:
# Writing artifacts to long-term storage
from finestore.store import DataStore
store = DataStore.open("/mnt/runs/run-001") # Initialize archive
uri = store.write_object("checkpoint.pt", model_bytes) # Create blob descriptor
store.flush() # Commit level-zero shard via CommitCoordinator
# Lazy retrieval without full archive scan
from finestore.reader import ReadView
view = ReadView("/mnt/runs/run-001") # Pin manifest snapshot
data = view.read_blob("checkpoint.pt") # Stream bytes on demand
# Using PersistentKvCache for lazy reads and async writes
from finestore.cache import PersistentKvCache
# Writer process
writer_cache = PersistentKvCache.for_prefix("experiment-1", is_writer=lambda: True)
writer_cache.store("metrics", metrics_blob) # Local memory + queued remote write
# Reader process
reader_cache = PersistentKvCache.for_prefix("experiment-1")
result = reader_cache.load("metrics") # Memory hit or lazy archive read
reader_cache.close() # Drain pending writes
Summary
- Finestore stores artifacts in immutable Parquet archives with append-only transactional writes, ensuring durable, versioned storage without in-place modifications.
- The
DataStoreclass manages buffered writes, shard creation, and background compaction viaCommitCoordinatorand maintenance threads defined inlib/finestore/src/finestore/store.py. - Lazy retrieval relies on
ReadViewsnapshots that pin manifests at construction time, streaming only required rows or blob chunks via iterators inlib/finestore/src/finestore/reader.py. - The
PersistentKvCacheprovides tiered access with in-process memory caching, lazy archive reads on misses, and asynchronous background writes for remote persistence inlib/finestore/src/finestore/cache.py.
Frequently Asked Questions
How does Finestore ensure data durability during writes?
Finestore ensures durability through an append-only transaction model. When DataStore.flush() is called, buffered rows convert to a ShardWriter and commit as a level-zero shard through the CommitCoordinator, which atomically records the operation in the immutable manifest. This guarantees that once a write completes, the data remains permanently stored in the Parquet shards regardless of subsequent operations.
What triggers background maintenance in Finestore?
A dedicated background thread (self._thread) runs self.maintain() at regular intervals to compact over-full shards and optionally seal the archive. This maintenance prevents the accumulation of small fragmentary shards by merging them into larger, more efficient files, ensuring long-term storage remains optimized for both space and read performance without blocking active writes.
How does ReadView provide consistent reads during concurrent writes?
ReadView implements snapshot isolation by pinning the manifest at construction time. This creates a consistent point-in-time view of the archive, allowing readers to call point(), iter_rows(), or read_blob() without seeing partial or concurrent modifications. The iterator streams from shards referenced in the pinned manifest while applying filters, ensuring deterministic results even as new shards are appended.
Can PersistentKvCache operate without remote storage?
Yes, PersistentKvCache functions effectively as a local in-memory cache even without remote storage configuration. When marin_temp_bucket is not specified, store() operations populate only the local dictionary, and load() operations serve memory hits or read directly from the local FineStore archive. The background write queue and _drain_remote_writes thread activate only when remote persistence is configured, making the cache adaptable to both single-node and distributed deployments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →