How EmbeddingStore Stores Visual Embeddings in Palmier Pro: Binary Format Deep Dive
The EmbeddingStore persists visual embeddings as compact binary files using a custom format with an 8-byte magic header, JSON metadata, and half-precision floating point vectors indexed by timestamps.
The EmbeddingStore in the palmier-pro repository provides a lightweight, file-based caching mechanism for visual embeddings generated by machine learning models. Understanding its binary storage format reveals how the system balances storage efficiency with fast retrieval of per-frame vector data for video assets.
File Location and Cache Key Generation
Each cached embedding file is identified by a deterministic key derived from the asset’s metadata. In Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift, the key(for:) method generates a 32-character prefix of a SHA-256 hash based on the asset’s file path, size, and modification date (lines 36-42).
This key serves as the filename within the application’s cache directory, specifically under …/Caches/<subsystem>/Embeddings as defined by the EmbeddingStore.directory property. This approach ensures that any change to the source video invalidates the cache automatically, while identical files map to the same storage location regardless of their path.
Binary File Format Structure
The on-disk format is deliberately simple yet self-describing, consisting of three distinct sections: a magic identifier, a variable-length JSON header, and a sequence of binary rows containing the actual embedding vectors.
Magic Header and JSON Metadata
Every embedding file begins with a fixed 8-byte magic string PALMEMB1 defined in EmbeddingStore.magic (lines 30-31). This allows the system to quickly validate file integrity before processing.
Immediately following the magic bytes, a 4-byte little-endian unsigned integer specifies the length of the JSON header. The EmbeddingStore.Header struct (lines 7-12) stores critical metadata including:
- The model name and version
- The sampler version
- The embedding dimension
- The total number of vectors stored
This metadata ensures that cached embeddings can be validated against the current model configuration and provides the necessary context to parse the binary data that follows.
Row Structure and Float16 Encoding
Each row corresponds to a video frame or shot boundary and stores temporal information alongside the embedding vector. The EmbeddingStore.Row struct (lines 15-19) defines three 64-bit double values: time, shotStart, and shotEnd.
Following the timestamps, the embedding vector itself is stored as Float16 (half-precision) values to minimize disk usage while maintaining sufficient precision for DSP operations. The complete binary layout follows this sequence:
- Magic (8 bytes):
"PALMEMB1" - Header length (4 bytes):
UInt32specifying JSON byte count - Header JSON (variable): UTF-8 encoded
EmbeddingStore.Header - Rows (repeated):
time (8 bytes) + shotStart (8 bytes) + shotEnd (8 bytes) + dim × Float16 (2 bytes per element)
Reading and Writing Embeddings
The EmbeddingStore exposes two primary interface methods for persistence: save(header:rows:vectors:key:) and load(key:).
Loading Cached Embeddings
The load(key:) method (lines 63-93) validates the magic header, extracts the JSON metadata, then iterates over the remaining bytes to reconstruct the embedding arrays. It decodes timestamps as Double values and converts the half-precision Float16 data to Float arrays suitable for vDSP or other mathematical operations.
do {
let asset = try EmbeddingStore.load(key: key)
print("Loaded \(asset.header.count) vectors of dim \(asset.header.dim)")
// `asset.vectors` is a flat Float array ready for vDSP or other math.
} catch {
print("Failed to load embeddings: \(error)")
}
Saving New Embeddings
The save(header:rows:vectors:key:) method (lines 96-112) assembles the binary blob by writing the magic bytes, header length, JSON payload, and then encoding each row’s timestamps followed by the vector values converted to Float16.
let header = EmbeddingStore.Header(
model: "clip-vit-base-patch32",
modelVersion: 1,
samplerVersion: 2,
dim: 512,
count: rows.count
)
try EmbeddingStore.save(
header: header,
rows: rows, // [EmbeddingStore.Row] with timing info
vectors: embeddingFlat, // [Float] flattened dim-wise
key: key
)
Versioning and Cache Invalidation
To prevent stale embeddings from being reused after model updates, the isCurrent(key:model:modelVersion:samplerVersion:) method compares the stored header metadata against the expected model identifiers. If the model name, version, or sampler version differs, the method returns false, triggering regeneration of the embeddings.
if EmbeddingStore.isCurrent(
key: key,
model: "clip-vit-base-patch32",
modelVersion: 1,
samplerVersion: 2) {
// reuse cached vectors
} else {
// recompute embeddings
}
Summary
- Deterministic key generation uses SHA-256 prefixes based on file attributes to map assets to cache files.
- Self-describing binary format combines an 8-byte magic string (
PALMEMB1), JSON metadata headers, and compactFloat16vector storage. - Temporal indexing stores
time,shotStart, andshotEndas 64-bit doubles alongside each embedding vector. - Automatic invalidation occurs through version checking in
isCurrent(), ensuring model updates trigger cache refreshes. - Memory-efficient storage uses half-precision floating point values that convert to
Floatarrays during loading for downstream processing.
Frequently Asked Questions
What file format does EmbeddingStore use for visual embeddings?
The EmbeddingStore uses a custom binary format rather than standard serialization. Each file starts with the magic bytes PALMEMB1, followed by a length-prefixed JSON header containing model metadata, and concludes with binary-encoded rows of timestamps and half-precision floating point vectors. This format minimizes storage overhead while preserving temporal alignment data.
How does EmbeddingStore generate cache keys for video files?
Cache keys are generated using the key(for:) method, which computes a 32-character prefix of a SHA-256 hash derived from the asset's file path, size, and modification timestamp. This ensures that any change to the source file automatically invalidates the cache, while identical files produce identical keys regardless of their location in the filesystem.
Why does EmbeddingStore use Float16 instead of Float32 for storage?
The system stores embedding vectors as Float16 (half-precision) values to reduce disk space by 50% compared to standard Float32 storage. During the loading process via load(key:), these values are converted back to Float (single-precision) arrays to maintain compatibility with DSP pipelines and mathematical operations that require higher precision.
How does EmbeddingStore handle model updates and cache invalidation?
The isCurrent(key:model:modelVersion:samplerVersion:) method validates cached embeddings against the current model configuration by comparing the stored Header struct values. If the model name, version, or sampler version has changed, the method returns false, signaling that the application should regenerate embeddings using the updated model parameters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →