Palmier Pro Visual Search Architecture: Three-Layer Pipeline Explained

Palmier Pro’s visual search architecture implements a three-layer pipeline that converts video and image assets into Float16 embeddings, persists them as binary .embed files keyed by file metadata, and executes vector similarity searches using Accelerate framework dot-products with best-per-shot deduplication.

The visual search system in palmier-io/palmier-pro transforms on-device media libraries into semantically searchable databases without requiring cloud processing. By combining a custom binary storage format, background indexing queues, and optimized matrix operations, the architecture delivers fast text-to-visual matching across entire projects.

The Three-Layer Pipeline

Palmier Pro’s visual search architecture separates concerns into distinct layers: durable embedding storage, asynchronous index creation, and real-time search coordination. Each layer operates through specific Swift classes that manage the transformation of raw pixels into searchable vectors.

Embedding Storage with EmbeddingStore

The EmbeddingStore class, implemented in Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift, defines the persistent storage format for visual embeddings. Rather than using a database, the system serializes vectors directly to disk as .embed binary files.

Each .embed file begins with a magic header, followed by a JSON-encoded Header containing the model name, version, dimension, and vector count. The payload stores temporal metadata as Float64 values representing (time, shotStart, shotEnd) tuples, with the actual embedding vectors compressed as Float16. This format minimizes disk usage while preserving millisecond-precision timing data.

The store generates stable cache keys via EmbeddingStore.key(for:) by hashing the asset’s file path, size, and modification date. This enables idempotent indexing—unchanged files reuse existing embeddings while modified assets trigger automatic re-indexing. The class exposes EmbeddingStore.load(key:) to hydrate indexes from disk and EmbeddingStore.save(...) to atomically write new embeddings.

Index Creation with VisualIndexer

VisualIndexer, located in Sources/PalmierPro/Search/Indexing/VisualIndexer.swift, orchestrates the conversion of media assets into embedding vectors. For each input file, the indexer invokes FrameSampler to extract representative frames, detects shot boundaries to segment the timeline, and passes images through VisualEmbedder.encode(image) to generate L2-normalized float vectors.

The indexer maintains parallel arrays: one containing the raw embedding vectors, and another of EmbeddingStore.Row objects tracking time codes and shot boundaries. After processing all frames, VisualIndexer delegates to EmbeddingStore.save to persist the compact binary representation under the asset’s stable key.

Static utility methods VisualIndexer.needsIndex(url:spec:) and VisualIndexer.indexImage(url:model:) provide synchronous checks and specialized handling for still images versus video assets, allowing the coordinator to skip unnecessary computation when existing .embed files remain valid.

Search Coordination with SearchIndexCoordinator

The SearchIndexCoordinator in Sources/PalmierPro/Search/SearchIndexCoordinator.swift acts as a singleton-style entry point that bridges the UI and indexing layers. It maintains an internal queue of assets requiring (re)indexing and dispatches background workers to invoke VisualIndexer without blocking the main thread.

When search(query:limit:within:) receives a text query, the coordinator first ensures the visual model is loaded via VisualModelLoader.swift. It then gathers candidate assets—respecting the optional within scope—and loads their indexes from disk using EmbeddingStore.load(key:) into AssetIndex objects that are cached per project.

The coordinator transforms the query text into a vector using VisualEmbedder.encode(text:), then delegates the computationally intensive ranking to VisualSearch.search. Additional responsibilities include handling export-pause logic to prevent indexing during critical operations and reporting progress fractions to the UI.

Similarity Ranking with VisualSearch

VisualSearch, defined in Sources/PalmierPro/Search/Query/VisualSearch.swift, implements the core similarity computation using the Accelerate framework. The VisualSearch.search method accepts a query vector and an array of (assetID, AssetIndex) tuples.

Using cblas_sgemv, it performs batch dot-products between the query vector and every stored embedding. Because the model L2-normalizes all vectors, these dot-products directly represent cosine similarity scores ranging from -1 to 1.

To prevent result flooding from multiple frames within the same scene, the algorithm applies bestPerShot filtering, retaining only the highest-scoring frame per shot boundary. It then sorts candidates by score, applies an absolute minScore threshold, and filters using a relativeCutoff of 0.85 to discard low-relevance matches. The final list is trimmed to the requested limit and returned as Hit objects containing asset IDs, timestamps, and similarity scores.

Data Flow and Binary Storage Format

The complete data flow illustrates how media traverses the three-layer architecture:


Asset (video/image)
   │
   ├─► FrameSampler → frames (time, image, shot flag)
   │
   └─► VisualEmbedder.encode(image) → float vector
          │
          └─► VisualIndexer → rows + vectors → EmbeddingStore.save → .embed file
               │
               └─► SearchIndexCoordinator loads AssetIndex on demand
                    │
                    └─► VisualSearch.search (dot-product, best-per-shot, cutoff) → Hit list

The EmbeddingStore binary format ensures efficient retrieval by storing header metadata separately from dense vector data. The Float16 quantization reduces storage footprint by 50% compared to Float32, while the magic header and JSON schema versioning enable forward compatibility as the model architecture evolves.

Implementation Examples

From a SwiftUI view or view model, invoke the coordinator to search across the entire project:

// Assume visual model preparation is handled by the app lifecycle
Task {
    let hits = await SearchIndexCoordinator.shared.search(
        query: "sunset over ocean",
        limit: 10,
        within: nil               // nil searches the whole project
    )
    for hit in hits {
        print("\(hit.assetID) @ \(hit.time)s → score \(hit.score)")
    }
}

Low-Level Search with Custom Indexes

For specialized use cases, manually load indexes and invoke the ranking engine directly:

let queryVector: [Float] = try visualEmbedder.encode(text: "mountain trail")
let assetIdx = try EmbeddingStore.load(key: someKey)      // load specific asset index
let results = VisualSearch.search(
    query: queryVector,
    indexes: [("myAsset", assetIdx)],
    limit: 5,
    relativeCutoff: 0.80,
    minScore: 0.4
)

for r in results {
    print("\(r.assetID) – \(r.time)s – score: \(r.score)")
}

Conditional Re-indexing

Force re-indexing only when necessary by checking the asset status:

if VisualIndexer.needsIndex(url: fileURL, spec: visualModel.spec) {
    try await VisualIndexer.index(
        url: fileURL,
        duration: assetDuration,
        model: visualModel,
        progress: { fraction in
            print("Indexing progress: \(fraction * 100)%")
        }
    )
}

Summary

  • Palmier Pro’s visual search architecture separates concerns into storage (EmbeddingStore), indexing (VisualIndexer), and search (SearchIndexCoordinator, VisualSearch) layers.
  • Binary persistence uses custom .embed files with Float16 vectors and Float64 temporal metadata, keyed by file path, size, and modification date for automatic cache invalidation.
  • Similarity computation leverages the Accelerate framework’s cblas_sgemv for optimized dot-products, with L2-normalized vectors enabling cosine similarity search without explicit normalization.
  • Result deduplication applies bestPerShot filtering and a default relativeCutoff of 0.85 to ensure diverse, high-quality results across different scenes.
  • Idempotent indexing guarantees that unchanged media files reuse existing embeddings while modified assets trigger surgical re-indexing via the coordinator’s background queue.

Frequently Asked Questions

How does Palmier Pro ensure that visual search indexes stay synchronized with media files?

The system uses the EmbeddingStore.key(for:) method to generate stable identifiers based on the asset’s file path, size, and modification date. When SearchIndexCoordinator encounters an asset, it compares the current file metadata against the cached key. If they match, the existing .embed file is loaded; if not, VisualIndexer generates fresh embeddings. This guarantees that renamed files trigger re-indexing while trivial metadata changes do not.

What similarity metric does Palmier Pro use for visual search ranking?

Palmier Pro uses cosine similarity computed via dot-products. Because the VisualEmbedder L2-normalizes all vectors during encoding, the dot-product calculated by cblas_sgemv in VisualSearch.search directly equals the cosine of the angle between vectors. Scores range from -1 to 1, with higher values indicating greater semantic similarity between the text query and visual content.

How does the system prevent duplicate results from the same video scene?

The VisualSearch class implements bestPerShot filtering during the ranking phase. After computing similarity scores for all frames, the algorithm examines the shotStart and shotEnd metadata stored in the EmbeddingStore.Row objects. It retains only the highest-scoring frame within each shot boundary, ensuring that the final result set contains diverse scenes rather than consecutive frames from the same moment.

What is the binary format of Palmier Pro's embedding storage files?

The .embed files use a custom binary protocol defined in Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift. The file begins with a magic header for format identification, followed by a JSON-encoded Header containing model metadata, dimensions, and vector count. The body stores temporal data as Float64 (timestamp, shot start, shot end) and embedding vectors as Float16 to optimize storage density while maintaining search accuracy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →