# How Palmier Pro Visual Search Indexes Video Content: Technical Architecture Explained

> Discover how Palmier Pro indexes video content with its technical architecture: frame sampling, Core ML encoding, vector caching, and accelerated searches for visual similarity. Learn more.

- Repository: [Palmier/palmier-pro](https://github.com/palmier-io/palmier-pro)
- Tags: architecture
- Published: 2026-06-23

---

**Palmier Pro indexes video content by sampling representative frames, encoding them into compact embeddings using Core ML models, caching the vectors to disk, and performing accelerated dot-product searches to find visually similar moments.**

Palmier Pro's visual search engine transforms video libraries into instantly searchable databases by analyzing frame-level content rather than relying solely on metadata. According to the palmier-io/palmier-pro source code, the system implements a multi-stage pipeline that extracts embeddings from video frames and stores them in a specialized binary format for millisecond-scale retrieval. This article examines the actual Swift implementation to explain how the Palmier Pro Visual Search index video architecture balances accuracy, storage efficiency, and search performance.

## The Three-Stage Indexing Pipeline

The indexing process follows a strict separation of concerns across three distinct components, each handling a specific transformation in the data pipeline.

### Frame Sampling with FrameSampler

The process begins in [`Sources/PalmierPro/Search/Indexing/FrameSampler.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Indexing/FrameSampler.swift), where the `FrameSampler` class extracts representative frames from video assets. Rather than processing every frame—which would be computationally prohibitive—the sampler extracts one frame per detected shot, plus optional additional frames for comprehensive coverage.

For each sampled frame, the system records:
- The exact **timestamp** within the video
- The **shot start and end times** defining the containing scene
- The **shot index** for grouping related frames

This shot-aware sampling ensures that the visual index maintains semantic boundaries, allowing searches to return the most relevant moment within a scene rather than arbitrary frames.

### Embedding Generation via VisualEmbedder

Once frames are sampled, [`Sources/PalmierPro/Search/Models/VisualEmbedder.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Models/VisualEmbedder.swift) handles the conversion of pixel data into searchable vectors. The `VisualEmbedder` loads a Core ML image encoder model (such as SigLIP) and processes each `CGImage` through a conversion pipeline:

1. The image is converted to a pixel buffer via `pixelBuffer(from:size:)`
2. The buffer feeds into the Core ML model
3. The resulting **embedding vector** is extracted from the "embedding" output using `vector(from:dim:)`

The vector dimensions (typically 512 in the reference implementation) are defined in the model's `Spec` configuration, ensuring consistency between indexing and query operations.

### Disk Caching in EmbeddingStore

To avoid re-processing video files on every app launch, [`Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift) persists embeddings to a binary cache. The store generates a unique key for each video file based on its size, modification time, and file path, then writes a binary `.embed` file containing:

- A **JSON header** (`Header`) with model metadata (name, version, sampler version), embedding dimension, and vector count
- A flat array of **Row structs** capturing `time`, `shotStart`, and `shotEnd` for each frame
- The embedding vectors themselves, stored as **Float16** for storage efficiency

The `save(header:rows:vectors:key:)` method writes this data atomically to `EmbeddingStore.directory`. When loading, `load(key:)` reconstructs an `AssetIndex` containing the header, rows, and a Float-32 copy of the vectors optimized for fast arithmetic operations.

## Orchestrating the Pipeline: VisualIndexer

The `VisualIndexer.index(...)` method in [`Sources/PalmierPro/Search/Indexing/VisualIndexer.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Indexing/VisualIndexer.swift) coordinates the entire indexing workflow:

```swift
let spec = model.spec
guard VisualIndexer.needsIndex(url: url, spec: spec) else { return }
for await frame in FrameSampler.frames(url: url, duration: duration) {
    vectors += try model.encode(image: frame.image)          // embedding generation
    times.append(frame.time)                                 // timestamp recording
    shotIndices.append(shotStarts.count - 1)
}
let rows = zip(times, shotIndices).map { … }              // Row construction
try save(rows: rows, vectors: vectors, spec: spec, key: key) // disk persistence

```

This async method first checks whether the video requires indexing (comparing the cached `.embed` file against current specifications), then iterates through sampled frames, accumulating vectors and metadata before triggering the atomic save operation.

## Querying the Index: VisualSearch Implementation

When users initiate a search, [`Sources/PalmierPro/Search/Query/VisualSearch.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Query/VisualSearch.swift) executes a high-performance similarity computation. The `VisualSearch.search` method receives a query vector—generated by the same `VisualEmbedder` used during indexing—and performs a **dense matrix-vector multiply** using the Accelerate framework's `cblas_sgemv` BLAS routine.

The search algorithm:
1. Loads the stored `AssetIndex` for each video asset
2. Computes dot-product similarity between the query vector and all stored frame embeddings
3. Reduces results to the best frame per shot to avoid duplicate recommendations
4. Filters by a relative cutoff score and optional minimum threshold
5. Returns top hits as `VisualSearch.Hit` structs containing asset ID, timestamp, shot range, and confidence score

This approach leverages CPU-optimized linear algebra to search thousands of frame embeddings in milliseconds without GPU requirements.

## Practical Implementation Examples

### Indexing a Video Asset

The following Swift code demonstrates the complete indexing workflow for a video file:

```swift
import PalmierPro

let videoURL = URL(fileURLWithPath: "/path/to/clip.mov")
let duration = 120.0                     // video length in seconds
let embedder = try VisualEmbedder(
    imageEncoderURL: Bundle.main.url(forResource: "siglip_image", withExtension: "mlmodelc")!,
    textEncoderURL: Bundle.main.url(forResource: "siglip_text", withExtension: "mlmodelc")!,
    tokenizer: TextTokenizer(),
    spec: VisualEmbedder.Spec(
        model: "siglip", version: 1, embeddingDim: 512,
        imageSize: 224, contextLength: 77
    )
)

// Asynchronously index (runs on a background actor)
Task {
    try await VisualIndexer.index(
        url: videoURL,
        duration: duration,
        model: embedder,
        progress: { fraction in
            print("Indexing progress: \(fraction * 100)%")
        })
}

```

### Searching for Visually Similar Frames

To query the index with a reference image:

```swift
import PalmierPro

// Assume we already have a query image (e.g. a thumbnail from the UI)
let queryImage: CGImage = …  

// Encode the query using the same model that was used for indexing
let queryVector = try embedder.encode(image: queryImage)

// Load cached indexes for all assets we want to search
let assetIndexes: [(String, EmbeddingStore.AssetIndex)] = try loadAllIndexes()

// Perform the search – returns the best hits across assets
let hits = VisualSearch.search(
    query: queryVector,
    indexes: assetIndexes,
    limit: 10,
    relativeCutoff: 0.85)

// Use the hits (asset ID, timestamp, shot range, score) to jump to the video
for hit in hits {
    print("\(hit.assetID) @ \(hit.time)s (score \(hit.score))")
}

```

### Loading Stored Indexes

This utility function loads all cached indexes from the embedding directory:

```swift
func loadAllIndexes() throws -> [(String, EmbeddingStore.AssetIndex)] {
    let dir = EmbeddingStore.directory
    let files = try FileManager.default.contentsOfDirectory(at: dir,
                                                            includingPropertiesForKeys: nil)
    return try files.compactMap { url in
        guard url.pathExtension == "embed",
              let key = url.deletingPathExtension().lastPathComponent as String?,
              let index = try? EmbeddingStore.load(key: key) else { return nil }
        // The asset ID is stored elsewhere; for demo purposes we use the file name
        return (url.lastPathComponent, index)
    }
}

```

## Summary

- **Shot-aware sampling**: The `FrameSampler` extracts one frame per shot from [`Sources/PalmierPro/Search/Indexing/FrameSampler.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Indexing/FrameSampler.swift), preserving scene boundaries while minimizing data volume.
- **Core ML embedding**: `VisualEmbedder` in [`Sources/PalmierPro/Search/Models/VisualEmbedder.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Models/VisualEmbedder.swift) converts frames to 512-dimensional vectors using models like SigLIP.
- **Compact binary storage**: `EmbeddingStore` persists vectors as Float16 in `.embed` files with JSON headers, enabling instant reloading via [`Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift).
- **Accelerated search**: `VisualSearch.search` uses `cblas_sgemv` from the Accelerate framework to perform dense matrix-vector multiplication for sub-millisecond similarity queries.
- **Per-shot deduplication**: Results are grouped by shot index to return the single best match per scene rather than redundant frames.

## Frequently Asked Questions

### What embedding models does Palmier Pro Visual Search support?

The implementation in [`Sources/PalmierPro/Search/Models/VisualEmbedder.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Models/VisualEmbedder.swift) supports any Core ML image encoder model that provides a named "embedding" output. The reference implementation uses SigLIP models (specifically `siglip_image.mlmodelc` and `siglip_text.mlmodelc`) with 512-dimensional embeddings and 224×224 input resolution, though the architecture allows swapping alternative vision models by updating the `Spec` configuration.

### How does Palmier Pro detect shot boundaries during indexing?

Shot detection occurs within `FrameSampler` in [`Sources/PalmierPro/Search/Indexing/FrameSampler.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Indexing/FrameSampler.swift). The implementation analyzes frame differences to identify shot transitions, then samples one representative frame per detected shot. This ensures that the visual index captures distinct scenes without storing redundant frames from the same continuous shot, significantly reducing storage requirements while maintaining search coverage.

### Why does Palmier Pro store embeddings as Float16 instead of Float32?

The `EmbeddingStore` in [`Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift) writes vectors as Float16 to reduce disk footprint by 50% compared to Float32 storage. When loading via `load(key:)`, the system automatically converts these back to Float32 in memory for optimal performance during the matrix-vector multiplication phase. This hybrid approach balances storage efficiency with computational speed.

### How fast is visual search in Palmier Pro?

Search performance relies on the Accelerate framework's `cblas_sgemv` implementation for dense matrix-vector multiplication, as implemented in [`Sources/PalmierPro/Search/Query/VisualSearch.swift`](https://github.com/palmier-io/palmier-pro/blob/main/Sources/PalmierPro/Search/Query/VisualSearch.swift). Because the operation uses optimized CPU BLAS routines rather than neural network inference, the system can score thousands of frame embeddings against a query vector in milliseconds, delivering real-time results even on large video libraries without requiring GPU activation.