How Palmier Pro Visual Search Indexes Video Content: Technical Architecture Explained
Palmier Pro indexes video content by sampling representative frames, encoding them into compact embeddings using Core ML models, caching the vectors to disk, and performing accelerated dot-product searches to find visually similar moments.
Palmier Pro's visual search engine transforms video libraries into instantly searchable databases by analyzing frame-level content rather than relying solely on metadata. According to the palmier-io/palmier-pro source code, the system implements a multi-stage pipeline that extracts embeddings from video frames and stores them in a specialized binary format for millisecond-scale retrieval. This article examines the actual Swift implementation to explain how the Palmier Pro Visual Search index video architecture balances accuracy, storage efficiency, and search performance.
The Three-Stage Indexing Pipeline
The indexing process follows a strict separation of concerns across three distinct components, each handling a specific transformation in the data pipeline.
Frame Sampling with FrameSampler
The process begins in Sources/PalmierPro/Search/Indexing/FrameSampler.swift, where the FrameSampler class extracts representative frames from video assets. Rather than processing every frame—which would be computationally prohibitive—the sampler extracts one frame per detected shot, plus optional additional frames for comprehensive coverage.
For each sampled frame, the system records:
- The exact timestamp within the video
- The shot start and end times defining the containing scene
- The shot index for grouping related frames
This shot-aware sampling ensures that the visual index maintains semantic boundaries, allowing searches to return the most relevant moment within a scene rather than arbitrary frames.
Embedding Generation via VisualEmbedder
Once frames are sampled, Sources/PalmierPro/Search/Models/VisualEmbedder.swift handles the conversion of pixel data into searchable vectors. The VisualEmbedder loads a Core ML image encoder model (such as SigLIP) and processes each CGImage through a conversion pipeline:
- The image is converted to a pixel buffer via
pixelBuffer(from:size:) - The buffer feeds into the Core ML model
- The resulting embedding vector is extracted from the "embedding" output using
vector(from:dim:)
The vector dimensions (typically 512 in the reference implementation) are defined in the model's Spec configuration, ensuring consistency between indexing and query operations.
Disk Caching in EmbeddingStore
To avoid re-processing video files on every app launch, Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift persists embeddings to a binary cache. The store generates a unique key for each video file based on its size, modification time, and file path, then writes a binary .embed file containing:
- A JSON header (
Header) with model metadata (name, version, sampler version), embedding dimension, and vector count - A flat array of Row structs capturing
time,shotStart, andshotEndfor each frame - The embedding vectors themselves, stored as Float16 for storage efficiency
The save(header:rows:vectors:key:) method writes this data atomically to EmbeddingStore.directory. When loading, load(key:) reconstructs an AssetIndex containing the header, rows, and a Float-32 copy of the vectors optimized for fast arithmetic operations.
Orchestrating the Pipeline: VisualIndexer
The VisualIndexer.index(...) method in Sources/PalmierPro/Search/Indexing/VisualIndexer.swift coordinates the entire indexing workflow:
let spec = model.spec
guard VisualIndexer.needsIndex(url: url, spec: spec) else { return }
for await frame in FrameSampler.frames(url: url, duration: duration) {
vectors += try model.encode(image: frame.image) // embedding generation
times.append(frame.time) // timestamp recording
shotIndices.append(shotStarts.count - 1)
}
let rows = zip(times, shotIndices).map { … } // Row construction
try save(rows: rows, vectors: vectors, spec: spec, key: key) // disk persistence
This async method first checks whether the video requires indexing (comparing the cached .embed file against current specifications), then iterates through sampled frames, accumulating vectors and metadata before triggering the atomic save operation.
Querying the Index: VisualSearch Implementation
When users initiate a search, Sources/PalmierPro/Search/Query/VisualSearch.swift executes a high-performance similarity computation. The VisualSearch.search method receives a query vector—generated by the same VisualEmbedder used during indexing—and performs a dense matrix-vector multiply using the Accelerate framework's cblas_sgemv BLAS routine.
The search algorithm:
- Loads the stored
AssetIndexfor each video asset - Computes dot-product similarity between the query vector and all stored frame embeddings
- Reduces results to the best frame per shot to avoid duplicate recommendations
- Filters by a relative cutoff score and optional minimum threshold
- Returns top hits as
VisualSearch.Hitstructs containing asset ID, timestamp, shot range, and confidence score
This approach leverages CPU-optimized linear algebra to search thousands of frame embeddings in milliseconds without GPU requirements.
Practical Implementation Examples
Indexing a Video Asset
The following Swift code demonstrates the complete indexing workflow for a video file:
import PalmierPro
let videoURL = URL(fileURLWithPath: "/path/to/clip.mov")
let duration = 120.0 // video length in seconds
let embedder = try VisualEmbedder(
imageEncoderURL: Bundle.main.url(forResource: "siglip_image", withExtension: "mlmodelc")!,
textEncoderURL: Bundle.main.url(forResource: "siglip_text", withExtension: "mlmodelc")!,
tokenizer: TextTokenizer(),
spec: VisualEmbedder.Spec(
model: "siglip", version: 1, embeddingDim: 512,
imageSize: 224, contextLength: 77
)
)
// Asynchronously index (runs on a background actor)
Task {
try await VisualIndexer.index(
url: videoURL,
duration: duration,
model: embedder,
progress: { fraction in
print("Indexing progress: \(fraction * 100)%")
})
}
Searching for Visually Similar Frames
To query the index with a reference image:
import PalmierPro
// Assume we already have a query image (e.g. a thumbnail from the UI)
let queryImage: CGImage = …
// Encode the query using the same model that was used for indexing
let queryVector = try embedder.encode(image: queryImage)
// Load cached indexes for all assets we want to search
let assetIndexes: [(String, EmbeddingStore.AssetIndex)] = try loadAllIndexes()
// Perform the search – returns the best hits across assets
let hits = VisualSearch.search(
query: queryVector,
indexes: assetIndexes,
limit: 10,
relativeCutoff: 0.85)
// Use the hits (asset ID, timestamp, shot range, score) to jump to the video
for hit in hits {
print("\(hit.assetID) @ \(hit.time)s (score \(hit.score))")
}
Loading Stored Indexes
This utility function loads all cached indexes from the embedding directory:
func loadAllIndexes() throws -> [(String, EmbeddingStore.AssetIndex)] {
let dir = EmbeddingStore.directory
let files = try FileManager.default.contentsOfDirectory(at: dir,
includingPropertiesForKeys: nil)
return try files.compactMap { url in
guard url.pathExtension == "embed",
let key = url.deletingPathExtension().lastPathComponent as String?,
let index = try? EmbeddingStore.load(key: key) else { return nil }
// The asset ID is stored elsewhere; for demo purposes we use the file name
return (url.lastPathComponent, index)
}
}
Summary
- Shot-aware sampling: The
FrameSamplerextracts one frame per shot fromSources/PalmierPro/Search/Indexing/FrameSampler.swift, preserving scene boundaries while minimizing data volume. - Core ML embedding:
VisualEmbedderinSources/PalmierPro/Search/Models/VisualEmbedder.swiftconverts frames to 512-dimensional vectors using models like SigLIP. - Compact binary storage:
EmbeddingStorepersists vectors as Float16 in.embedfiles with JSON headers, enabling instant reloading viaSources/PalmierPro/Search/Indexing/EmbeddingStore.swift. - Accelerated search:
VisualSearch.searchusescblas_sgemvfrom the Accelerate framework to perform dense matrix-vector multiplication for sub-millisecond similarity queries. - Per-shot deduplication: Results are grouped by shot index to return the single best match per scene rather than redundant frames.
Frequently Asked Questions
What embedding models does Palmier Pro Visual Search support?
The implementation in Sources/PalmierPro/Search/Models/VisualEmbedder.swift supports any Core ML image encoder model that provides a named "embedding" output. The reference implementation uses SigLIP models (specifically siglip_image.mlmodelc and siglip_text.mlmodelc) with 512-dimensional embeddings and 224×224 input resolution, though the architecture allows swapping alternative vision models by updating the Spec configuration.
How does Palmier Pro detect shot boundaries during indexing?
Shot detection occurs within FrameSampler in Sources/PalmierPro/Search/Indexing/FrameSampler.swift. The implementation analyzes frame differences to identify shot transitions, then samples one representative frame per detected shot. This ensures that the visual index captures distinct scenes without storing redundant frames from the same continuous shot, significantly reducing storage requirements while maintaining search coverage.
Why does Palmier Pro store embeddings as Float16 instead of Float32?
The EmbeddingStore in Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift writes vectors as Float16 to reduce disk footprint by 50% compared to Float32 storage. When loading via load(key:), the system automatically converts these back to Float32 in memory for optimal performance during the matrix-vector multiplication phase. This hybrid approach balances storage efficiency with computational speed.
How fast is visual search in Palmier Pro?
Search performance relies on the Accelerate framework's cblas_sgemv implementation for dense matrix-vector multiplication, as implemented in Sources/PalmierPro/Search/Query/VisualSearch.swift. Because the operation uses optimized CPU BLAS routines rather than neural network inference, the system can score thousands of frame embeddings against a query vector in milliseconds, delivering real-time results even on large video libraries without requiring GPU activation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →