Building Visual Search with MLX Embeddings and Vector Similarity in Palmier Pro
Palmier Pro implements visual search by converting image frames and text queries into 768-dimensional vectors using on-device MLX models, then computing cosine similarity against cached embeddings for fully offline search.
The palmier-io/palmier-pro repository provides a complete macOS-native implementation for indexing visual media and performing semantic search using Apple's MLX framework. This article examines the Swift implementation that embeds images and text into shareable vector spaces, persists them to disk, and executes high-performance similarity queries without network dependencies.
Architecture Overview
The visual search pipeline in Palmier Pro consists of four coordinated stages: MLX runtime management, embedding generation, asset indexing, and similarity search. Each stage is isolated into distinct Swift components that handle specific responsibilities while maintaining thread safety across Metal GPU operations.
MLX Runtime and Concurrency Management
The MLXRuntime class in Sources/PalmierPro/Utilities/MLXRuntime.swift centralizes the on-device MLX lifecycle and prevents concurrent Metal kernel crashes through a global operation gate. When the app initializes, it creates an MLXOperationGate (defined in Sources/PalmierPro/Utilities/MLXOperationGate.swift) that serializes all model inferences using a semaphore-like counter.
The runtime provides termination handlers to ensure graceful shutdown:
MLXRuntime.beginTermination()blocks new inference requestsMLXRuntime.waitUntilIdle()waits for the current gate to clear
If the bundled mlx.metallib is missing, the runtime throws an Unavailable error, preventing initialization of dependent components.
Generating Image and Text Embeddings
The VisualEmbedder class in Sources/PalmierPro/Search/Models/VisualEmbedder.swift wraps dual Core ML models—an image encoder and a text encoder—to produce normalized 768-dimensional vectors. It coordinates with ModelDownloader to resolve model binaries and specifications.
Key methods:
encode(image:)– Converts aCGImageto a BGRACVPixelBuffersized according tospec.imageSize, runs the image encoder, and extracts the"embedding"multi-arrayencode(text:)– Tokenizes the string usingTextTokenizer, feeds tokens to the text encoder, and returns a[Float]array
Both methods return flat [Float] arrays of length 768 (configurable via SearchIndexConfig), compatible with cosine similarity comparisons.
Indexing and Persistence
VisualIndexer in Sources/PalmierPro/Search/Indexing/VisualIndexer.swift orchestrates the extraction of embeddings from media assets. For video files, it uses a FrameSampler (defaulting to one frame per second) to select representative stills; for static images, it uses the single available frame.
The indexer writes vectors to EmbeddingStore in Sources/PalmierPro/Search/Indexing/EmbeddingStore.swift, which provides disk-based caching keyed by:
- File identity hash
- Model version
- Embedding dimension
This design makes indexing idempotent—re-indexing the same asset with unchanged parameters skips redundant computation.
Implementing Similarity Search
When executing a query, Palmier Pro loads stored vectors from EmbeddingStore.shared.loadAllVectors() and computes similarity against the query embedding using a pure Swift implementation of cosine similarity:
func cosine(_ a: [Float], _ b: [Float]) -> Float {
let dot = zip(a, b).map(*).reduce(0, +)
let normA = sqrt(a.map { $0*$0 }.reduce(0, +))
let normB = sqrt(b.map { $0*$0 }.reduce(0, +))
return dot / (normA * normB)
}
The SearchIndexCoordinator sorts results by descending score and returns the top-N matches to the UI layer. This entire pipeline operates on macOS 26 arm64 devices without requiring network latency or cloud API calls.
Code Examples
Initializing the Embedder
Load the bundled model specification and instantiate the embedder with both encoder URLs:
import PalmierPro
import CoreGraphics
// Load spec from the bundled downloader (populated at first launch)
let spec = try ModelDownloader.shared.loadSpec(for: .visual)
// Initialize with image and text encoder URLs
let embedder = try VisualEmbedder(
imageEncoderURL: spec.imageEncoderURL,
textEncoderURL: spec.textEncoderURL,
tokenizer: TextTokenizer(),
spec: spec,
computeUnits: .all
)
// Generate embedding from a CGImage thumbnail
let cgImage: CGImage = // obtain from AVAsset or NSImage
let embedding = try embedder.encode(image: cgImage) // → [Float] (768 elements)
Indexing a Media Asset
Create a VisualIndexer with a frame sampling strategy and populate the on-disk cache:
let asset = MediaAsset(url: videoURL)
let indexer = VisualIndexer(
embedder: embedder,
sampler: FrameSampler(interval: .seconds(1)),
store: EmbeddingStore.shared
)
// Idempotent operation: writes to EmbeddingStore
try await indexer.index(asset: asset)
Querying with Image Similarity
Compute cosine similarity between a query image and all indexed vectors:
let queryImage: CGImage = // user-selected image
let queryVector = try embedder.encode(image: queryImage)
// Load cached vectors (binary format for performance)
let allVectors = EmbeddingStore.shared.loadAllVectors()
// Compute scores and sort
let results = allVectors
.map { (asset: $0.asset, score: cosine($0.vector, queryVector)) }
.sorted { $0.score > $1.score }
.prefix(5)
Text-to-Image Search
Use the text encoder to search for images matching a natural language description:
let queryText = "sunset over a lake"
let textVector = try embedder.encode(text: queryText)
// Compare against stored image embeddings using the same cosine function
let textResults = allVectors
.map { (asset: $0.asset, score: cosine($0.vector, textVector)) }
.sorted { $0.score > $1.score }
Key Configuration and Model Management
The SearchIndexConfig struct in Sources/PalmierPro/Search/SearchIndexConfig.swift centralizes index parameters, defaulting to 768 dimensions to match the SigLIP-style model architecture. The ModelDownloader class resolves version-specific specs including embeddingDim, imageSize, and contextLength, ensuring the embedder receives compatible model binaries.
Summary
- MLXRuntime and MLXOperationGate serialize Metal kernel usage to prevent GPU contention and crashes during concurrent inference.
- VisualEmbedder provides unified
encode(image:)andencode(text:)methods that return 768-dimensional[Float]vectors using Core ML models. - VisualIndexer samples video frames and persists embeddings to
EmbeddingStore, which caches vectors by file hash and model version for idempotent operations. - Cosine similarity is computed in pure Swift against loaded vectors, enabling offline text-to-image and image-to-image search.
- The architecture targets macOS 26 arm64 devices with zero network dependencies after initial model download.
Frequently Asked Questions
How does Palmier Pro prevent Metal GPU crashes during concurrent embedding generation?
The MLXOperationGate class in Sources/PalmierPro/Utilities/MLXOperationGate.swift implements an @unchecked Sendable semaphore pattern that allows only one inference operation at a time. MLXRuntime manages this gate globally, blocking new requests during app termination and throwing Unavailable errors if the MLX Metal library is missing.
What format do the embeddings use and how are they stored?
Embeddings are flat [Float] arrays of length 768 (configurable via SearchIndexConfig). The EmbeddingStore persists these as binary data on disk, keyed by file identity hash, model version, and dimension. This allows the VisualIndexer to skip re-processing assets that have already been indexed with the same parameters.
Can I search using text queries even though the index contains only image embeddings?
Yes. The VisualEmbedder uses separate but aligned encoders for images and text—both projecting into the same 768-dimensional vector space. When you call embedder.encode(text:), the resulting vector can be compared directly against stored image embeddings using cosine similarity, enabling natural language search across visual content.
What is the minimum macOS version required for MLX embedding support?
According to the source implementation, Palmier Pro targets macOS 26 arm64 devices with Apple Silicon. The MLX runtime requires the mlx.metallib Metal library bundled with the application, and the VisualEmbedder requires Core ML compute units available only on these newer systems.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →