MappedImage Structure and Vector Storage in NekoImageGallery: A Complete Technical Guide

The MappedImage Pydantic model stores image metadata and dual 768-dimensional vectors separately from the payload, with vectors persisted in Qdrant as dense fields while metadata lives in JSON-compatible payloads.

The MappedImage class serves as the core data structure in the NekoImageGallery open-source image search engine, bridging high-dimensional neural embeddings with searchable metadata. According to the hv0905/nekoimagegallery source code, this Pydantic model implements a strict separation between vector data and payload fields to optimize storage and retrieval performance in Qdrant.

MappedImage Core Structure and Fields

The MappedImage definition resides in app/Models/mapped_image.py (lines 9-34) and extends Pydantic’s BaseModel to enforce type safety across the application.

Primary Metadata Fields

The model captures comprehensive image metadata through the following fields:

  • id: A UUID generated from the image digest, serving as the primary key
  • url and thumbnail_url: Optional strings for remote image locations
  • ocr_text: Optional string containing text extracted via OCR
  • index_date: datetime marking when the image was indexed
  • width, height, aspect_ratio: Optional dimensional metadata
  • starred: Optional boolean flag for user favorites
  • categories: Optional list of string tags
  • local and local_thumbnail: Boolean flags indicating local storage status
  • format: Optional string required for S3-compatible storage backends
  • comments: Optional free-form user notes

Vector Storage Fields

Critically, the model maintains two distinct embedding vectors:

  • image_vector: Optional ndarray containing the 768-dimensional vision embedding
  • text_contain_vector: Optional ndarray containing the 768-dimensional OCR/text embedding

Both fields store high-dimensional float vectors as NumPy arrays during runtime, though they are intentionally excluded from the serialized payload as implemented in app/Models/mapped_image.py (lines 45-50).

How Vectors Are Stored in Qdrant

The separation between vectors and metadata is enforced by the VectorDbContext class in app/Services/vector_db_context.py. This architecture allows Qdrant to index the dense vectors for similarity search while keeping the payload lightweight.

Vector Preparation and Point Construction

The private method _get_point_from_mapped_image (lines 61-67) constructs a Qdrant PointStruct by combining the JSON payload with vector data:

  1. Payload extraction: Calls MappedImage.payload to retrieve all metadata fields excluding vectors
  2. Vector conversion: Invokes _get_vector_from_img_data (lines 50-56) which extracts the two NumPy arrays, converts them to Python lists via .tolist(), and returns a dictionary with keys image_vector and text_contain_vector
  3. Point assembly: Creates a PointStruct with the UUID as the point ID, the filtered payload, and the vector dictionary

Insertion and Upsert Operations

The insert_items method (lines 75-82) handles batch persistence by calling self._client.upsert with the prepared PointStruct objects. This operation atomically writes both the dense vectors (for neural search) and the JSON payload (for metadata filtering) into Qdrant collections.

Vector Retrieval and Reconstruction

When reading from Qdrant, the _get_mapped_image_from_point method (lines 70-77) performs the inverse transformation:

  • Extracts the payload dictionary to populate MappedImage metadata fields
  • Detects if the point contains a vector dictionary, then converts each stored list back into a NumPy ndarray with dtype=float32
  • Reattaches the reconstructed image_vector and text_contain_vector to the model instance

This round-trip conversion ensures that vectors remain in 768-dimensional floating-point format suitable for machine learning operations while being serialized as JSON-compatible lists for the vector database.

Practical Implementation Examples

Creating a MappedImage with Embeddings

import uuid
import numpy as np
from datetime import datetime
from app.Models.mapped_image import MappedImage

# Generate sample 768-dimensional embeddings

vision_embedding = np.random.rand(768).astype(np.float32)
text_embedding = np.random.rand(768).astype(np.float32)

img = MappedImage(
    id=uuid.uuid4(),
    url="https

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →