# MappedImage Structure and Vector Storage in NekoImageGallery: A Complete Technical Guide

> Understand the MappedImage structure for image metadata and dual 768-dimensional vector storage in NekoImageGallery. Learn how vectors are persisted in Qdrant.

- Repository: [EdgeNeko/nekoimagegallery](https://github.com/hv0905/nekoimagegallery)
- Tags: deep-dive
- Published: 2026-03-03

---

**The `MappedImage` Pydantic model stores image metadata and dual 768-dimensional vectors separately from the payload, with vectors persisted in Qdrant as dense fields while metadata lives in JSON-compatible payloads.**

The `MappedImage` class serves as the core data structure in the NekoImageGallery open-source image search engine, bridging high-dimensional neural embeddings with searchable metadata. According to the hv0905/nekoimagegallery source code, this Pydantic model implements a strict separation between vector data and payload fields to optimize storage and retrieval performance in Qdrant.

## MappedImage Core Structure and Fields

The `MappedImage` definition resides in [`app/Models/mapped_image.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Models/mapped_image.py) (lines 9-34) and extends Pydantic’s `BaseModel` to enforce type safety across the application.

### Primary Metadata Fields

The model captures comprehensive image metadata through the following fields:

- **`id`**: A `UUID` generated from the image digest, serving as the primary key
- **`url`** and **`thumbnail_url`**: Optional strings for remote image locations
- **`ocr_text`**: Optional string containing text extracted via OCR
- **`index_date`**: `datetime` marking when the image was indexed
- **`width`**, **`height`**, **`aspect_ratio`**: Optional dimensional metadata
- **`starred`**: Optional boolean flag for user favorites
- **`categories`**: Optional list of string tags
- **`local`** and **`local_thumbnail`**: Boolean flags indicating local storage status
- **`format`**: Optional string required for S3-compatible storage backends
- **`comments`**: Optional free-form user notes

### Vector Storage Fields

Critically, the model maintains two distinct embedding vectors:

- **`image_vector`**: Optional `ndarray` containing the 768-dimensional vision embedding
- **`text_contain_vector`**: Optional `ndarray` containing the 768-dimensional OCR/text embedding

Both fields store high-dimensional float vectors as NumPy arrays during runtime, though they are intentionally excluded from the serialized payload as implemented in [`app/Models/mapped_image.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Models/mapped_image.py) (lines 45-50).

## How Vectors Are Stored in Qdrant

The separation between vectors and metadata is enforced by the `VectorDbContext` class in [`app/Services/vector_db_context.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/vector_db_context.py). This architecture allows Qdrant to index the dense vectors for similarity search while keeping the payload lightweight.

### Vector Preparation and Point Construction

The private method `_get_point_from_mapped_image` (lines 61-67) constructs a Qdrant `PointStruct` by combining the JSON payload with vector data:

1. **Payload extraction**: Calls `MappedImage.payload` to retrieve all metadata fields excluding vectors
2. **Vector conversion**: Invokes `_get_vector_from_img_data` (lines 50-56) which extracts the two NumPy arrays, converts them to Python lists via `.tolist()`, and returns a dictionary with keys `image_vector` and `text_contain_vector`
3. **Point assembly**: Creates a `PointStruct` with the UUID as the point ID, the filtered payload, and the vector dictionary

### Insertion and Upsert Operations

The `insert_items` method (lines 75-82) handles batch persistence by calling `self._client.upsert` with the prepared `PointStruct` objects. This operation atomically writes both the dense vectors (for neural search) and the JSON payload (for metadata filtering) into Qdrant collections.

### Vector Retrieval and Reconstruction

When reading from Qdrant, the `_get_mapped_image_from_point` method (lines 70-77) performs the inverse transformation:

- Extracts the payload dictionary to populate `MappedImage` metadata fields
- Detects if the point contains a vector dictionary, then converts each stored list back into a NumPy `ndarray` with `dtype=float32`
- Reattaches the reconstructed `image_vector` and `text_contain_vector` to the model instance

This round-trip conversion ensures that vectors remain in 768-dimensional floating-point format suitable for machine learning operations while being serialized as JSON-compatible lists for the vector database.

## Practical Implementation Examples

### Creating a MappedImage with Embeddings

```python
import uuid
import numpy as np
from datetime import datetime
from app.Models.mapped_image import MappedImage

# Generate sample 768-dimensional embeddings

vision_embedding = np.random.rand(768).astype(np.float32)
text_embedding = np.random.rand(768).astype(np.float32)

img = MappedImage(
    id=uuid.uuid4(),
    url="https