# How NekoImageGallery's Indexing Pipeline Processes and Vectorizes Images

> Discover how NekoImageGallery's async indexing pipeline processes images through four stages: upload, preparation, CLIP/BERT vector generation, and Qdrant persistence, creating searchable embeddings.

- Repository: [EdgeNeko/nekoimagegallery](https://github.com/hv0905/nekoimagegallery)
- Tags: internals
- Published: 2026-03-03

---

**NekoImageGallery uses an async four-stage pipeline—upload queue, background preparation, dual vector generation via CLIP and optional BERT, and Qdrant persistence—to convert raw images into 768-dimensional searchable embeddings.**

The image indexing pipeline in the [hv0905/nekoimagegallery](https://github.com/hv0905/nekoimagegallery) repository transforms uploaded files into high-dimensional vectors stored in a Qdrant database. This pipeline supports both visual similarity search through CLIP embeddings and text-based retrieval when OCR extracts readable content from images.

## The Four Stages of the Image Indexing Pipeline

The pipeline processes images asynchronously through distinct logical stages, ensuring the event loop remains responsive during heavy ML inference tasks.

### Stage 1: Upload and Queue

When an image enters the system, `UploadService.queue_upload_image` creates a `MappedImage` Pydantic model and places it on an `asyncio.Queue`. The queue length is controlled by `config.admin_index_queue_max_length` to prevent memory overload. A background worker (`_upload_worker`) continuously pulls items from this queue and forwards them to `_upload_task` in [`app/Services/upload_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/upload_service.py).

### Stage 2: Background Preparation

The worker thread invokes `IndexService._prepare_image` in [`app/Services/index_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/index_service.py) to normalize the PIL image. This method records image dimensions, forces RGB conversion, and extracts two potential vectors: a **visual vector** from the CLIP model and an optional **OCR-text vector** from a BERT model when text is detected.

### Stage 3: Vector Generation

Vector extraction occurs through specialized service methods:

- **Visual embedding**: `TransformersService.get_image_vector` generates a 768-dimensional CLIP embedding from the processed PIL image
- **Text embedding**: `TransformersService.get_bert_vector` converts OCR-extracted text into dense vectors using BERT

Both methods are implemented in [`app/Services/transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/transformers_service.py) and handle model inference asynchronously to prevent blocking.

### Stage 4: Persistence to Qdrant

The populated `MappedImage` model, now containing both metadata and vectors, is upserted into Qdrant via `VectorDbContext.insert_items` in [`app/Services/vector_db_context.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/vector_db_context.py). This method constructs Qdrant `PointStruct` objects from the payload and executes async upserts to store the visual and text vectors for hybrid search capabilities.

## Deep Dive into the Pipeline Execution Flow

The indexing workflow begins with a deterministic UUID generation. `UploadService.assign_image_id` creates identifiers derived from the image's binary digest, guaranteeing repeatable IDs for duplicate detection.

For batch processing, the [`scripts/local_indexing.py`](https://github.com/hv0905/nekoimagegallery/blob/main/scripts/local_indexing.py) CLI tool enumerates files recursively and processes them concurrently:

```python
files = [p for root in root_directory for p in glob_local_files(root, '**/*')]
await index_task(item, categories, starred, thumbnail_mode)

```

The `UploadService._upload_task` opens raw bytes with Pillow, constructs storage URLs (local or S3), optionally generates thumbnails, and delegates heavy processing to `IndexService.index_image`. This call executes in a thread pool (`background=True`) to maintain async responsiveness.

Inside `IndexService.index_image`, the system checks for duplicate IDs (unless disabled via configuration) before calling `_prepare_image`. This critical method:

1. Records dimensions and aspect ratio on the `MappedImage` instance
2. Forces RGB conversion and copies the image to prevent side effects
3. Extracts the CLIP embedding via `self._transformers_service.get_image_vector(image)`
4. Conditionally runs OCR and BERT encoding if `config.ocr_search.enable` is active

After vector attachment, `IndexService` triggers `VectorDbContext.insert_items`, which builds Qdrant points from the mapped image payload and executes the async upsert operation.

## Code Implementation Examples

### Batch Indexing via CLI

Run the local indexing script to process entire directories:

```bash
python -m scripts.local_indexing \
    /path/to/images \
    --categories "cat,portrait" \
    --starred false \
    --thumbnail-mode always

```

This creates a `ServiceProvider`, walks the directory tree, and pushes each file through the complete pipeline.

### Programmatic Single-Image Indexing

Index a single `PIL.Image` directly without using the queue:

```python
from pathlib import Path
from PIL import Image
from app.Services.provider import ServiceProvider
from app.Models.mapped_image import MappedImage
from app.Models.api_models.admin_query_params import UploadImageThumbnailMode
from datetime import datetime
import asyncio

async def index_one(file_path: Path):
    # Initialize service provider

    provider = ServiceProvider()
    await provider.onload()
    
    # Generate deterministic UUID

    img_id = await provider.upload_service.assign_image_id(file_path)
    
    # Create mapped image record

    mapped = MappedImage(
        id=img_id,
        local=True,
        categories=["demo"],
        starred=False,
        format=file_path.suffix[1:],
        index_date=datetime.now(),
    )
    
    # Read and index

    raw = file_path.read_bytes()
    await provider.upload_service.sync_upload_image(
        mapped,
        raw,
        skip_ocr=False,
        thumbnail_mode=UploadImageThumbnailMode.ALWAYS,
    )

asyncio.run(index_one(Path("example.jpg")))

```

This bypasses the async queue and calls the indexing chain directly: `UploadService → IndexService → TransformersService`.

### Retrieving Stored Vectors

Debug stored embeddings by fetching them from Qdrant:

```python
from app.Services.provider import ServiceProvider
import asyncio

async def show_vectors(img_id: str):
    provider = ServiceProvider()
    await provider.onload()
    img = await provider.db_context.retrieve_by_id(img_id, with_vectors=True)
    print("Visual vector:", img.image_vector[:5], "...")
    if img.text_contain_vector is not None:
        print("OCR vector:", img.text_contain_vector[:5], "...")

asyncio.run(show_vectors("3fa85f64-5717-4562-b3fc-2c963f66afa6"))

```

## Key Components and Source Files

- **[`app/Services/index_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/index_service.py)**: Orchestrates image preparation, vector extraction, duplicate checks, and database insertion through `index_image` and `_prepare_image`
- **[`app/Services/transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/transformers_service.py)**: Loads CLIP and BERT models, exposing `get_image_vector`, `get_text_vector`, and `get_bert_vector` for embedding generation
- **[`app/Services/upload_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/upload_service.py)**: Manages the async upload queue (`queue_upload_image`), UUID assignment, and thumbnail generation before delegating to indexing
- **[`app/Services/vector_db_context.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/vector_db_context.py)**: Wraps the async Qdrant client, handling `insert_items` for upserts and query operations for hybrid search
- **[`scripts/local_indexing.py`](https://github.com/hv0905/nekoimagegallery/blob/main/scripts/local_indexing.py)**: Convenience CLI for batch directory processing
- **[`app/Models/mapped_image.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Models/mapped_image.py)**: Pydantic model storing metadata, vectors, and payload conversion logic for Qdrant compatibility

## Summary

- The pipeline uses **asyncio queues** with configurable length limits to manage backpressure during bulk uploads
- **CLIP models** generate 768-dimensional visual embeddings while **BERT models** encode OCR-extracted text
- **Deterministic UUIDs** derived from binary digests prevent duplicate entries and enable idempotent operations
- **Thread pool execution** keeps the async event loop responsive during CPU-intensive ML inference
- **Qdrant vector database** stores both visual and text vectors, supporting hybrid similarity search
- The entire workflow is accessible via CLI batch processing or direct Python API calls

## Frequently Asked Questions

### How does NekoImageGallery prevent duplicate image entries?

The system generates deterministic UUIDs using `UploadService.assign_image_id`, which derives identifiers from the image's binary digest. Before indexing, `IndexService.index_image` checks for existing IDs in Qdrant (unless duplicate checking is disabled in configuration), ensuring the same file produces consistent identifiers across multiple indexing runs.

### What vector dimensions does the indexing pipeline produce?

According to the source code in [`app/Services/index_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/index_service.py), the CLIP visual embeddings are **768-dimensional** vectors. When OCR is enabled and text is detected, BERT generates additional dense embeddings for the extracted text content, enabling multi-modal search across both visual and textual features.

### Can I index images without using the async queue?

Yes. While `queue_upload_image` adds images to the background processing queue, you can call `sync_upload_image` directly on the `UploadService` instance. This method bypasses the queue and immediately processes the image through `IndexService`, suitable for programmatic indexing or real-time applications requiring synchronous completion.

### Does the pipeline support cloud storage backends?

Yes. The `UploadService._upload_task` in [`app/Services/upload_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/upload_service.py) constructs storage URLs that support both local filesystem and S3-compatible backends. The pipeline reads raw bytes from these sources before vectorization, making it storage-agnostic regarding the final vector generation process.