How NekoImageGallery's Indexing Pipeline Processes and Vectorizes Images

NekoImageGallery uses an async four-stage pipeline—upload queue, background preparation, dual vector generation via CLIP and optional BERT, and Qdrant persistence—to convert raw images into 768-dimensional searchable embeddings.

The image indexing pipeline in the hv0905/nekoimagegallery repository transforms uploaded files into high-dimensional vectors stored in a Qdrant database. This pipeline supports both visual similarity search through CLIP embeddings and text-based retrieval when OCR extracts readable content from images.

The Four Stages of the Image Indexing Pipeline

The pipeline processes images asynchronously through distinct logical stages, ensuring the event loop remains responsive during heavy ML inference tasks.

Stage 1: Upload and Queue

When an image enters the system, UploadService.queue_upload_image creates a MappedImage Pydantic model and places it on an asyncio.Queue. The queue length is controlled by config.admin_index_queue_max_length to prevent memory overload. A background worker (_upload_worker) continuously pulls items from this queue and forwards them to _upload_task in app/Services/upload_service.py.

Stage 2: Background Preparation

The worker thread invokes IndexService._prepare_image in app/Services/index_service.py to normalize the PIL image. This method records image dimensions, forces RGB conversion, and extracts two potential vectors: a visual vector from the CLIP model and an optional OCR-text vector from a BERT model when text is detected.

Stage 3: Vector Generation

Vector extraction occurs through specialized service methods:

  • Visual embedding: TransformersService.get_image_vector generates a 768-dimensional CLIP embedding from the processed PIL image
  • Text embedding: TransformersService.get_bert_vector converts OCR-extracted text into dense vectors using BERT

Both methods are implemented in app/Services/transformers_service.py and handle model inference asynchronously to prevent blocking.

Stage 4: Persistence to Qdrant

The populated MappedImage model, now containing both metadata and vectors, is upserted into Qdrant via VectorDbContext.insert_items in app/Services/vector_db_context.py. This method constructs Qdrant PointStruct objects from the payload and executes async upserts to store the visual and text vectors for hybrid search capabilities.

Deep Dive into the Pipeline Execution Flow

The indexing workflow begins with a deterministic UUID generation. UploadService.assign_image_id creates identifiers derived from the image's binary digest, guaranteeing repeatable IDs for duplicate detection.

For batch processing, the scripts/local_indexing.py CLI tool enumerates files recursively and processes them concurrently:

files = [p for root in root_directory for p in glob_local_files(root, '**/*')]
await index_task(item, categories, starred, thumbnail_mode)

The UploadService._upload_task opens raw bytes with Pillow, constructs storage URLs (local or S3), optionally generates thumbnails, and delegates heavy processing to IndexService.index_image. This call executes in a thread pool (background=True) to maintain async responsiveness.

Inside IndexService.index_image, the system checks for duplicate IDs (unless disabled via configuration) before calling _prepare_image. This critical method:

  1. Records dimensions and aspect ratio on the MappedImage instance
  2. Forces RGB conversion and copies the image to prevent side effects
  3. Extracts the CLIP embedding via self._transformers_service.get_image_vector(image)
  4. Conditionally runs OCR and BERT encoding if config.ocr_search.enable is active

After vector attachment, IndexService triggers VectorDbContext.insert_items, which builds Qdrant points from the mapped image payload and executes the async upsert operation.

Code Implementation Examples

Batch Indexing via CLI

Run the local indexing script to process entire directories:

python -m scripts.local_indexing \
    /path/to/images \
    --categories "cat,portrait" \
    --starred false \
    --thumbnail-mode always

This creates a ServiceProvider, walks the directory tree, and pushes each file through the complete pipeline.

Programmatic Single-Image Indexing

Index a single PIL.Image directly without using the queue:

from pathlib import Path
from PIL import Image
from app.Services.provider import ServiceProvider
from app.Models.mapped_image import MappedImage
from app.Models.api_models.admin_query_params import UploadImageThumbnailMode
from datetime import datetime
import asyncio

async def index_one(file_path: Path):
    # Initialize service provider

    provider = ServiceProvider()
    await provider.onload()
    
    # Generate deterministic UUID

    img_id = await provider.upload_service.assign_image_id(file_path)
    
    # Create mapped image record

    mapped = MappedImage(
        id=img_id,
        local=True,
        categories=["demo"],
        starred=False,
        format=file_path.suffix[1:],
        index_date=datetime.now(),
    )
    
    # Read and index

    raw = file_path.read_bytes()
    await provider.upload_service.sync_upload_image(
        mapped,
        raw,
        skip_ocr=False,
        thumbnail_mode=UploadImageThumbnailMode.ALWAYS,
    )

asyncio.run(index_one(Path("example.jpg")))

This bypasses the async queue and calls the indexing chain directly: UploadService → IndexService → TransformersService.

Retrieving Stored Vectors

Debug stored embeddings by fetching them from Qdrant:

from app.Services.provider import ServiceProvider
import asyncio

async def show_vectors(img_id: str):
    provider = ServiceProvider()
    await provider.onload()
    img = await provider.db_context.retrieve_by_id(img_id, with_vectors=True)
    print("Visual vector:", img.image_vector[:5], "...")
    if img.text_contain_vector is not None:
        print("OCR vector:", img.text_contain_vector[:5], "...")

asyncio.run(show_vectors("3fa85f64-5717-4562-b3fc-2c963f66afa6"))

Key Components and Source Files

Summary

  • The pipeline uses asyncio queues with configurable length limits to manage backpressure during bulk uploads
  • CLIP models generate 768-dimensional visual embeddings while BERT models encode OCR-extracted text
  • Deterministic UUIDs derived from binary digests prevent duplicate entries and enable idempotent operations
  • Thread pool execution keeps the async event loop responsive during CPU-intensive ML inference
  • Qdrant vector database stores both visual and text vectors, supporting hybrid similarity search
  • The entire workflow is accessible via CLI batch processing or direct Python API calls

Frequently Asked Questions

How does NekoImageGallery prevent duplicate image entries?

The system generates deterministic UUIDs using UploadService.assign_image_id, which derives identifiers from the image's binary digest. Before indexing, IndexService.index_image checks for existing IDs in Qdrant (unless duplicate checking is disabled in configuration), ensuring the same file produces consistent identifiers across multiple indexing runs.

What vector dimensions does the indexing pipeline produce?

According to the source code in app/Services/index_service.py, the CLIP visual embeddings are 768-dimensional vectors. When OCR is enabled and text is detected, BERT generates additional dense embeddings for the extracted text content, enabling multi-modal search across both visual and textual features.

Can I index images without using the async queue?

Yes. While queue_upload_image adds images to the background processing queue, you can call sync_upload_image directly on the UploadService instance. This method bypasses the queue and immediately processes the image through IndexService, suitable for programmatic indexing or real-time applications requiring synchronous completion.

Does the pipeline support cloud storage backends?

Yes. The UploadService._upload_task in app/Services/upload_service.py constructs storage URLs that support both local filesystem and S3-compatible backends. The pipeline reads raw bytes from these sources before vectorization, making it storage-agnostic regarding the final vector generation process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →