How NekoImageGallery's Indexing Pipeline Processes and Vectorizes Images
NekoImageGallery uses an async four-stage pipeline—upload queue, background preparation, dual vector generation via CLIP and optional BERT, and Qdrant persistence—to convert raw images into 768-dimensional searchable embeddings.
The image indexing pipeline in the hv0905/nekoimagegallery repository transforms uploaded files into high-dimensional vectors stored in a Qdrant database. This pipeline supports both visual similarity search through CLIP embeddings and text-based retrieval when OCR extracts readable content from images.
The Four Stages of the Image Indexing Pipeline
The pipeline processes images asynchronously through distinct logical stages, ensuring the event loop remains responsive during heavy ML inference tasks.
Stage 1: Upload and Queue
When an image enters the system, UploadService.queue_upload_image creates a MappedImage Pydantic model and places it on an asyncio.Queue. The queue length is controlled by config.admin_index_queue_max_length to prevent memory overload. A background worker (_upload_worker) continuously pulls items from this queue and forwards them to _upload_task in app/Services/upload_service.py.
Stage 2: Background Preparation
The worker thread invokes IndexService._prepare_image in app/Services/index_service.py to normalize the PIL image. This method records image dimensions, forces RGB conversion, and extracts two potential vectors: a visual vector from the CLIP model and an optional OCR-text vector from a BERT model when text is detected.
Stage 3: Vector Generation
Vector extraction occurs through specialized service methods:
- Visual embedding:
TransformersService.get_image_vectorgenerates a 768-dimensional CLIP embedding from the processed PIL image - Text embedding:
TransformersService.get_bert_vectorconverts OCR-extracted text into dense vectors using BERT
Both methods are implemented in app/Services/transformers_service.py and handle model inference asynchronously to prevent blocking.
Stage 4: Persistence to Qdrant
The populated MappedImage model, now containing both metadata and vectors, is upserted into Qdrant via VectorDbContext.insert_items in app/Services/vector_db_context.py. This method constructs Qdrant PointStruct objects from the payload and executes async upserts to store the visual and text vectors for hybrid search capabilities.
Deep Dive into the Pipeline Execution Flow
The indexing workflow begins with a deterministic UUID generation. UploadService.assign_image_id creates identifiers derived from the image's binary digest, guaranteeing repeatable IDs for duplicate detection.
For batch processing, the scripts/local_indexing.py CLI tool enumerates files recursively and processes them concurrently:
files = [p for root in root_directory for p in glob_local_files(root, '**/*')]
await index_task(item, categories, starred, thumbnail_mode)
The UploadService._upload_task opens raw bytes with Pillow, constructs storage URLs (local or S3), optionally generates thumbnails, and delegates heavy processing to IndexService.index_image. This call executes in a thread pool (background=True) to maintain async responsiveness.
Inside IndexService.index_image, the system checks for duplicate IDs (unless disabled via configuration) before calling _prepare_image. This critical method:
- Records dimensions and aspect ratio on the
MappedImageinstance - Forces RGB conversion and copies the image to prevent side effects
- Extracts the CLIP embedding via
self._transformers_service.get_image_vector(image) - Conditionally runs OCR and BERT encoding if
config.ocr_search.enableis active
After vector attachment, IndexService triggers VectorDbContext.insert_items, which builds Qdrant points from the mapped image payload and executes the async upsert operation.
Code Implementation Examples
Batch Indexing via CLI
Run the local indexing script to process entire directories:
python -m scripts.local_indexing \
/path/to/images \
--categories "cat,portrait" \
--starred false \
--thumbnail-mode always
This creates a ServiceProvider, walks the directory tree, and pushes each file through the complete pipeline.
Programmatic Single-Image Indexing
Index a single PIL.Image directly without using the queue:
from pathlib import Path
from PIL import Image
from app.Services.provider import ServiceProvider
from app.Models.mapped_image import MappedImage
from app.Models.api_models.admin_query_params import UploadImageThumbnailMode
from datetime import datetime
import asyncio
async def index_one(file_path: Path):
# Initialize service provider
provider = ServiceProvider()
await provider.onload()
# Generate deterministic UUID
img_id = await provider.upload_service.assign_image_id(file_path)
# Create mapped image record
mapped = MappedImage(
id=img_id,
local=True,
categories=["demo"],
starred=False,
format=file_path.suffix[1:],
index_date=datetime.now(),
)
# Read and index
raw = file_path.read_bytes()
await provider.upload_service.sync_upload_image(
mapped,
raw,
skip_ocr=False,
thumbnail_mode=UploadImageThumbnailMode.ALWAYS,
)
asyncio.run(index_one(Path("example.jpg")))
This bypasses the async queue and calls the indexing chain directly: UploadService → IndexService → TransformersService.
Retrieving Stored Vectors
Debug stored embeddings by fetching them from Qdrant:
from app.Services.provider import ServiceProvider
import asyncio
async def show_vectors(img_id: str):
provider = ServiceProvider()
await provider.onload()
img = await provider.db_context.retrieve_by_id(img_id, with_vectors=True)
print("Visual vector:", img.image_vector[:5], "...")
if img.text_contain_vector is not None:
print("OCR vector:", img.text_contain_vector[:5], "...")
asyncio.run(show_vectors("3fa85f64-5717-4562-b3fc-2c963f66afa6"))
Key Components and Source Files
app/Services/index_service.py: Orchestrates image preparation, vector extraction, duplicate checks, and database insertion throughindex_imageand_prepare_imageapp/Services/transformers_service.py: Loads CLIP and BERT models, exposingget_image_vector,get_text_vector, andget_bert_vectorfor embedding generationapp/Services/upload_service.py: Manages the async upload queue (queue_upload_image), UUID assignment, and thumbnail generation before delegating to indexingapp/Services/vector_db_context.py: Wraps the async Qdrant client, handlinginsert_itemsfor upserts and query operations for hybrid searchscripts/local_indexing.py: Convenience CLI for batch directory processingapp/Models/mapped_image.py: Pydantic model storing metadata, vectors, and payload conversion logic for Qdrant compatibility
Summary
- The pipeline uses asyncio queues with configurable length limits to manage backpressure during bulk uploads
- CLIP models generate 768-dimensional visual embeddings while BERT models encode OCR-extracted text
- Deterministic UUIDs derived from binary digests prevent duplicate entries and enable idempotent operations
- Thread pool execution keeps the async event loop responsive during CPU-intensive ML inference
- Qdrant vector database stores both visual and text vectors, supporting hybrid similarity search
- The entire workflow is accessible via CLI batch processing or direct Python API calls
Frequently Asked Questions
How does NekoImageGallery prevent duplicate image entries?
The system generates deterministic UUIDs using UploadService.assign_image_id, which derives identifiers from the image's binary digest. Before indexing, IndexService.index_image checks for existing IDs in Qdrant (unless duplicate checking is disabled in configuration), ensuring the same file produces consistent identifiers across multiple indexing runs.
What vector dimensions does the indexing pipeline produce?
According to the source code in app/Services/index_service.py, the CLIP visual embeddings are 768-dimensional vectors. When OCR is enabled and text is detected, BERT generates additional dense embeddings for the extracted text content, enabling multi-modal search across both visual and textual features.
Can I index images without using the async queue?
Yes. While queue_upload_image adds images to the background processing queue, you can call sync_upload_image directly on the UploadService instance. This method bypasses the queue and immediately processes the image through IndexService, suitable for programmatic indexing or real-time applications requiring synchronous completion.
Does the pipeline support cloud storage backends?
Yes. The UploadService._upload_task in app/Services/upload_service.py constructs storage URLs that support both local filesystem and S3-compatible backends. The pipeline reads raw bytes from these sources before vectorization, making it storage-agnostic regarding the final vector generation process.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →