How the CLIP Model Generates Image Vectors for Vision-Based Search in NekoImageGallery

NekoImageGallery uses OpenAI's CLIP model to convert images into 768-dimensional, L2-normalized vectors through preprocessing, feature extraction, and normalization, enabling efficient cosine-similarity search in Qdrant.

NekoImageGallery is an open-source image management system that implements semantic search capabilities through vector embeddings. The repository leverages the openai/clip-vit-large-patch14 model to generate dense vector representations of images, which serve as the search basis for vision-only queries via the /search/image endpoint.

CLIP Vector Generation Pipeline

The transformation of a raw image into a searchable vector occurs in five distinct stages within the TransformersService class.

Model Initialization and Loading

When the application starts, TransformersService loads the CLIP model and processor into memory based on configuration settings defined in app/config.py. The model identifier is specified by config.model.clip (defaulting to openai/clip-vit-large-patch14), and the device is automatically selected between CPU and GPU.


# app/Services/transformers_service.py

self._clip_model = CLIPModel.from_pretrained(config.model.clip).to(self.device)
self._clip_processor = CLIPProcessor.from_pretrained(config.model.clip)

This initialization occurs once at startup, ensuring subsequent vector generation calls operate on pre-loaded model weights.

Image Preprocessing

The get_image_vector method first ensures the input Pillow Image is in RGB mode. It then invokes the CLIPProcessor, which applies the specific normalization parameters required by the CLIP vision encoder:

  • Resize and center-crop to 224×224 pixels
  • Normalize with mean [0.48145466, 0.4578275, 0.40821073] and standard deviation [0.26862954, 0.26130258, 0.27577711]

# app/Services/transformers_service.py

inputs = self._clip_processor(images=image, return_tensors="pt").to(self.device)

The processor returns a PyTorch tensor dictionary moved to the configured device (CPU or CUDA).

Feature Extraction and Normalization

The preprocessed tensor passes through the CLIP vision encoder via the get_image_features method, which outputs a raw 768-dimensional embedding (for the vit-large-patch14 variant).


# app/Services/transformers_service.py

outputs: FloatTensor = self._clip_model.get_image_features(**inputs)

To ensure that cosine similarity calculations reduce to simple dot products, the implementation applies L2 normalization:


# app/Services/transformers_service.py

outputs /= outputs.norm(dim=-1, keepdim=True)

Finally, the normalized torch tensor converts to a NumPy array and flattens to shape (768,):


# app/Services/transformers_service.py

return outputs.numpy(force=True).reshape(-1)

The generated vectors power the /search/image endpoint defined in app/Controllers/search.py. When a user uploads an image file, the controller opens it as a Pillow object and invokes TransformersService.get_image_vector to produce the search vector.


# app/Controllers/search.py (excerpt)

image = Image.open(BytesIO(image_bytes))
image_vector = services.transformers_service.get_image_vector(img)

This vector wraps into a DbQueryCriteriaVector object and queries the Qdrant vector database via VectorDbContext. Qdrant performs a cosine-similarity nearest-neighbor search to retrieve visually similar images.

The identical vector generation path executes during image indexing in app/Services/index_service.py, ensuring that stored embeddings share the same normalization and dimensionality as search queries.

Practical Code Examples

Direct Python Usage

from app.Services.transformers_service import TransformersService
from PIL import Image

svc = TransformersService()

# Load any image

img = Image.open("example.jpg")

# Obtain 768-dim CLIP embedding (unit-norm)

vector = svc.get_image_vector(img)          # → np.ndarray, shape (768,)

print(vector.shape, vector[:5])             # (768,) [0.0123 …]

Calling the Public Image-Search Endpoint

curl -X POST "http://localhost:8000/search/image" \
  -F "image=@cat_0.jpg" \
  -H "Content-Type: multipart/form-data"

The server internally executes get_image_vector and returns the top-k visually similar images.

Unit-Test Verification

def test_get_image_vector(self):
    vec1 = self.transformers_service.get_image_vector(
        Image.open(assets_path / "test_images/cat_0.jpg"))
    vec2 = self.transformers_service.get_image_vector(
        Image.open(assets_path / "test_images/cat_1.jpg"))
    assert vec1.shape == (768,)
    assert vec2.shape == (768,)
    # Similar cats → high cosine similarity

    assert calculate_vectors_cosine(vec1, vec2) > 0.8

Key Implementation Files

File Role Link
app/Services/transformers_service.py Core CLIP & BERT loading + vector generation (get_image_vector, get_text_vector, get_bert_vector) https://github.com/hv0905/nekoimagegallery/blob/master/app/Services/transformers_service.py
app/config.py Holds the model identifier (clip: "openai/clip-vit-large-patch14"), device selection, and OCR toggle https://github.com/hv0905/nekoimagegallery/blob/master/app/config.py
app/Controllers/search.py Exposes /search/image endpoint that consumes the image vector for vision-only search https://github.com/hv0905/nekoimagegallery/blob/master/app/Controllers/search.py
app/Services/index_service.py Generates the vector during indexing of new uploads (stores it in Qdrant) https://github.com/hv0905/nekoimagegallery/blob/master/app/Services/index_service.py
tests/unit/test_transformers_service.py Tests that CLIP vectors are correctly produced and are semantically consistent https://github.com/hv0905/nekoimagegallery/blob/master/tests/unit/test_transformers_service.py
app/util/calculate_vectors_cosine.py Helper for measuring cosine similarity in tests and optional post-processing https://github.com/hv0905/nekoimagegallery/blob/master/app/util/calculate_vectors_cosine.py

Summary

  • CLIP Model: NekoImageGallery uses openai/clip-vit-large-patch14 loaded via TransformersService to generate 768-dimensional dense vectors.
  • Preprocessing: Images are resized, center-cropped, and normalized using CLIP-specific mean and standard deviation values before tensor conversion.
  • Feature Extraction: The get_image_features method produces raw embeddings, which undergo L2 normalization to enable cosine-similarity search via dot product.
  • Search Integration: The /search/image endpoint converts uploaded images to vectors using get_image_vector, then queries Qdrant for nearest neighbors.
  • Extensibility: Developers can swap the CLIP checkpoint by modifying config.model.clip without changing the vector generation logic.

Frequently Asked Questions

What is the dimensionality of CLIP vectors in NekoImageGallery?

The system generates 768-dimensional vectors when using the default openai/clip-vit-large-patch14 model. This shape is hardcoded in the model architecture and is preserved through the L2 normalization and NumPy conversion steps in transformers_service.py.

How does the system ensure consistent similarity scoring across different images?

After extracting raw features via CLIPModel.get_image_features, the implementation applies L2 normalization by dividing the vector by its Euclidean norm (outputs /= outputs.norm(dim=-1, keepdim=True)). This ensures all vectors have unit length, making cosine similarity mathematically equivalent to a simple dot product and guaranteeing consistent scoring regardless of input image size or brightness.

Can I replace the default CLIP model with a different vision encoder?

Yes. The model identifier is configurable via config.model.clip in app/config.py. You can specify any Hugging Face Transformers-compatible CLIP checkpoint (such as openai/clip-vit-base-patch16 for smaller vectors or domain-specific variants) and the TransformersService will automatically load it during initialization, provided the model follows the CLIP architecture for get_image_features.

Where is the image vector generation triggered during a search request?

The generation occurs in the /search/image endpoint defined in app/Controllers/search.py. When a user uploads an image file, the controller reads the bytes into a Pillow Image object and calls services.transformers_service.get_image_vector(img). The resulting NumPy array is then wrapped in a DbQueryCriteriaVector and passed to the Qdrant client to perform the nearest-neighbor lookup.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →