# How the CLIP Model Generates Image Vectors for Vision-Based Search in NekoImageGallery

> Learn how the CLIP model generates 768-dimensional image vectors for NekoImageGallery's vision-based search using preprocessing, feature extraction, and normalization for Qdrant. Explore efficient cosine-similarity search.

- Repository: [EdgeNeko/nekoimagegallery](https://github.com/hv0905/nekoimagegallery)
- Tags: deep-dive
- Published: 2026-03-03

---

**NekoImageGallery uses OpenAI's CLIP model to convert images into 768-dimensional, L2-normalized vectors through preprocessing, feature extraction, and normalization, enabling efficient cosine-similarity search in Qdrant.**

NekoImageGallery is an open-source image management system that implements semantic search capabilities through vector embeddings. The repository leverages the `openai/clip-vit-large-patch14` model to generate dense vector representations of images, which serve as the search basis for vision-only queries via the `/search/image` endpoint.

## CLIP Vector Generation Pipeline

The transformation of a raw image into a searchable vector occurs in five distinct stages within the `TransformersService` class.

### Model Initialization and Loading

When the application starts, `TransformersService` loads the CLIP model and processor into memory based on configuration settings defined in [`app/config.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/config.py). The model identifier is specified by `config.model.clip` (defaulting to `openai/clip-vit-large-patch14`), and the device is automatically selected between CPU and GPU.

```python

# app/Services/transformers_service.py

self._clip_model = CLIPModel.from_pretrained(config.model.clip).to(self.device)
self._clip_processor = CLIPProcessor.from_pretrained(config.model.clip)

```

This initialization occurs once at startup, ensuring subsequent vector generation calls operate on pre-loaded model weights.

### Image Preprocessing

The `get_image_vector` method first ensures the input Pillow `Image` is in RGB mode. It then invokes the `CLIPProcessor`, which applies the specific normalization parameters required by the CLIP vision encoder:

- **Resize and center-crop** to 224×224 pixels
- **Normalize** with mean `[0.48145466, 0.4578275, 0.40821073]` and standard deviation `[0.26862954, 0.26130258, 0.27577711]`

```python

# app/Services/transformers_service.py

inputs = self._clip_processor(images=image, return_tensors="pt").to(self.device)

```

The processor returns a PyTorch tensor dictionary moved to the configured device (CPU or CUDA).

### Feature Extraction and Normalization

The preprocessed tensor passes through the CLIP vision encoder via the `get_image_features` method, which outputs a raw 768-dimensional embedding (for the `vit-large-patch14` variant).

```python

# app/Services/transformers_service.py

outputs: FloatTensor = self._clip_model.get_image_features(**inputs)

```

To ensure that cosine similarity calculations reduce to simple dot products, the implementation applies **L2 normalization**:

```python

# app/Services/transformers_service.py

outputs /= outputs.norm(dim=-1, keepdim=True)

```

Finally, the normalized torch tensor converts to a NumPy array and flattens to shape `(768,)`:

```python

# app/Services/transformers_service.py

return outputs.numpy(force=True).reshape(-1)

```

## Integration with Vision-Based Search

The generated vectors power the `/search/image` endpoint defined in [`app/Controllers/search.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Controllers/search.py). When a user uploads an image file, the controller opens it as a Pillow object and invokes `TransformersService.get_image_vector` to produce the search vector.

```python

# app/Controllers/search.py (excerpt)

image = Image.open(BytesIO(image_bytes))
image_vector = services.transformers_service.get_image_vector(img)

```

This vector wraps into a `DbQueryCriteriaVector` object and queries the Qdrant vector database via `VectorDbContext`. Qdrant performs a cosine-similarity nearest-neighbor search to retrieve visually similar images.

The identical vector generation path executes during image indexing in [`app/Services/index_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/index_service.py), ensuring that stored embeddings share the same normalization and dimensionality as search queries.

## Practical Code Examples

### Direct Python Usage

```python
from app.Services.transformers_service import TransformersService
from PIL import Image

svc = TransformersService()

# Load any image

img = Image.open("example.jpg")

# Obtain 768-dim CLIP embedding (unit-norm)

vector = svc.get_image_vector(img)          # → np.ndarray, shape (768,)

print(vector.shape, vector[:5])             # (768,) [0.0123 …]

```

### Calling the Public Image-Search Endpoint

```bash
curl -X POST "http://localhost:8000/search/image" \
  -F "image=@cat_0.jpg" \
  -H "Content-Type: multipart/form-data"

```

The server internally executes `get_image_vector` and returns the top-k visually similar images.

### Unit-Test Verification

```python
def test_get_image_vector(self):
    vec1 = self.transformers_service.get_image_vector(
        Image.open(assets_path / "test_images/cat_0.jpg"))
    vec2 = self.transformers_service.get_image_vector(
        Image.open(assets_path / "test_images/cat_1.jpg"))
    assert vec1.shape == (768,)
    assert vec2.shape == (768,)
    # Similar cats → high cosine similarity

    assert calculate_vectors_cosine(vec1, vec2) > 0.8

```

## Key Implementation Files

| File | Role | Link |
|------|------|------|
| **[`app/Services/transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/transformers_service.py)** | Core CLIP & BERT loading + vector generation (`get_image_vector`, `get_text_vector`, `get_bert_vector`) | <https://github.com/hv0905/nekoimagegallery/blob/master/app/Services/transformers_service.py> |
| **[`app/config.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/config.py)** | Holds the model identifier (`clip: "openai/clip-vit-large-patch14"`), device selection, and OCR toggle | <https://github.com/hv0905/nekoimagegallery/blob/master/app/config.py> |
| **[`app/Controllers/search.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Controllers/search.py)** | Exposes `/search/image` endpoint that consumes the image vector for vision-only search | <https://github.com/hv0905/nekoimagegallery/blob/master/app/Controllers/search.py> |
| **[`app/Services/index_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/index_service.py)** | Generates the vector during indexing of new uploads (stores it in Qdrant) | <https://github.com/hv0905/nekoimagegallery/blob/master/app/Services/index_service.py> |
| **[`tests/unit/test_transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/tests/unit/test_transformers_service.py)** | Tests that CLIP vectors are correctly produced and are semantically consistent | <https://github.com/hv0905/nekoimagegallery/blob/master/tests/unit/test_transformers_service.py> |
| **[`app/util/calculate_vectors_cosine.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/util/calculate_vectors_cosine.py)** | Helper for measuring cosine similarity in tests and optional post-processing | <https://github.com/hv0905/nekoimagegallery/blob/master/app/util/calculate_vectors_cosine.py> |

## Summary

- **CLIP Model**: NekoImageGallery uses `openai/clip-vit-large-patch14` loaded via `TransformersService` to generate 768-dimensional dense vectors.
- **Preprocessing**: Images are resized, center-cropped, and normalized using CLIP-specific mean and standard deviation values before tensor conversion.
- **Feature Extraction**: The `get_image_features` method produces raw embeddings, which undergo L2 normalization to enable cosine-similarity search via dot product.
- **Search Integration**: The `/search/image` endpoint converts uploaded images to vectors using `get_image_vector`, then queries Qdrant for nearest neighbors.
- **Extensibility**: Developers can swap the CLIP checkpoint by modifying `config.model.clip` without changing the vector generation logic.

## Frequently Asked Questions

### What is the dimensionality of CLIP vectors in NekoImageGallery?

The system generates **768-dimensional** vectors when using the default `openai/clip-vit-large-patch14` model. This shape is hardcoded in the model architecture and is preserved through the L2 normalization and NumPy conversion steps in [`transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/transformers_service.py).

### How does the system ensure consistent similarity scoring across different images?

After extracting raw features via `CLIPModel.get_image_features`, the implementation applies **L2 normalization** by dividing the vector by its Euclidean norm (`outputs /= outputs.norm(dim=-1, keepdim=True)`). This ensures all vectors have unit length, making cosine similarity mathematically equivalent to a simple dot product and guaranteeing consistent scoring regardless of input image size or brightness.

### Can I replace the default CLIP model with a different vision encoder?

Yes. The model identifier is configurable via `config.model.clip` in [`app/config.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/config.py). You can specify any Hugging Face Transformers-compatible CLIP checkpoint (such as `openai/clip-vit-base-patch16` for smaller vectors or domain-specific variants) and the `TransformersService` will automatically load it during initialization, provided the model follows the CLIP architecture for `get_image_features`.

### Where is the image vector generation triggered during a search request?

The generation occurs in the `/search/image` endpoint defined in [`app/Controllers/search.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Controllers/search.py). When a user uploads an image file, the controller reads the bytes into a Pillow `Image` object and calls `services.transformers_service.get_image_vector(img)`. The resulting NumPy array is then wrapped in a `DbQueryCriteriaVector` and passed to the Qdrant client to perform the nearest-neighbor lookup.