# How TransformersService Loads and Manages CLIP and BERT Models in NekoImageGallery

> Discover how NekoImageGallery's TransformersService loads and manages CLIP and BERT models for efficient vector extraction. Learn about its startup initialization and attribute storage.

- Repository: [EdgeNeko/nekoimagegallery](https://github.com/hv0905/nekoimagegallery)
- Tags: internals
- Published: 2026-03-03

---

**The `TransformersService` initializes CLIP and conditional BERT models on the configured device during startup, storing them as private attributes for efficient vector extraction across image and text inputs.**

NekoImageGallery is an open-source image management system that relies on transformer models to power semantic search capabilities. The `TransformersService` acts as the central inference engine, handling the lifecycle of CLIP and BERT models while providing standardized methods for converting images and text into comparable vector embeddings. Understanding how this service loads and manages these models is essential for customizing deployment settings or optimizing resource usage.

## Device Selection and Initialization

The service begins by determining the compute target through `config.device`, which defaults to *auto* selection. When set to *auto*, the implementation checks for CUDA availability and falls back to CPU if no GPU is present, ensuring portability across different hardware configurations.

This device selection occurs in [`app/Services/transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/transformers_service.py) lines 17–20, where the logic inspects `torch.cuda.is_available()` before finalizing the target device. The chosen device string is stored for subsequent model loading operations.

## Loading the CLIP Model for Image-Text Similarity

CLIP (Contrastive Language-Image Pre-training) serves as the primary backbone for cross-modal similarity search within the gallery.

### Model Configuration

The specific CLIP variant is controlled by `config.model.clip`, which defaults to `openai/clip-vit-large-patch14` as defined in [`app/config.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/config.py) lines 30–33. This configuration key allows operators to swap in different CLIP architectures without modifying source code.

### Processor and Model Instantiation

The service loads the model weights using `CLIPModel.from_pretrained` and immediately moves the parameters to the previously selected device. Simultaneously, `CLIPProcessor.from_pretrained` instantiates the matching tokenizer and image preprocessor. Both objects are cached as private instance attributes `_clip_model` and `_clip_processor` to eliminate repeated loading overhead during inference requests.

```python

# From app/Services/transformers_service.py

_clip_model = CLIPModel.from_pretrained(config.model.clip).to(self._device)
_clip_processor = CLIPProcessor.from_pretrained(config.model.clip)

```

This initialization sequence appears in lines 22–24 of the service implementation.

## Conditional BERT Loading for OCR Search

Unlike CLIP, which loads unconditionally, BERT initialization depends on feature flags to conserve memory when OCR capabilities are not required.

### OCR Configuration Check

The service inspects `config.ocr_search.enable` before attempting to load BERT. When this boolean is `False`, the service skips BERT initialization entirely, leaving `_bert_model` and `_bert_tokenizer` undefined. This conditional logic prevents wasted GPU memory on deployments that do not require Chinese text OCR capabilities.

### BERT Model Initialization

When OCR search is enabled, the service loads `bert-base-chinese` (or the configured alternative) using `BertModel.from_pretrained` and `BertTokenizer.from_pretrained`. These are stored in `_bert_model` and `_bert_tokenizer` respectively, as shown in lines 25–30 of [`transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/transformers_service.py).

## Vector Extraction Methods

The service exposes three primary vectorization methods that wrap the loaded models with preprocessing and postprocessing logic.

### Image Vector Generation with CLIP

The `get_image_vector` method handles the complete pipeline from raw image to normalized embedding. It converts input images to RGB format, processes them through `_clip_processor`, extracts features via `_clip_model.get_image_features`, applies L2 normalization, and returns a NumPy array. This implementation occupies lines 33–43 and produces 768-dimensional vectors matching the default CLIP model's output size.

```python
from PIL import Image

service = TransformersService()
img = Image.open("photo.jpg")
vector = service.get_image_vector(img)  # Returns ndarray shape (768,)

```

### Text Vector Generation with CLIP

For text-based image retrieval, `get_text_vector` follows an analogous flow using `_clip_model.get_text_features`. The method tokenizes input strings through the shared `_clip_processor`, runs inference on the target device, normalizes the resulting tensor, and converts it to a NumPy array (lines 46–54).

### Text Vector Generation with BERT

When OCR search is active, `get_bert_vector` provides Chinese-language text embeddings. The method tokenizes input using `_bert_tokenizer`, forwards through `_bert_model`, averages the hidden states across the sequence dimension to create a fixed-size representation, and returns the vector (lines 56–64). This averaging strategy captures semantic meaning while maintaining consistent output dimensions regardless of input length.

## Lifecycle Management and Utilities

The service inherits from `LifespanService`, defined in [`app/Services/lifespan_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/lifespan_service.py), which provides empty `on_load` and `on_exit` hooks for future resource cleanup implementations. This architectural choice allows for graceful model unloading or memory pool reset when the application shuts down.

For testing and fallback scenarios, the class provides a static utility `get_random_vector` that generates dummy 768-dimensional vectors using seeded random number generation (lines 66–70).

```python

# Generate a deterministic random vector for testing

dummy_vec = TransformersService.get_random_vector(seed=42)

```

## Summary

- **Device auto-detection**: The service automatically selects CUDA when available, falling back to CPU, based on `config.device` settings.
- **CLIP always loads**: The CLIP model and processor initialize unconditionally from `config.model.clip` and cache as `_clip_model` and `_clip_processor`.
- **BERT conditional loading**: BERT only loads when `config.ocr_search.enable` is `True`, storing instances in `_bert_model` and `_bert_tokenizer`.
- **Unified vector interface**: Three methods—`get_image_vector`, `get_text_vector`, and `get_bert_vector`—handle preprocessing, inference, normalization, and format conversion.
- **Resource efficiency**: By loading models once during initialization and reusing instances, the service minimizes I/O overhead and GPU memory fragmentation.

## Frequently Asked Questions

### How does TransformersService handle GPU availability?

The service reads the `device` configuration value in [`app/Services/transformers_service.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/transformers_service.py) lines 17–20. When set to *auto* (the default), it checks `torch.cuda.is_available()` and selects **CUDA** if true, otherwise **CPU**. Explicit values like *cuda:0* or *cpu* bypass auto-detection and force the specified target.

### Can I use a different CLIP model than the default?

Yes. Modify the `model.clip` value in [`app/config.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/config.py) (lines 30–33) to specify any Hugging Face CLIP model identifier. The service passes this string directly to `CLIPModel.from_pretrained` and `CLIPProcessor.from_pretrained` during initialization.

### Why does BERT not load even though I have a GPU?

BERT initialization depends on the `ocr_search.enable` configuration flag. If this boolean is `False` in your configuration, the service intentionally skips loading `bert-base-chinese` to conserve memory. Set `ocr_search.enable` to `True` and restart the service to activate BERT embeddings.

### What is the output dimension of the generated vectors?

All vector methods return 768-dimensional NumPy arrays. This matches the hidden size of the default `openai/clip-vit-large-patch14` and `bert-base-chinese` models. The `get_random_vector` static method also generates 768-dimensional vectors for consistency with the production embedding space.