# SearchBasisEnum.vision vs SearchBasisEnum.ocr: Visual vs Text Search in NekoImageGallery

> Understand SearchBasisEnum vision vs ocr. Discover how CLIP visual search differs from OCR text search in NekoImageGallery for effective image matching.

- Repository: [EdgeNeko/nekoimagegallery](https://github.com/hv0905/nekoimagegallery)
- Tags: deep-dive
- Published: 2026-03-03

---

**SearchBasisEnum.vision uses CLIP embeddings to match images by visual similarity, while SearchBasisEnum.ocr uses BERT embeddings to match images by their OCR-extracted text content.**

NekoImageGallery is an AI-powered image search engine that supports dual-modal retrieval capabilities. The `SearchBasisEnum` defines how search vectors are generated, determining whether queries match against visual features or textual content extracted via OCR. Understanding the distinction between these modes is essential for optimizing search relevance in the hv0905/nekoimagegallery repository.

## Core Differences Between Vision and OCR Search Modes

The two search modes differ in three fundamental aspects: the embedding model used, the source of the vector data, and the database field queried.

### Vector Sources and Embedding Models

- **SearchBasisEnum.vision**: Generates **image-level CLIP embeddings** using OpenAI's CLIP model (`clip-vit-large-patch14`). For text queries, it uses the CLIP text encoder to create vectors that match visual concepts.

- **SearchBasisEnum.ocr**: Generates **text-level BERT embeddings** using a Chinese BERT model (`bert-base-chinese`). These vectors represent the textual content extracted from images via OCR, not the visual appearance.

### Database Index Fields

According to [`app/Services/vector_db_context.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/vector_db_context.py), each mode maps to a distinct Qdrant vector field:

- **vision** queries the `IMG_VECTOR` field, which stores pre-computed CLIP embeddings for every image
- **ocr** queries the `TEXT_VECTOR` field, which stores BERT embeddings of OCR-extracted text (populated only when OCR is enabled in configuration)

## Implementation in the Source Code

### Enum Definition

The search basis is declared as a string enum in [`app/Models/api_models/search_api_model.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Models/api_models/search_api_model.py):

```python
class SearchBasisEnum(str, Enum):
    vision = "vision"
    ocr = "ocr"

```

### Vector Field Mapping

The `VectorDbContext` class translates the enum into the correct database field. In [`app/Services/vector_db_context.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/vector_db_context.py), the `vector_name_for_basis` method handles this mapping:

```python
@classmethod
def vector_name_for_basis(cls, basis: SearchBasisEnum) -> str:
    match basis:
        case SearchBasisEnum.vision:
            return cls.IMG_VECTOR
        case SearchBasisEnum.ocr:
            return cls.TEXT_VECTOR

```

### Controller Logic and Endpoint Handling

The search controller in [`app/Controllers/search.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Controllers/search.py) selects the appropriate embedding method based on the requested basis. For text searches at `/search/text`, it branches between CLIP and BERT:

```python
text_vector = services.transformers_service.get_text_vector(prompt) \
             if basis.basis == SearchBasisEnum.vision \
             else services.transformers_service.get_bert_vector(prompt)

```

For image searches at `/search/image`, the system always uses **vision** mode because the query is already an image:

```python
image_vector = services.transformers_service.get_image_vector(img)

```

In hybrid search scenarios, the controller automatically swaps the secondary basis to combine both modalities. When the primary basis is `ocr`, the secondary uses `vision` (and vice versa):

```python
match basis.basis:
    case SearchBasisEnum.ocr:
        second_basis = SearchBasisEnum.vision
        second_vector = services.transformers_service.get_text_vector(model.extra_prompt)
    case SearchBasisEnum.vision:
        second_basis = SearchBasisEnum.ocr
        second_vector = services.transformers_service.get_bert_vector(model.extra_prompt)

```

## Practical Usage Examples

### HTTP API Calls

Query the text search endpoint with different basis parameters to switch between visual concept matching and text content matching:

```bash

# Vision-based search: finds images visually similar to "a cute cat"

curl -X POST "http://localhost:8000/api/v1/search/text?basis=vision" \
     -H "Content-Type: application/json" \
     -d '{"prompt":"a cute cat"}'

# OCR-based search: finds images containing Chinese text matching "黑猫"

curl -X POST "http://localhost:8000/api/v1/search/text?basis=ocr" \
     -H "Content-Type: application/json" \
     -d '{"prompt":"黑猫"}'

```

### Python Client Implementation

Use the `requests` library to programmatically select the search mode:

```python
import requests

API = "http://localhost:8000/api/v1/search/text"

def search(prompt: str, basis: str = "vision"):
    resp = requests.post(
        f"{API}?basis={basis}",
        json={"prompt": prompt},
    )
    resp.raise_for_status()
    return resp.json()

# Visual similarity search

print(search("sunset over the ocean", basis="vision"))

# OCR text content search

print(search("海岸线", basis="ocr"))

```

### Internal Service Integration

For direct service access (as used in the test suite), generate vectors explicitly using `TransformersService`:

```python
from app.Services.transformers_service import TransformersService

svc = TransformersService()

# CLIP embedding for visual search

vision_vec = svc.get_text_vector("a red sports car")

# BERT embedding for OCR text search  

ocr_vec = svc.get_bert_vector("红色跑车")

```

## Summary

- **SearchBasisEnum.vision** leverages OpenAI CLIP (`clip-vit-large-patch14`) to match queries against the `IMG_VECTOR` field for visual similarity
- **SearchBasisEnum.ocr** leverages Chinese BERT (`bert-base-chinese`) to match queries against the `TEXT_VECTOR` field for text content similarity
- The enum is defined in [`app/Models/api_models/search_api_model.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Models/api_models/search_api_model.py) and mapped to database fields in [`app/Services/vector_db_context.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Services/vector_db_context.py)
- Controllers in [`app/Controllers/search.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/Controllers/search.py) route requests to `get_text_vector` (CLIP) or `get_bert_vector` (BERT) based on the selected basis
- Image uploads always use vision mode, while text searches can toggle between modes using the `basis` query parameter

## Frequently Asked Questions

### Can I use OCR search if OCR is disabled in the configuration?

No. The `TEXT_VECTOR` field is only populated when OCR is enabled via the `ocr_search.enable` flag in [`app/config.py`](https://github.com/hv0905/nekoimagegallery/blob/main/app/config.py). If disabled, the OCR search basis will return empty results because no text vectors exist in the database.

### Which search mode should I use for non-Chinese text?

While `SearchBasisEnum.ocr` uses `bert-base-chinese` optimized for Chinese characters, it can still process English text. However, for English or multilingual visual concepts, `SearchBasisEnum.vision` typically provides better results because CLIP is trained on diverse multilingual image-text pairs and matches semantic visual concepts rather than literal text strings.

### How does hybrid search handle both SearchBasisEnum values?

Hybrid searches automatically combine both modalities by using your selected basis as the primary search and automatically switching to the opposite basis for the secondary search. If you select `ocr` as the primary basis, the system uses `vision` for the secondary vector comparison, ensuring both visual and textual relevance are considered.

### What embedding function is used when searching by image upload?

Image uploads via `/search/image` always use `SearchBasisEnum.vision` logic. The controller calls `get_image_vector` to generate CLIP embeddings directly from the uploaded image pixels, bypassing the text embedding functions entirely.