SearchBasisEnum.vision vs SearchBasisEnum.ocr: Visual vs Text Search in NekoImageGallery

SearchBasisEnum.vision uses CLIP embeddings to match images by visual similarity, while SearchBasisEnum.ocr uses BERT embeddings to match images by their OCR-extracted text content.

NekoImageGallery is an AI-powered image search engine that supports dual-modal retrieval capabilities. The SearchBasisEnum defines how search vectors are generated, determining whether queries match against visual features or textual content extracted via OCR. Understanding the distinction between these modes is essential for optimizing search relevance in the hv0905/nekoimagegallery repository.

Core Differences Between Vision and OCR Search Modes

The two search modes differ in three fundamental aspects: the embedding model used, the source of the vector data, and the database field queried.

Vector Sources and Embedding Models

  • SearchBasisEnum.vision: Generates image-level CLIP embeddings using OpenAI's CLIP model (clip-vit-large-patch14). For text queries, it uses the CLIP text encoder to create vectors that match visual concepts.

  • SearchBasisEnum.ocr: Generates text-level BERT embeddings using a Chinese BERT model (bert-base-chinese). These vectors represent the textual content extracted from images via OCR, not the visual appearance.

Database Index Fields

According to app/Services/vector_db_context.py, each mode maps to a distinct Qdrant vector field:

  • vision queries the IMG_VECTOR field, which stores pre-computed CLIP embeddings for every image
  • ocr queries the TEXT_VECTOR field, which stores BERT embeddings of OCR-extracted text (populated only when OCR is enabled in configuration)

Implementation in the Source Code

Enum Definition

The search basis is declared as a string enum in app/Models/api_models/search_api_model.py:

class SearchBasisEnum(str, Enum):
    vision = "vision"
    ocr = "ocr"

Vector Field Mapping

The VectorDbContext class translates the enum into the correct database field. In app/Services/vector_db_context.py, the vector_name_for_basis method handles this mapping:

@classmethod
def vector_name_for_basis(cls, basis: SearchBasisEnum) -> str:
    match basis:
        case SearchBasisEnum.vision:
            return cls.IMG_VECTOR
        case SearchBasisEnum.ocr:
            return cls.TEXT_VECTOR

Controller Logic and Endpoint Handling

The search controller in app/Controllers/search.py selects the appropriate embedding method based on the requested basis. For text searches at /search/text, it branches between CLIP and BERT:

text_vector = services.transformers_service.get_text_vector(prompt) \
             if basis.basis == SearchBasisEnum.vision \
             else services.transformers_service.get_bert_vector(prompt)

For image searches at /search/image, the system always uses vision mode because the query is already an image:

image_vector = services.transformers_service.get_image_vector(img)

In hybrid search scenarios, the controller automatically swaps the secondary basis to combine both modalities. When the primary basis is ocr, the secondary uses vision (and vice versa):

match basis.basis:
    case SearchBasisEnum.ocr:
        second_basis = SearchBasisEnum.vision
        second_vector = services.transformers_service.get_text_vector(model.extra_prompt)
    case SearchBasisEnum.vision:
        second_basis = SearchBasisEnum.ocr
        second_vector = services.transformers_service.get_bert_vector(model.extra_prompt)

Practical Usage Examples

HTTP API Calls

Query the text search endpoint with different basis parameters to switch between visual concept matching and text content matching:


# Vision-based search: finds images visually similar to "a cute cat"

curl -X POST "http://localhost:8000/api/v1/search/text?basis=vision" \
     -H "Content-Type: application/json" \
     -d '{"prompt":"a cute cat"}'

# OCR-based search: finds images containing Chinese text matching "黑猫"

curl -X POST "http://localhost:8000/api/v1/search/text?basis=ocr" \
     -H "Content-Type: application/json" \
     -d '{"prompt":"黑猫"}'

Python Client Implementation

Use the requests library to programmatically select the search mode:

import requests

API = "http://localhost:8000/api/v1/search/text"

def search(prompt: str, basis: str = "vision"):
    resp = requests.post(
        f"{API}?basis={basis}",
        json={"prompt": prompt},
    )
    resp.raise_for_status()
    return resp.json()

# Visual similarity search

print(search("sunset over the ocean", basis="vision"))

# OCR text content search

print(search("海岸线", basis="ocr"))

Internal Service Integration

For direct service access (as used in the test suite), generate vectors explicitly using TransformersService:

from app.Services.transformers_service import TransformersService

svc = TransformersService()

# CLIP embedding for visual search

vision_vec = svc.get_text_vector("a red sports car")

# BERT embedding for OCR text search  

ocr_vec = svc.get_bert_vector("红色跑车")

Summary

  • SearchBasisEnum.vision leverages OpenAI CLIP (clip-vit-large-patch14) to match queries against the IMG_VECTOR field for visual similarity
  • SearchBasisEnum.ocr leverages Chinese BERT (bert-base-chinese) to match queries against the TEXT_VECTOR field for text content similarity
  • The enum is defined in app/Models/api_models/search_api_model.py and mapped to database fields in app/Services/vector_db_context.py
  • Controllers in app/Controllers/search.py route requests to get_text_vector (CLIP) or get_bert_vector (BERT) based on the selected basis
  • Image uploads always use vision mode, while text searches can toggle between modes using the basis query parameter

Frequently Asked Questions

Can I use OCR search if OCR is disabled in the configuration?

No. The TEXT_VECTOR field is only populated when OCR is enabled via the ocr_search.enable flag in app/config.py. If disabled, the OCR search basis will return empty results because no text vectors exist in the database.

Which search mode should I use for non-Chinese text?

While SearchBasisEnum.ocr uses bert-base-chinese optimized for Chinese characters, it can still process English text. However, for English or multilingual visual concepts, SearchBasisEnum.vision typically provides better results because CLIP is trained on diverse multilingual image-text pairs and matches semantic visual concepts rather than literal text strings.

How does hybrid search handle both SearchBasisEnum values?

Hybrid searches automatically combine both modalities by using your selected basis as the primary search and automatically switching to the opposite basis for the secondary search. If you select ocr as the primary basis, the system uses vision for the secondary vector comparison, ensuring both visual and textual relevance are considered.

What embedding function is used when searching by image upload?

Image uploads via /search/image always use SearchBasisEnum.vision logic. The controller calls get_image_vector to generate CLIP embeddings directly from the uploaded image pixels, bypassing the text embedding functions entirely.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →