NekoImageGallery Model Selection and Device Allocation: A Complete Configuration Guide

NekoImageGallery exposes four environment variables—APP_MODEL__CLIP, APP_MODEL__BERT, APP_MODEL__EASYPADDLEOCR, and APP_DEVICE—that control which deep-learning models load for image encoding, text embedding, and OCR, as well as whether inference runs on CPU, GPU, or auto-selected hardware.

NekoImageGallery model selection and device allocation are governed by the Config class in app/config.py (lines 15‑99), which aggregates settings from environment variables prefixed with APP_ or from .env files in the config/ directory. These configuration options determine the specific transformer models used for CLIP-based image similarity, BERT-based text search, and optional EasyPaddleOCR recognition, plus the PyTorch execution device.

Model Architecture Selection

The ModelsSettings pydantic model in app/config.py (lines 30‑34) defines three tunable model paths. All values accept HuggingFace model identifiers or local filesystem paths.

  • APP_MODEL__CLIP – Controls the CLIP vision encoder used for image-to-text similarity search. According to the hv0905/nekoimagegallery source code, the default value is openai/clip-vit-large-patch14, referenced in ModelsSettings.clip.

  • APP_MODEL__BERT – Specifies the BERT model for OCR-derived text embeddings and Chinese-language text processing. The default is bert-base-chinese, defined in ModelsSettings.bert.

  • APP_MODEL__EASYPADDLEOCR – Optional custom path or HuggingFace identifier for the EasyPaddleOCR engine. When set to None (the default), the system auto-downloads the standard model. This is implemented in ModelsSettings.easypaddleocr.

Hardware Device Allocation

The APP_DEVICE setting controls PyTorch hardware placement for all inference workloads. Valid values include:

  • auto – Automatically selects CUDA if available, otherwise falls back to CPU.
  • cpu – Forces CPU-only execution regardless of GPU availability.
  • cuda or cuda:0, cuda:1, etc. – Targets specific NVIDIA devices.

This logic resides in Config.device within app/config.py and defaults to auto.

Configuration File Locations and Environment Overrides

NekoImageGallery reads configuration from three sources, with later sources overriding earlier ones: built-in defaults, config/default.env, and config/local.env (or exported shell variables).

1. Per-project overrides via config/local.env:

Create config/local.env in your repository root to persist settings across restarts:


# Use a smaller CLIP variant and force GPU allocation

APP_MODEL__CLIP=laion/CLIP-ViT-B-32
APP_DEVICE=cuda

2. Runtime environment variables:

Pass configuration directly via shell export or Docker -e flags for ephemeral changes:

export APP_MODEL__CLIP=google/vit-base-patch16-224
export APP_MODEL__BERT=distilbert-base-uncased
export APP_DEVICE=cpu   # Force CPU even when GPU is present

uvicorn main:app --host 0.0.0.0 --port 8000

3. Enabling custom OCR models:

To activate EasyPaddleOCR with a custom model directory, set both the model path and the enablement flag:

APP_MODEL__EASYPADDLEOCR=/path/to/custom/ocr/model
APP_OCR_SEARCH__ENABLE=True

Programmatic Configuration Access

Python code can inspect the resolved configuration at runtime by importing the singleton config instance. This exposes typed attributes for all model identifiers and device strings:

from app.config import config

clip_model_name = config.model.clip          # e.g., "openai/clip-vit-large-patch14"

bert_model_name = config.model.bert          # e.g., "bert-base-chinese"

device = config.device                     # "auto", "cpu", or "cuda"

print(f"Using CLIP={clip_model_name}, BERT={bert_model_name} on {device}")

Summary

  • Four core environment variables control NekoImageGallery inference: APP_MODEL__CLIP, APP_MODEL__BERT, APP_MODEL__EASYPADDLEOCR, and APP_DEVICE.
  • Defaults are hardcoded in app/config.py: openai/clip-vit-large-patch14 for vision, bert-base-chinese for text, and auto for device selection.
  • Configuration hierarchy follows Pydantic-settings logic: defaults → config/default.env → config/local.env → exported shell variables.
  • All changes take effect on application startup; the Config class initializes once and provides immutable settings to the CLIP, BERT, and OCR pipeline constructors.

Frequently Asked Questions

What is the default CLIP model used by NekoImageGallery?

The default CLIP vision encoder is openai/clip-vit-large-patch14, as specified in the ModelsSettings.clip attribute inside app/config.py (lines 30‑34). You can override this with any HuggingFace-compatible CLIP identifier via the APP_MODEL__CLIP environment variable.

Can I force CPU-only inference even when a GPU is available?

Yes. Set APP_DEVICE=cpu in your environment or config/local.env. This bypasses the automatic GPU detection logic in Config.device (lines 87‑89) and forces PyTorch to load all tensors and models onto the CPU.

How do I specify a custom EasyPaddleOCR model instead of the auto-downloaded one?

Export APP_MODEL__EASYPADDLEOCR=/absolute/path/to/model and ensure APP_OCR_SEARCH__ENABLE=True. If APP_MODEL__EASYPADDLEOCR remains unset (the default None), NekoImageGallery automatically downloads the standard pretrained weights on first use.

Where should I place environment variables for production deployments?

Place persistent settings in config/local.env (git‑ignored by default) or inject them via your orchestration layer (Docker Compose environment: blocks, Kubernetes ConfigMaps, or systemd service files). The application parses these through Pydantic’s SettingsConfig in app/config.py, which prefixes all variables with APP_ and supports Unix-style shell expansion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →