Which LLM Models Are Compatible with LLM Wiki’s Vector Search?

Any embedding model that implements the OpenAI-compatible /v1/embeddings endpoint is compatible with LLM Wiki, including models from OpenAI, Azure, Google Vertex AI, Doubao, Volcengine, and self-hosted Ollama instances.

LLM Wiki’s vector search system converts text into dense float vectors using external embedding providers and indexes them in LanceDB for similarity search. According to the nashsu/llm_wiki source code, the application supports any model capable of returning a JSON object with a valid embedding array of f32 values, provided the provider is explicitly configured in the embedding settings.

Supported Embedding Providers and Models

The compatibility detection logic resides in src-tauri/src/commands/search.rs, where provider-specific functions validate configuration types. The system supports five primary provider categories:

OpenAI and Azure OpenAI

The is_openai_embedding_config function detects providers where provider == "openai" or provider == "azure". Compatible models include:

  • text-embedding-ada-002
  • text-embedding-3-large
  • text-embedding-3-small
  • Any custom model deployed on Azure OpenAI

These models use the standard /v1/embeddings endpoint and return vectors with dimensions ranging from 384 to 3072 depending on the specific model configuration.

Google Vertex AI

Detected via is_google_embedding_config when provider == "google", this integration supports:

  • textembedding-gecko@001
  • textembedding-gecko-multilingual@001

The implementation sends requests to the Vertex AI embedding endpoint. Batch mode is not supported for Google providers—the code explicitly rejects batch calls to prevent API errors.

Doubao Multimodal

The is_doubao_multimodal_embedding_config function identifies Doubao configurations when provider == "doubao" and the model specifies a multimodal variant. Supported models include:

  • doubao-multimodal-embedding-v1

This provider uses only the single-text embedding endpoint; batch processing is disabled to accommodate the API’s limitations.

Volcengine

When provider == "volcengine", the system selects the volcengine_embedding_endpoint code path. This supports any Volcengine-hosted embedding model, such as:

  • bge-large-zh

Unlike Google and Doubao implementations, the Volcengine provider supports batch requests, allowing multiple texts to be embedded in a single API call for improved throughput.

Ollama (Self-Hosted)

Self-hosted models work implicitly when you configure the endpoint to point to an Ollama server implementing the OpenAI-compatible /v1/embeddings API. Compatible models include:

  • llama2-embedding
  • mistral-embed

No special code path exists for Ollama—the system treats these as generic OpenAI providers, making any locally-hosted embedding model accessible without modification to the core codebase.

Technical Implementation and Validation

The embedding pipeline relies on three core components in the nashsu/llm_wiki repository:

  1. src-tauri/src/commands/search.rs (lines 1060–1190): Contains provider detection functions and validates that returned embeddings are non-empty Vec<f32> arrays containing finite values. If a provider returns an unexpected shape or dimensionality, the system raises a clear error: "embedding dim … does not match …".

  2. src-tauri/src/commands/vectorstore.rs (lines 150–250): Implements vector_upsert for storage and search_by_embedding for similarity queries against the LanceDB index.

  3. src-tauri/src/types/wiki.rs: Defines the data structures for page embeddings and metadata used throughout the vector search system.

Configuration Guide

To configure a compatible model, update your settings.json or the UI’s Embedding Settings with the following structure:

{
  "provider": "openai",
  "model": "text-embedding-ada-002",
  "endpoint": "https://api.openai.com/v1/embeddings",
  "apiKey": "YOUR_KEY",
  "outputDimensionality": 1536
}

The provider field accepts: "openai", "azure", "google", "doubao", or "volcengine". The outputDimensionality parameter is optional and overrides the model’s default dimension when specified.

Code Examples

Upserting Page Embeddings

To store embeddings programmatically, use the vector_upsert function from src-tauri/src/commands/vectorstore.rs:

use llm_wiki::commands::vectorstore::vector_upsert;
use std::path::PathBuf;

#[tokio::main]
async fn main() {
    let project_path = PathBuf::from("/my/project");
    let page_id = "intro".to_string();
    let embedding = llm_wiki::commands::search::fake_embedding(42, 1536);

    match vector_upsert(project_path, page_id, embedding).await {
        Ok(_) => println!("✅ Embedding upserted"),
        Err(e) => eprintln!("❌ Failed: {e}"),
    }
}

Searching by Vector Similarity

The search_by_embedding function in src-tauri/src/commands/search.rs performs similarity searches:

use llm_wiki::commands::search::search_by_embedding;

#[tokio::main]
async fn main() {
    let project_path = "/my/project".into();
    let query_emb = llm_wiki::commands::search::fake_embedding(1, 1536);

    let results = search_by_embedding(project_path, query_emb, 10).await.unwrap();
    for hit in results {
        println!("• {} (score: {:.2})", hit.page_id, hit.score);
    }
}

Command-Line Usage

For batch operations, use the built-in CLI:


# Generate embeddings for entire project

llm-wiki embed --project /my/project

# Search with natural language

llm-wiki search "What is vector search?" --top 5

Summary

  • LLM Wiki’s vector search requires embedding models that implement the OpenAI-compatible /v1/embeddings API format.
  • Five provider categories are explicitly supported: OpenAI/Azure, Google Vertex AI, Doubao, Volcengine, and Ollama (via compatibility layer).
  • Validation occurs in search.rs, ensuring all embeddings are finite f32 vectors with consistent dimensions.
  • Configuration requires specifying the provider, model name, endpoint, and API key in the embedding settings.
  • Batch processing is supported for Volcengine and OpenAI-compatible providers, but explicitly disabled for Google and Doubao integrations.

Frequently Asked Questions

Can I use local embedding models with LLM Wiki?

Yes. You can configure any self-hosted model that exposes an OpenAI-compatible /v1/embeddings endpoint, such as those running on Ollama. Set the provider to "openai" and point the endpoint to your local server (e.g., http://localhost:11434/v1/embeddings). The system will treat it as a standard OpenAI provider without requiring code changes.

Why do I get a "dimension mismatch" error when switching models?

This error occurs when the new model’s output dimensions differ from the existing vectors in your LanceDB index. For example, switching from text-embedding-ada-002 (1536 dimensions) to textembedding-gecko@001 (768 dimensions) requires rebuilding the vector index. Clear the existing embeddings or create a new project to resolve the mismatch.

Does LLM Wiki support batch embedding requests?

Batch support depends on the provider. Volcengine supports batch requests for improved throughput, while Google Vertex AI and Doubao explicitly disable batch mode in the code to prevent API compatibility issues. OpenAI-compatible providers generally support batching if the underlying API implementation allows it.

How do I migrate from OpenAI to Google Vertex AI embeddings?

Update your settings.json to change provider to "google" and set the model to a Vertex AI embedding model such as textembedding-gecko@001. Ensure your endpoint points to the Vertex AI prediction URL and that you have configured Google Cloud authentication. Note that batch operations will be automatically disabled for this provider.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →