# Which LLM Models Are Compatible with LLM Wiki’s Vector Search?

> Discover which LLM models work with LLM Wiki's vector search. Explore compatibility with OpenAI, Azure, Google Vertex AI, Ollama, and more.

- Repository: [nash_su/llm_wiki](https://github.com/nashsu/llm_wiki)
- Tags: compatibility
- Published: 2026-09-13

---

**Any embedding model that implements the OpenAI-compatible `/v1/embeddings` endpoint is compatible with LLM Wiki, including models from OpenAI, Azure, Google Vertex AI, Doubao, Volcengine, and self-hosted Ollama instances.**

LLM Wiki’s vector search system converts text into dense float vectors using external embedding providers and indexes them in LanceDB for similarity search. According to the nashsu/llm_wiki source code, the application supports any model capable of returning a JSON object with a valid `embedding` array of `f32` values, provided the provider is explicitly configured in the embedding settings.

## Supported Embedding Providers and Models

The compatibility detection logic resides in [`src-tauri/src/commands/search.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs), where provider-specific functions validate configuration types. The system supports five primary provider categories:

### OpenAI and Azure OpenAI

The `is_openai_embedding_config` function detects providers where `provider == "openai"` or `provider == "azure"`. Compatible models include:

- `text-embedding-ada-002`
- `text-embedding-3-large`
- `text-embedding-3-small`
- Any custom model deployed on Azure OpenAI

These models use the standard `/v1/embeddings` endpoint and return vectors with dimensions ranging from 384 to 3072 depending on the specific model configuration.

### Google Vertex AI

Detected via `is_google_embedding_config` when `provider == "google"`, this integration supports:

- `textembedding-gecko@001`
- `textembedding-gecko-multilingual@001`

The implementation sends requests to the Vertex AI embedding endpoint. **Batch mode is not supported** for Google providers—the code explicitly rejects batch calls to prevent API errors.

### Doubao Multimodal

The `is_doubao_multimodal_embedding_config` function identifies Doubao configurations when `provider == "doubao"` and the model specifies a multimodal variant. Supported models include:

- `doubao-multimodal-embedding-v1`

This provider uses only the single-text embedding endpoint; batch processing is disabled to accommodate the API’s limitations.

### Volcengine

When `provider == "volcengine"`, the system selects the `volcengine_embedding_endpoint` code path. This supports any Volcengine-hosted embedding model, such as:

- `bge-large-zh`

Unlike Google and Doubao implementations, the Volcengine provider **supports batch requests**, allowing multiple texts to be embedded in a single API call for improved throughput.

### Ollama (Self-Hosted)

Self-hosted models work implicitly when you configure the `endpoint` to point to an Ollama server implementing the OpenAI-compatible `/v1/embeddings` API. Compatible models include:

- `llama2-embedding`
- `mistral-embed`

No special code path exists for Ollama—the system treats these as generic OpenAI providers, making any locally-hosted embedding model accessible without modification to the core codebase.

## Technical Implementation and Validation

The embedding pipeline relies on three core components in the nashsu/llm_wiki repository:

1. **[`src-tauri/src/commands/search.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs)** (lines 1060–1190): Contains provider detection functions and validates that returned embeddings are non-empty `Vec<f32>` arrays containing finite values. If a provider returns an unexpected shape or dimensionality, the system raises a clear error: `"embedding dim … does not match …"`.

2. **[`src-tauri/src/commands/vectorstore.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/vectorstore.rs)** (lines 150–250): Implements `vector_upsert` for storage and `search_by_embedding` for similarity queries against the LanceDB index.

3. **[`src-tauri/src/types/wiki.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/types/wiki.rs)**: Defines the data structures for page embeddings and metadata used throughout the vector search system.

## Configuration Guide

To configure a compatible model, update your [`settings.json`](https://github.com/nashsu/llm_wiki/blob/main/settings.json) or the UI’s Embedding Settings with the following structure:

```json
{
  "provider": "openai",
  "model": "text-embedding-ada-002",
  "endpoint": "https://api.openai.com/v1/embeddings",
  "apiKey": "YOUR_KEY",
  "outputDimensionality": 1536
}

```

The `provider` field accepts: `"openai"`, `"azure"`, `"google"`, `"doubao"`, or `"volcengine"`. The `outputDimensionality` parameter is optional and overrides the model’s default dimension when specified.

## Code Examples

### Upserting Page Embeddings

To store embeddings programmatically, use the `vector_upsert` function from [`src-tauri/src/commands/vectorstore.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/vectorstore.rs):

```rust
use llm_wiki::commands::vectorstore::vector_upsert;
use std::path::PathBuf;

#[tokio::main]
async fn main() {
    let project_path = PathBuf::from("/my/project");
    let page_id = "intro".to_string();
    let embedding = llm_wiki::commands::search::fake_embedding(42, 1536);

    match vector_upsert(project_path, page_id, embedding).await {
        Ok(_) => println!("✅ Embedding upserted"),
        Err(e) => eprintln!("❌ Failed: {e}"),
    }
}

```

### Searching by Vector Similarity

The `search_by_embedding` function in [`src-tauri/src/commands/search.rs`](https://github.com/nashsu/llm_wiki/blob/main/src-tauri/src/commands/search.rs) performs similarity searches:

```rust
use llm_wiki::commands::search::search_by_embedding;

#[tokio::main]
async fn main() {
    let project_path = "/my/project".into();
    let query_emb = llm_wiki::commands::search::fake_embedding(1, 1536);

    let results = search_by_embedding(project_path, query_emb, 10).await.unwrap();
    for hit in results {
        println!("• {} (score: {:.2})", hit.page_id, hit.score);
    }
}

```

### Command-Line Usage

For batch operations, use the built-in CLI:

```bash

# Generate embeddings for entire project

llm-wiki embed --project /my/project

# Search with natural language

llm-wiki search "What is vector search?" --top 5

```

## Summary

- **LLM Wiki’s vector search** requires embedding models that implement the OpenAI-compatible `/v1/embeddings` API format.
- **Five provider categories** are explicitly supported: OpenAI/Azure, Google Vertex AI, Doubao, Volcengine, and Ollama (via compatibility layer).
- **Validation occurs** in [`search.rs`](https://github.com/nashsu/llm_wiki/blob/main/search.rs), ensuring all embeddings are finite `f32` vectors with consistent dimensions.
- **Configuration** requires specifying the provider, model name, endpoint, and API key in the embedding settings.
- **Batch processing** is supported for Volcengine and OpenAI-compatible providers, but explicitly disabled for Google and Doubao integrations.

## Frequently Asked Questions

### Can I use local embedding models with LLM Wiki?

Yes. You can configure any self-hosted model that exposes an OpenAI-compatible `/v1/embeddings` endpoint, such as those running on Ollama. Set the `provider` to `"openai"` and point the `endpoint` to your local server (e.g., `http://localhost:11434/v1/embeddings`). The system will treat it as a standard OpenAI provider without requiring code changes.

### Why do I get a "dimension mismatch" error when switching models?

This error occurs when the new model’s output dimensions differ from the existing vectors in your LanceDB index. For example, switching from `text-embedding-ada-002` (1536 dimensions) to `textembedding-gecko@001` (768 dimensions) requires rebuilding the vector index. Clear the existing embeddings or create a new project to resolve the mismatch.

### Does LLM Wiki support batch embedding requests?

Batch support depends on the provider. **Volcengine** supports batch requests for improved throughput, while **Google Vertex AI** and **Doubao** explicitly disable batch mode in the code to prevent API compatibility issues. OpenAI-compatible providers generally support batching if the underlying API implementation allows it.

### How do I migrate from OpenAI to Google Vertex AI embeddings?

Update your [`settings.json`](https://github.com/nashsu/llm_wiki/blob/main/settings.json) to change `provider` to `"google"` and set the `model` to a Vertex AI embedding model such as `textembedding-gecko@001`. Ensure your `endpoint` points to the Vertex AI prediction URL and that you have configured Google Cloud authentication. Note that batch operations will be automatically disabled for this provider.