How to Use Custom Embedder Engines in AnythingLLM: A Developer’s Guide

To use a custom embedder engine in AnythingLLM, create a new class in server/utils/EmbeddingEngines/ that implements embedTextInput() and embedChunks(), register it in the getEmbeddingEngineSelection() switch statement, and set the EMBEDDING_ENGINE environment variable to your engine’s key.

AnythingLLM abstracts vector generation through a plugin architecture that lets you swap embedding providers without modifying database logic. All embedders live in server/utils/EmbeddingEngines/ and expose a common interface consumed by vector store drivers such as LanceDB or Chroma. By implementing four core methods and updating the central dispatcher, you can integrate proprietary APIs or self-hosted models into the document ingestion pipeline.

Understanding the Embedding Engine Architecture

The selection logic resides in server/utils/helpers/index.js inside the getEmbeddingEngineSelection() function. This helper reads the EMBEDDING_ENGINE environment variable and instantiates the matching class:

// server/utils/helpers/index.js (excerpt)
function getEmbeddingEngineSelection() {
  const { NativeEmbedder } = require("../EmbeddingEngines/native");
  const engineSelection = process.env.EMBEDDING_ENGINE;
  switch (engineSelection) {
    case "openai":          return new OpenAiEmbedder();
    case "azure":           return new AzureOpenAiEmbedder();
    case "localai":         return new LocalAiEmbedder();
    case "ollama":          return new OllamaEmbedder();
    case "native":          return new NativeEmbedder();
    default:                return new NativeEmbedder();
  }
}

Every engine must satisfy a minimal contract:

  • constructor() – Initializes the client, reads engine-specific environment variables, and sets maxConcurrentChunks and embeddingMaxChunkLength.
  • embedTextInput(text) – Normalizes a single string or array and delegates to embedChunks.
  • embedChunks(chunks) – Performs the actual API calls or local inference and returns an array of embedding vectors.
  • log() – (Optional) Provides color-coded console output for debugging.

Creating a Custom Embedder Engine

Step 1 – Create the Engine Class

Create a new directory and file at server/utils/EmbeddingEngines/myEngine/index.js. Copy the skeleton from an existing engine (such as openAi or native) and implement your custom logic:

// server/utils/EmbeddingEngines/myEngine/index.js
const { toChunks } = require("../../helpers");

class MyEngineEmbedder {
  constructor() {
    this.className = "MyEngineEmbedder";
    this.apiKey = process.env.MYENGINE_API_KEY;
    this.baseUrl = process.env.MYENGINE_BASE_URL;
    this.maxConcurrentChunks = 500;
    this.embeddingMaxChunkLength = 8192;
  }

  log(msg, ...args) {
    console.log(`\x1b[36m[${this.className}]\x1b[0m ${msg}`, ...args);
  }

  async embedTextInput(text) {
    const result = await this.embedChunks(
      Array.isArray(text) ? text : [text]
    );
    return result?.[0] || [];
  }

  async embedChunks(textChunks = []) {
    this.log(`Embedding ${textChunks.length} chunks...`);
    const embeddingRequests = [];

    for (const chunk of toChunks(textChunks, this.maxConcurrentChunks)) {
      embeddingRequests.push(
        new Promise(async (resolve) => {
          try {
            const resp = await fetch(`${this.baseUrl}/embeddings`, {
              method: "POST",
              headers: {
                "Content-Type": "application/json",
                "Authorization": `Bearer ${this.apiKey}`,
              },
              body: JSON.stringify({ input: chunk, model: "my-model" }),
            });
            const { data } = await resp.json();
            resolve({ data, error: null });
          } catch (e) {
            resolve({ data: [], error: e });
          }
        })
      );
    }

    const { data = [], error = null } = await Promise.all(embeddingRequests).then(r => ({
      data: r.map(res => res.data).flat(),
      error: r.find(res => res.error)?.error,
    }));

    if (error) throw new Error(`MyEngine failed: ${error.message}`);
    return data.map(d => d.embedding);
  }
}

module.exports = { MyEngineEmbedder };

Step 2 – Register the Engine in the Selector

Edit server/utils/helpers/index.js (around line 260) and append a new case to the switch statement:

// server/utils/helpers/index.js
case "myengine":
  const { MyEngineEmbedder } = require("../EmbeddingEngines/myEngine");
  return new MyEngineEmbedder();

Step 3 – Configure Environment Variables

Add your engine’s configuration to .env or .env.example:

EMBEDDING_ENGINE=myengine
MYENGINE_API_KEY=sk-xxxxxxxxxxxx
MYENGINE_BASE_URL=https://api.myengine.com/v1

Restart the server after saving the environment file. AnythingLLM reads EMBEDDING_ENGINE at startup and routes all vectorization through your new class.

How the Engine Integrates with Vector Storage

When documents are ingested, vector store drivers (LanceDB, Chroma, Pinecone, etc.) retrieve the active embedder via getEmbeddingEngineSelection() and call embedChunks():

// server/utils/vectorDbProviders/lance/index.js (excerpt, lines 340-351)
const EmbedderEngine = getEmbeddingEngineSelection();
// ...
const vectorValues = await EmbedderEngine.embedChunks(textChunks);

Because every provider respects the same interface, you do not need to modify database-specific code. As long as your class returns an array of float arrays from embedChunks(), the vector DB layer will persist the embeddings correctly.

Summary

  • Plugin location: Place custom engines in server/utils/EmbeddingEngines/<name>/index.js.
  • Required methods: Implement constructor(), embedTextInput(), and embedChunks() to satisfy the interface.
  • Registration: Add a case to getEmbeddingEngineSelection() in server/utils/helpers/index.js.
  • Configuration: Set EMBEDDING_ENGINE and any provider-specific variables in .env.
  • Abstraction: Vector DB drivers in server/utils/vectorDbProviders/*/index.js consume the engine generically, ensuring portability across storage backends.

Frequently Asked Questions

What is the maximum chunk length I can configure for a custom embedder?

Set this.embeddingMaxChunkLength in your class constructor to match your provider’s token limit. For example, OpenAI’s text-embedding-3-small supports 8,192 tokens, while other APIs may differ. This value prevents overflow errors during batch requests.

Can I use a local model instead of an API for custom embeddings?

Yes. Reference the native embedder in server/utils/EmbeddingEngines/native/index.js, which loads transformer models via Xenova Transformers. Replace the fetch logic in embedChunks() with local inference calls, ensuring you still return an array of embedding vectors.

Why does AnythingLLM use a switch statement instead of dynamic imports for engine selection?

The switch statement in getEmbeddingEngineSelection() provides explicit control over initialization order and allows conditional loading of dependencies. This prevents unused embedder libraries from being required at runtime, reducing memory footprint and startup time.

How do I debug embedding failures in a custom engine?

Implement the log() method with color-coded output (using ANSI codes like \x1b[36m) to trace batch sizes and API responses. AnythingLLM will surface errors thrown by embedChunks() in the server logs, including stack traces that point to your custom file in server/utils/EmbeddingEngines/.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →