How to Use Custom Embedder Engines in AnythingLLM: A Developer’s Guide
To use a custom embedder engine in AnythingLLM, create a new class in server/utils/EmbeddingEngines/ that implements embedTextInput() and embedChunks(), register it in the getEmbeddingEngineSelection() switch statement, and set the EMBEDDING_ENGINE environment variable to your engine’s key.
AnythingLLM abstracts vector generation through a plugin architecture that lets you swap embedding providers without modifying database logic. All embedders live in server/utils/EmbeddingEngines/ and expose a common interface consumed by vector store drivers such as LanceDB or Chroma. By implementing four core methods and updating the central dispatcher, you can integrate proprietary APIs or self-hosted models into the document ingestion pipeline.
Understanding the Embedding Engine Architecture
The selection logic resides in server/utils/helpers/index.js inside the getEmbeddingEngineSelection() function. This helper reads the EMBEDDING_ENGINE environment variable and instantiates the matching class:
// server/utils/helpers/index.js (excerpt)
function getEmbeddingEngineSelection() {
const { NativeEmbedder } = require("../EmbeddingEngines/native");
const engineSelection = process.env.EMBEDDING_ENGINE;
switch (engineSelection) {
case "openai": return new OpenAiEmbedder();
case "azure": return new AzureOpenAiEmbedder();
case "localai": return new LocalAiEmbedder();
case "ollama": return new OllamaEmbedder();
case "native": return new NativeEmbedder();
default: return new NativeEmbedder();
}
}
Every engine must satisfy a minimal contract:
constructor()– Initializes the client, reads engine-specific environment variables, and setsmaxConcurrentChunksandembeddingMaxChunkLength.embedTextInput(text)– Normalizes a single string or array and delegates toembedChunks.embedChunks(chunks)– Performs the actual API calls or local inference and returns an array of embedding vectors.log()– (Optional) Provides color-coded console output for debugging.
Creating a Custom Embedder Engine
Step 1 – Create the Engine Class
Create a new directory and file at server/utils/EmbeddingEngines/myEngine/index.js. Copy the skeleton from an existing engine (such as openAi or native) and implement your custom logic:
// server/utils/EmbeddingEngines/myEngine/index.js
const { toChunks } = require("../../helpers");
class MyEngineEmbedder {
constructor() {
this.className = "MyEngineEmbedder";
this.apiKey = process.env.MYENGINE_API_KEY;
this.baseUrl = process.env.MYENGINE_BASE_URL;
this.maxConcurrentChunks = 500;
this.embeddingMaxChunkLength = 8192;
}
log(msg, ...args) {
console.log(`\x1b[36m[${this.className}]\x1b[0m ${msg}`, ...args);
}
async embedTextInput(text) {
const result = await this.embedChunks(
Array.isArray(text) ? text : [text]
);
return result?.[0] || [];
}
async embedChunks(textChunks = []) {
this.log(`Embedding ${textChunks.length} chunks...`);
const embeddingRequests = [];
for (const chunk of toChunks(textChunks, this.maxConcurrentChunks)) {
embeddingRequests.push(
new Promise(async (resolve) => {
try {
const resp = await fetch(`${this.baseUrl}/embeddings`, {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${this.apiKey}`,
},
body: JSON.stringify({ input: chunk, model: "my-model" }),
});
const { data } = await resp.json();
resolve({ data, error: null });
} catch (e) {
resolve({ data: [], error: e });
}
})
);
}
const { data = [], error = null } = await Promise.all(embeddingRequests).then(r => ({
data: r.map(res => res.data).flat(),
error: r.find(res => res.error)?.error,
}));
if (error) throw new Error(`MyEngine failed: ${error.message}`);
return data.map(d => d.embedding);
}
}
module.exports = { MyEngineEmbedder };
Step 2 – Register the Engine in the Selector
Edit server/utils/helpers/index.js (around line 260) and append a new case to the switch statement:
// server/utils/helpers/index.js
case "myengine":
const { MyEngineEmbedder } = require("../EmbeddingEngines/myEngine");
return new MyEngineEmbedder();
Step 3 – Configure Environment Variables
Add your engine’s configuration to .env or .env.example:
EMBEDDING_ENGINE=myengine
MYENGINE_API_KEY=sk-xxxxxxxxxxxx
MYENGINE_BASE_URL=https://api.myengine.com/v1
Restart the server after saving the environment file. AnythingLLM reads EMBEDDING_ENGINE at startup and routes all vectorization through your new class.
How the Engine Integrates with Vector Storage
When documents are ingested, vector store drivers (LanceDB, Chroma, Pinecone, etc.) retrieve the active embedder via getEmbeddingEngineSelection() and call embedChunks():
// server/utils/vectorDbProviders/lance/index.js (excerpt, lines 340-351)
const EmbedderEngine = getEmbeddingEngineSelection();
// ...
const vectorValues = await EmbedderEngine.embedChunks(textChunks);
Because every provider respects the same interface, you do not need to modify database-specific code. As long as your class returns an array of float arrays from embedChunks(), the vector DB layer will persist the embeddings correctly.
Summary
- Plugin location: Place custom engines in
server/utils/EmbeddingEngines/<name>/index.js. - Required methods: Implement
constructor(),embedTextInput(), andembedChunks()to satisfy the interface. - Registration: Add a case to
getEmbeddingEngineSelection()inserver/utils/helpers/index.js. - Configuration: Set
EMBEDDING_ENGINEand any provider-specific variables in.env. - Abstraction: Vector DB drivers in
server/utils/vectorDbProviders/*/index.jsconsume the engine generically, ensuring portability across storage backends.
Frequently Asked Questions
What is the maximum chunk length I can configure for a custom embedder?
Set this.embeddingMaxChunkLength in your class constructor to match your provider’s token limit. For example, OpenAI’s text-embedding-3-small supports 8,192 tokens, while other APIs may differ. This value prevents overflow errors during batch requests.
Can I use a local model instead of an API for custom embeddings?
Yes. Reference the native embedder in server/utils/EmbeddingEngines/native/index.js, which loads transformer models via Xenova Transformers. Replace the fetch logic in embedChunks() with local inference calls, ensuring you still return an array of embedding vectors.
Why does AnythingLLM use a switch statement instead of dynamic imports for engine selection?
The switch statement in getEmbeddingEngineSelection() provides explicit control over initialization order and allows conditional loading of dependencies. This prevents unused embedder libraries from being required at runtime, reducing memory footprint and startup time.
How do I debug embedding failures in a custom engine?
Implement the log() method with color-coded output (using ANSI codes like \x1b[36m) to trace batch sizes and API responses. AnythingLLM will surface errors thrown by embedChunks() in the server logs, including stack traces that point to your custom file in server/utils/EmbeddingEngines/.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →