How to Set Up and Use Custom LLM Providers in AnythingLLM
You can integrate any language model backend into AnythingLLM either by configuring the built-in generic-openai provider for OpenAI-compatible endpoints or by implementing a new provider class in server/utils/AiProviders/ and registering it in the provider factory.
AnythingLLM discovers and routes chat requests through a provider abstraction that lives in the Mintplex-Labs/anything-llm repository. By leveraging the factory pattern in server/utils/helpers/index.js, you can connect self-hosted models, third-party APIs, or custom endpoints while preserving the standard workspace UI and chat workflow.
Understanding the Provider Architecture
The system selects LLM implementations via getLLMProvider() in server/utils/helpers/index.js (lines 131-150). This factory reads the LLM_PROVIDER environment variable—or an override from workspace settings—and instantiates the matching class:
// server/utils/helpers/index.js
function getLLMProvider({ provider = null, model = null } = {}) {
const LLMSelection = provider ?? process.env.LLM_PROVIDER ?? "openai";
const embedder = getEmbeddingEngineSelection();
switch (LLMSelection) {
case "openai": /* … */ break;
case "azure": /* … */ break;
case "generic-openai": /* … */ break;
case "anthropic": /* … */ break;
// … additional built-in providers
default:
throw new Error(
`ENV: No valid LLM_PROVIDER value found in environment! Using ${process.env.LLM_PROVIDER}`
);
}
}
Each built-in provider ships as a dedicated class under server/utils/AiProviders/<name>/index.js. When you add a custom provider, you extend this switch statement to map a new string key to your implementation.
Using the Built-In Generic OpenAI Provider
For any service exposing an OpenAI-compatible HTTP API—such as local LLM servers, LiteLLM proxies, or alternative hosting providers—you can use the generic-openai provider without writing code.
This provider class resides in server/utils/AiProviders/genericOpenAi/index.js and constructs a client that targets your custom base URL:
// server/utils/AiProviders/genericOpenAi/index.js
class GenericOpenAiLLM {
constructor(embedder = null, modelPreference = null) {
if (!process.env.GENERIC_OPEN_AI_BASE_PATH)
throw new Error("GenericOpenAI must have a valid base path to use for the api.");
this.basePath = process.env.GENERIC_OPEN_AI_BASE_PATH;
this.openai = new OpenAIApi({
baseURL: this.basePath,
apiKey: process.env.GENERIC_OPEN_AI_API_KEY ?? null,
defaultHeaders: {
"User-Agent": getAnythingLLMUserAgent(),
...GenericOpenAiLLM.parseCustomHeaders(),
},
});
this.model = modelPreference ?? process.env.GENERIC_OPEN_AI_MODEL_PREF ?? null;
}
}
Configuration Environment Variables
Add these entries to your .env or .env.development file:
LLM_PROVIDER=generic-openai– Activates the generic provider.GENERIC_OPEN_AI_BASE_PATH– Base URL of your custom endpoint (e.g.,http://localhost:8000/v1).GENERIC_OPEN_AI_API_KEY– API key if your endpoint requires authentication.GENERIC_OPEN_AI_MODEL_PREF– Default model name when none is selected in the UI.GENERIC_OPEN_AI_MODEL_TOKEN_LIMIT– Context window size (defaults to 4096).GENERIC_OPEN_AI_MAX_TOKENS– Maximum completion tokens per request.GENERIC_OPEN_AI_CUSTOM_HEADERS– Additional HTTP headers in CSV format:"Header-Name:Value,Other:Val".
Example Configuration
# .env
LLM_PROVIDER=generic-openai
GENERIC_OPEN_AI_BASE_PATH=https://api.mycustomllm.com/v1
GENERIC_OPEN_AI_API_KEY=sk-custom-key-12345
GENERIC_OPEN_AI_MODEL_PREF=llama-3.1-70b
GENERIC_OPEN_AI_MODEL_TOKEN_LIMIT=8192
GENERIC_OPEN_AI_MAX_TOKENS=2048
GENERIC_OPEN_AI_CUSTOM_HEADERS="X-Org-ID:42,X-Request-Source:anything-llm"
After restarting the server, the UI will expose "Generic OpenAI" in the LLM preference dropdown, routing all chat traffic to your specified endpoint.
Building a Completely Custom LLM Provider
When your backend follows a non-OpenAI protocol—such as a private gRPC service, AWS Bedrock, or a specialized JSON schema—you must implement a new provider class.
Step 1: Create the Provider Class
Create a folder under server/utils/AiProviders/<your-provider>/ and export a class that implements the interface defined in server/utils/agents/aibitat/providers/ai-provider.js. Required public methods include:
streamingEnabled()– Returns boolean indicating streaming support.promptWindowLimit()– Returns the model’s context window size.isValidChatCompletionModel(modelName)– Validates model selection.constructPrompt({ systemPrompt, contextTexts, chatHistory, userPrompt })– Builds the request payload.getChatCompletion(messages, opts)– Executes the request and returns{ textResponse, metrics }.streamGetChatCompletion(messages, opts)– Optional streaming variant returning a measured stream.
Step 2: Register in the Provider Factory
Add a case to the switch statement in server/utils/helpers/index.js around line 165:
case "mycoolai":
const { MyCoolAiLLM } = require("../AiProviders/mycoolai");
return new MyCoolAiLLM(embedder, model);
Step 3: Expose Environment Variables
Edit server/utils/helpers/updateENV.js to surface configuration fields in General Settings → LLM Preference:
{
envKey: "MYCOOLAI_BASE_PATH",
label: "MyCoolAI Base URL",
description: "Full URL to the MyCoolAI chat completions endpoint",
isPassword: false,
},
{
envKey: "MYCOOLAI_API_KEY",
label: "MyCoolAI API Key",
description: "Secret key for authenticating requests",
isPassword: true,
},
{
envKey: "MYCOOLAI_MODEL_PREF",
label: "Default Model",
description: "Model name to use when none is selected in UI",
isPassword: false,
}
Step 4: Add the UI Entry
Update frontend/src/pages/GeneralSettings/LLMPreference.vue to include your provider in AVAILABLE_LLM_PROVIDERS:
{
label: "MyCoolAI",
value: "mycoolai",
requiredConfig: ["MYCOOLAI_BASE_PATH", "MYCOOLAI_API_KEY", "MYCOOLAI_MODEL_PREF"],
}
The hasMissingCredentials() function (lines 46-58 in server/utils/helpers/index.js) uses thisrequiredConfig` array to validate that all necessary credentials are present before allowing the provider to be selected.
Complete Implementation Example
Here is a skeleton for server/utils/AiProviders/mycoolai/index.js targeting a hypothetical REST API:
const { LLMPerformanceMonitor } = require("../../helpers/chat/LLMPerformanceMonitor");
const {
formatChatHistory,
writeResponseChunk,
clientAbortedHandler,
} = require("../../helpers/chat/responses");
const fetch = require("node-fetch");
class MyCoolAiLLM {
constructor(embedder = null, modelPreference = null) {
this.className = "MyCoolAiLLM";
this.basePath = process.env.MYCOOLAI_BASE_PATH;
this.apiKey = process.env.MYCOOLAI_API_KEY ?? null;
this.model = modelPreference ?? process.env.MYCOOLAI_MODEL_PREF ?? null;
if (!this.basePath || !this.model) {
throw new Error("MyCoolAI requires MYCOOLAI_BASE_PATH and MYCOOLAI_MODEL_PREF env vars");
}
this.embedder = embedder ?? new (require("../../EmbeddingEngines/native")).NativeEmbedder();
}
log(txt, ...args) {
console.log(`\x1b[35m[${this.className}]\x1b[0m ${txt}`, ...args);
}
streamingEnabled() {
return false; // Set to true if implementing streamGetChatCompletion
}
static promptWindowLimit() {
return Number(process.env.MYCOOLAI_TOKEN_LIMIT || 4096);
}
isValidChatCompletionModel(model = "") {
return model.includes("mycool"); // Adjust validation logic
}
constructPrompt({ systemPrompt = "", contextTexts = [], chatHistory = [], userPrompt = "" }) {
const messages = [
{
role: "system",
content: systemPrompt + contextTexts.map(t => `\nContext: ${t}`).join("")
},
...formatChatHistory(chatHistory),
{ role: "user", content: userPrompt },
];
return { model: this.model, messages };
}
async getChatCompletion(messages = null, { temperature = 0.7 }) {
const payload = this.constructPrompt({
userPrompt: messages?.[messages.length - 1]?.content
});
const result = await LLMPerformanceMonitor.measureAsyncFunction(
fetch(`${this.basePath}/v1/chat/completions`, {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: this.apiKey ? `Bearer ${this.apiKey}` : undefined,
},
body: JSON.stringify({ ...payload, temperature, max_tokens: 1024 }),
}).then(r => r.json())
);
return {
textResponse: result.output?.choices?.[0]?.message?.content ?? "",
metrics: {
prompt_tokens: result.usage?.prompt_tokens,
completion_tokens: result.usage?.completion_tokens,
total_tokens: result.usage?.total_tokens,
outputTps: result.usage?.completion_tokens / (result.latency || 1),
latency: result.latency,
},
};
}
}
module.exports = { MyCoolAiLLM };
How the UI Drives Provider Selection
The frontend provider selector (frontend/src/components/WorkspaceChat/ChatContainer/PromptInput/LLMSelector/utils.js) filters the AVAILABLE_LLM_PROVIDERS array based on which configurations have non-empty credentials. When a user selects a provider in the workspace settings, the value is stored in workspace.chatProvider.
At request time, the backend invokes getLLMProvider({ provider: workspace.chatProvider }), which instantiates the correct class according to the logic in the factory. This architecture keeps the UI agnostic to the underlying transport protocol while ensuring each workspace can target a different backend.
Summary
- Provider Factory:
getLLMProvider()inserver/utils/helpers/index.jsroutes requests based on theLLM_PROVIDERenvironment variable or workspace setting. - Generic OpenAI: Use
LLM_PROVIDER=generic-openaiwithGENERIC_OPEN_AI_BASE_PATHto connect any OpenAI-compatible endpoint without code changes. - Custom Providers: Implement a class under
server/utils/AiProviders/<name>/with standard methods (constructPrompt,getChatCompletion, etc.), register it in the factory switch statement, and expose env vars inupdateENV.js. - UI Integration: Add entries to
AVAILABLE_LLM_PROVIDERSin the frontend settings page and definerequiredConfigarrays for credential validation. - Credential Validation: The
hasMissingCredentials()helper checks that all required environment variables are populated before surfacing a provider in the dropdown.
Frequently Asked Questions
Can I use multiple custom providers simultaneously in different workspaces?
Yes. The getLLMProvider() factory accepts a provider parameter that overrides the global LLM_PROVIDER environment variable. Each workspace stores its selection in workspace.chatProvider, allowing Workspace A to use generic-openai pointing to a local model while Workspace B uses a custom Anthropic provider, all within the same AnythingLLM instance.
What is the difference between generic-openai and creating a new provider class?
generic-openai requires only environment variable configuration and works immediately for any HTTP API matching the OpenAI chat completions schema. Creating a new provider class is necessary when your backend uses a different authentication mechanism, request format, or response structure that cannot be mapped to OpenAI’s conventions. The custom class approach also allows you to implement provider-specific optimizations or streaming protocols.
How do I validate that my custom provider credentials are working?
AnythingLLM validates credentials through the hasMissingCredentials() function in server/utils/helpers/index.js. Ensure your provider entry in AVAILABLE_LLM_PROVIDERS lists all required environment keys in its requiredConfig array. The UI will display a warning if any are missing. For runtime validation, check the server logs for initialization errors thrown by your provider’s constructor, such as missing BASE_PATH or failed connection tests.
Is streaming supported for custom LLM providers?
Yes. Implement the streamGetChatCompletion() method in your provider class and have streamingEnabled() return true. This method should return a measured stream that yields text chunks using the writeResponseChunk helper from server/utils/helpers/chat/responses.js. Refer to existing streaming implementations like the OpenAI or Azure providers in server/utils/AiProviders/ for the exact pattern to follow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →