How to Use the AnythingLLM Developer API for Custom Integrations

The AnythingLLM Developer API provides an OpenAI-compatible REST interface that authenticates via Bearer tokens and exposes workspace-specific chat completion endpoints at /api/v1/openai/*.

AnythingLLM ships a RESTful Developer API built on Express that allows external services to interact with running instances programmatically. As implemented in the Mintplex-Labs/anything-llm repository, this interface follows the OpenAI SDK specification, enabling you to reuse existing client libraries or direct HTTP calls to integrate custom applications with AnythingLLM workspaces.

Authenticating API Requests

Every request to the Developer API must include a valid Bearer API key in the Authorization header. The validApiKey middleware in server/utils/middleware/validApiKey.js validates the token against the ApiKey model stored in the api_keys table. If the key is missing or invalid, the middleware aborts the request with a 403 Forbidden response.

To generate a key, navigate to Settings → API Keys in the AnythingLLM web UI. The secret is persisted via server/models/apiKeys.js and must be passed as Authorization: Bearer <KEY> in all subsequent calls.

Core Architecture and Endpoints

The API contract is fully documented in server/swagger/openapi.json, which describes all /v1/* routes, request schemas, and response formats. The base path for all API calls is /api, making the complete URL pattern https://your-instance.com/api/v1/openai/<endpoint>.

The two primary endpoints are implemented in server/endpoints/api/openai/index.js:

  1. GET /v1/openai/models — Returns a list of available workspace slugs that function as model identifiers.
  2. POST /v1/openai/chat/completions — Accepts OpenAI-compatible chat payloads and returns responses via JSON or Server-Sent Events (SSE).

Listing Workspaces as Models

Before sending chat requests, you must identify which workspace to target. The models endpoint iterates over all workspaces defined in server/models/workspace.js and returns them as OpenAI-style model objects.

curl -s -X GET https://your-instance.com/api/v1/openai/models \
  -H "Authorization: Bearer $API_KEY"

The response contains workspace slugs in the id field:

{
  "object": "list",
  "data": [
    {
      "id": "sales-assistant",
      "object": "model",
      "created": 1691234567,
      "owned_by": "openai-gpt-4"
    }
  ]
}

Use the id value (e.g., sales-assistant) as the model parameter in chat completion requests.

Sending Chat Completions

The chat endpoint at POST /v1/openai/chat/completions accepts standard OpenAI payloads including model, messages, temperature, and stream. The handler in server/endpoints/api/openai/index.js (lines 84-199) extracts the last user message, constructs the system prompt and conversation history, then delegates execution to the OpenAICompatibleChat engine located in server/utils/chats/openaiCompatible.js.

Non-Streaming Requests

For synchronous responses, set stream: false (or omit the field). The server runs the LLM provider configured for the specified workspace and returns a complete JSON payload.

curl -s -X POST https://your-instance.com/api/v1/openai/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sales-assistant",
    "messages": [
      {"role": "system", "content": "You are a helpful sales assistant."},
      {"role": "user", "content": "What does our pricing plan include?"}
    ],
    "temperature": 0.7,
    "stream": false
  }'

The response follows the standard OpenAI schema including choices, usage statistics, and finish_reason.

Streaming with Server-Sent Events

When stream: true, the endpoint returns an SSE feed with Content-Type: text/event-stream. The OpenAICompatibleChat.streamChat method writes incremental deltas to the response until the generation completes, followed by a [DONE] message.

Below is a Node.js example using the native fetch API (Node ≥ 18) to consume the stream:

// streamChat.js
const API_URL = 'https://your-instance.com/api/v1/openai/chat/completions';
const API_KEY = process.env.API_KEY;

async function streamChat() {
  const response = await fetch(API_URL, {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${API_KEY}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify({
      model: 'dev-ops-helper',
      messages: [
        { role: 'system', content: 'You are a DevOps assistant.' },
        { role: 'user', content: 'How do I rotate a Kubernetes secret?' }
      ],
      temperature: 0.6,
      stream: true,
    }),
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder('utf-8');
  let buffer = '';

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });

    // SSE events are separated by double newlines
    const parts = buffer.split('\n\n');
    buffer = parts.pop();

    for (const part of parts) {
      if (part.startsWith('data:')) {
        const json = part.slice(5).trim();
        if (json === '[DONE]') return;
        const evt = JSON.parse(json);
        const delta = evt.choices?.[0]?.delta?.content;
        if (delta) process.stdout.write(delta);
      }
    }
  }
}

streamChat().catch(console.error);

Using OpenAI Client Libraries

Because the API maintains compatibility with the OpenAI specification, you can point official SDKs at your AnythingLLM instance by overriding the base URL.

Python Example

import openai

openai.api_key = "YOUR_API_KEY"
openai.base_url = "https://your-instance.com/api/v1/openai"

response = openai.ChatCompletion.create(
    model="sales-assistant",
    messages=[
        {"role": "system", "content": "You are a concise helpdesk bot."},
        {"role": "user", "content": "How do I reset my password?"}
    ],
    temperature=0.5,
    stream=False,
)

print(response.choices[0].message.content)

The SDK automatically routes requests to /chat/completions and handles the JSON serialization exactly as the Express server expects.

Telemetry and Error Handling

After each chat completion, the endpoint triggers Telemetry.sendTelemetry and EventLogs.logEvent to record usage statistics and audit events. These calls are internal and do not affect the API response.

All routes wrap their logic in try/catch blocks. Unhandled exceptions return a 500 Internal Server Error status while logging details to the console for server-side diagnostics.

Key Source Files

Summary

  • Authenticate all requests with a Bearer token generated in the AnythingLLM UI, verified by the validApiKey middleware.
  • Discover workspaces by calling GET /api/v1/openai/models, which returns workspace slugs usable as model IDs.
  • Send chats to POST /api/v1/openai/chat/completions using standard OpenAI payload structures.
  • Choose streaming by setting stream: true to receive SSE deltas, or stream: false for complete JSON responses.
  • Reuse existing code by pointing OpenAI SDKs at https://your-instance.com/api/v1/openai.

Frequently Asked Questions

How do I generate an API key for the AnythingLLM Developer API?

Navigate to Settings → API Keys in the AnythingLLM web interface and create a new key. The secret is stored in the api_keys table via server/models/apiKeys.js and must be included as a Bearer token in the Authorization header of every request.

What is the difference between a workspace slug and a model ID in the API?

In the AnythingLLM Developer API, workspace slugs function as model identifiers. When you call GET /v1/openai/models, the id field of each returned object is the workspace slug (e.g., sales-assistant). You pass this slug as the model parameter in chat completion requests to route the query to that specific workspace's configuration and LLM provider.

Can I use the official OpenAI Python or Node.js SDK with AnythingLLM?

Yes. The API is OpenAI-compatible, so you can set openai.base_url (Python) or the equivalent configuration in Node.js to https://your-instance.com/api/v1/openai. The SDK will communicate with the AnythingLLM endpoints exactly as if it were talking to OpenAI's servers, allowing seamless integration with existing codebases.

How does streaming work in the chat completions endpoint?

When you set stream: true in the request payload, server/endpoints/api/openai/index.js delegates to OpenAICompatibleChat.streamChat, which sets SSE headers (Content-Type: text/event-stream) and writes incremental response chunks to the client. Each chunk is a JSON object containing delta content, terminated by a data: [DONE] message.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →