FreeLLMAPI OpenAI-Compatible Endpoint: Base URL and Authentication Guide

The OpenAI-compatible endpoint for FreeLLMAPI is http://localhost:3001/v1, which exposes standard OpenAI API routes like /v1/chat/completions and authenticates using a single unified API key from the dashboard.

FreeLLMAPI acts as a request router that aggregates multiple free-tier LLM providers behind one OpenAI-compatible interface. This architecture allows you to point existing OpenAI SDKs and tools at the FreeLLMAPI server without code modifications, while the system internally handles provider selection, rate-limit management, and credential rotation. The primary keyword refers to this unified entry point that makes the diverse provider pool appear as a single standard API.

Base URL and Available Routes

When the FreeLLMAPI server runs locally with default configuration, it exposes the OpenAI-compatible endpoint at:


http://localhost:3001/v1

Under this base path, all standard OpenAI API routes are available, including:

  • /v1/chat/completions for chat-based inference
  • /v1/completions for legacy text completion
  • /v1/models for listing available models

According to the docs/api.md source file, these endpoints accept the exact same request bodies as the official OpenAI API. The server/src/services/router.ts component intercepts incoming requests at these paths and transparently forwards them to the appropriate free-tier backend based on current availability and load.

Authentication with Unified API Keys

Authentication uses a unified API key generated from the dashboard's Keys page. This single credential replaces individual provider API keys and must be included in the Authorization header using the Bearer scheme:

api_key="freellmapi-your-unified-key"

The router validates this key before processing requests. Unlike direct provider integrations, you do not need to manage separate credentials for each backend service; the unified key grants access to the entire pool of available models.

Routing Architecture and Provider Selection

The core logic resides in server/src/services/router.ts, which implements the mapping layer between OpenAI-protocol requests and specific free-tier providers. This service handles:

  • Automatic model selection when requests specify "model": "auto", choosing the best available provider
  • Rate-limit management across multiple backend accounts to prevent quota exhaustion
  • Key rotation for load distribution and failover between providers

Because the router maintains strict protocol compatibility, responses return in the standard OpenAI format regardless of which underlying provider (such as Groq, Cerebras, or others) processes the request.

Integration Code Examples

You can integrate with the OpenAI-compatible endpoint using any standard OpenAI client library by changing only the base_url parameter.

Python (OpenAI SDK)

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="freellmapi-your-unified-key",
)

response = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Explain the routing architecture."}]
)

print(response.choices[0].message.content)

cURL

curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "What is the OpenAI-compatible endpoint?"}]
  }'

Node.js (OpenAI SDK)

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:3001/v1",
  apiKey: "freellmapi-your-unified-key",
});

const response = await client.chat.completions.create({
  model: "auto",
  messages: [{ role: "user", content: "How does the router work?" }]
});

console.log(response.choices[0].message.content);

URL Tokens for Headerless Access

For clients that cannot easily modify headers (such as embedded devices or certain mobile environments), server/src/services/url-tokens.ts implements revocable URL tokens. These tokens embed authentication credentials directly into the request URL, allowing access to the OpenAI-compatible endpoint without requiring the standard Authorization header. This provides a flexible alternative when header injection is impractical.

Health Checks and API Discovery

The server/src/services/health.ts file supplies health monitoring endpoints and exposes the OpenAPI specification at:


http://localhost:3001/v1/openapi.json

This route returns the complete OpenAPI schema describing all available paths and parameters, enabling automatic client generation and API discovery tools to recognize the compatible interface.

Summary

  • The OpenAI-compatible endpoint for FreeLLMAPI is located at http://localhost:3001/v1 when running the default local server configuration.
  • Authentication requires a unified API key from the dashboard, passed as a Bearer token in the Authorization header.
  • Standard routes like /v1/chat/completions function identically to OpenAI's official API, with server/src/services/router.ts handling provider selection and failover.
  • Client integration requires only changing the base_url parameter in existing OpenAI SDK implementations.
  • Alternative access methods via URL tokens are supported through server/src/services/url-tokens.ts for scenarios where header modification is not possible.

Frequently Asked Questions

What is the exact base URL for the FreeLLMAPI OpenAI-compatible endpoint?

The base URL is http://localhost:3001/v1 when running the server locally with default settings on port 3001. All standard OpenAI API paths, such as /v1/chat/completions, append to this base URL to form complete endpoints like http://localhost:3001/v1/chat/completions.

Do I need separate API keys for each provider when using FreeLLMAPI?

No. FreeLLMAPI uses a single unified API key that you generate from the dashboard's Keys page. The internal router manages authentication with individual free-tier providers automatically, so you only maintain one credential for all backend access.

How does the router determine which model or provider to use?

When you specify "model": "auto" in your request body, the server/src/services/router.ts component applies internal logic to select the optimal available provider based on current rate limits and latency. You may also specify explicit model identifiers if you require a specific backend.

What are URL tokens used for in FreeLLMAPI?

URL tokens provide revocable access credentials that can be embedded directly in request URLs. As implemented in server/src/services/url-tokens.ts, these tokens allow clients that cannot easily set HTTP headers to authenticate with the OpenAI-compatible endpoint, offering flexibility for constrained environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →