How to Connect an OpenAI-Compatible Client to FreeLLMAPI: Complete Setup Guide
Point any standard OpenAI SDK or HTTP client to http://localhost:3000/v1 and authenticate with a FreeLLMAPI-generated bearer token to route requests through the self-hosted OpenAI-compatible proxy.
FreeLLMAPI is a self-hosted server that exposes an OpenAI-compatible HTTP API, allowing you to connect an OpenAI-compatible client to FreeLLMAPI and access diverse LLM backends—including Groq, Cloudflare, and OpenRouter—without modifying your application code. The project, available at tashfeenahmed/freellmapi, automatically translates requests and normalizes responses through its provider abstraction layer. This guide details the exact configuration parameters, authentication headers, and code patterns required to integrate any OpenAI SDK with your local instance.
How the OpenAI-Compatible Layer Works
The integration relies on the OpenAICompatProvider class located in server/src/providers/openai-compat.ts. This provider acts as a bidirectional adapter: it accepts standard OpenAI-formatted requests and converts them into upstream-specific payloads, then normalizes responses back into the OpenAI schema.
When you instantiate a client, two critical parameters route traffic through FreeLLMAPI:
- Base URL: The
baseURLparameter (or raw HTTP endpoint) must point to your running FreeLLMAPI instance, typicallyhttp://localhost:3000/v1. The provider constructs final request URLs by appending OpenAI-style paths tothis.baseUrlinside theOpenAICompatProviderconstructor andchatCompletion()method【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L63-L66】【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L41-L44】. - API Key: FreeLLMAPI requires a bearer token for authentication. The
authHeader()method automatically injectsAuthorization: Bearer <key>into outgoing requests unless the upstream provider is configured as "keyless"【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L45-L48】. The server validates this key viavalidateKey()before proxying to upstream services【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L63-L64】.
Server Setup and Base URL Configuration
Before configuring clients, start the FreeLLMAPI server and obtain authentication credentials.
Start the Server
You can launch FreeLLMAPI using Docker (recommended) or Node.js directly. The server listens on port 3000 by default and exposes OpenAI-compatible endpoints under the /v1 path (e.g., http://localhost:3000/v1/chat/completions).
Using Docker:
docker compose up -d
Using npm:
npm install
npm run start
Obtain an API Key
Generate an API key through the FreeLLMAPI web UI, or set the FREELLMAPI_KEY environment variable in a .env file. The server stores keys in the api_keys table and validates them on every request through the validateKey() method in OpenAICompatProvider.
Client Configuration Examples
Any client that supports custom baseURL and apiKey parameters can connect to FreeLLMAPI. The following examples demonstrate Node.js, Python, and cURL configurations.
Node.js / JavaScript
Use the official openai npm package. Set baseURL to your FreeLLMAPI instance and apiKey to your generated key:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "http://localhost:3000/v1", // 👈 FreeLLMAPI endpoint
apiKey: process.env.FREELLMAPI_KEY, // 👈 Your generated key
});
const result = await client.chat.completions.create({
model: "openai/gpt-4.1", // Supported model identifier
messages: [{ role: "user", content: "Hello, world!" }],
});
console.log(result.choices[0].message.content);
The chatCompletion() method in server/src/providers/openai-compat.ts handles the underlying request transformation and forwarding to the appropriate upstream provider【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L41-L44】.
Python
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3000/v1",
api_key=os.environ.get("FREELLMAPI_KEY")
)
response = client.chat.completions.create(
model="openai/gpt-4.1",
messages=[{"role": "user", "content": "Explain quantum entanglement"}]
)
print(response.choices[0].message.content)
cURL
For direct HTTP testing, include the Authorization header with your bearer token:
curl http://localhost:3000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $FREELLMAPI_KEY" \
-d '{
"model": "openai/gpt-4.1",
"messages": [{"role":"user","content":"Hello"}]
}'
Selecting Models and Advanced Parameters
FreeLLMAPI exposes available models at the GET /v1/models endpoint. Query this to see which identifiers—such as openai/gpt-4.1, groq/llama-3, or openrouter/claude-3—are currently configured:
curl http://localhost:3000/v1/models -H "Authorization: Bearer $FREELLMAPI_KEY"
All standard OpenAI request fields—including temperature, max_tokens, top_p, and tools—are supported. The provider sanitizes platform-specific quirks in the resolveParallelToolCalls() method, which adjusts parameters like forced parallel_tool_calls for NVIDIA NIM compatibility【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L90-L92】.
Streaming Responses and Function Calling
FreeLLMAPI supports streaming and tool use through the same chatCompletion() pipeline.
Streaming Example (Node.js)
const stream = await client.chat.completions.create({
model: "openai/gpt-4.1",
messages: [{ role: "user", content: "Write a haiku about AI." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0].delta?.content || "");
}
Tool/Function Calling Example
const response = await client.chat.completions.create({
model: "openai/gpt-4.1",
messages: [{ role: "user", content: "What is the weather in London?" }],
tools: [
{
type: "function",
function: {
name: "get_weather",
description: "Fetch current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
},
],
tool_choice: "auto",
});
Both patterns work because the OpenAICompatProvider normalizes the request payload before forwarding it to the selected upstream backend.
Summary
- Base URL Configuration: Set your client’s
baseURLtohttp://localhost:3000/v1(or your deployed instance URL) to route requests through FreeLLMAPI. - Authentication: Provide a FreeLLMAPI-generated key as the
Authorization: Bearerheader; theauthHeader()method inserver/src/providers/openai-compat.tshandles injection. - Request Handling: The
OpenAICompatProviderclass translates OpenAI-formatted requests to upstream provider formats inchatCompletion()and normalizes responses back to the OpenAI schema. - Feature Support: Streaming, function calling, and standard parameters work out-of-the-box, with platform-specific adjustments handled automatically in
resolveParallelToolCalls(). - Model Discovery: Query
GET /v1/modelsto see available model identifiers mapped by your FreeLLMAPI instance.
Frequently Asked Questions
Do I need to modify my existing OpenAI integration code to use FreeLLMAPI?
No. Any client that allows overriding the baseURL and apiKey parameters can connect to FreeLLMAPI without code changes. Simply redirect the endpoint to http://localhost:3000/v1 and use your FreeLLMAPI key as the API key. The OpenAICompatProvider handles all translation between OpenAI's request/response format and the upstream provider formats.
Where does FreeLLMAPI validate the API key?
The server validates bearer tokens in the validateKey() method inside server/src/providers/openai-compat.ts【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L63-L64】. This method checks the key against the api_keys table before proxying the request to the selected LLM provider. If validation fails, the request returns a 401 error before reaching any upstream service.
Can I use streaming and function calling with FreeLLMAPI?
Yes. Both streaming responses and tool/function calling are fully supported through the standard OpenAI SDK interfaces. The chatCompletion() method processes these requests and manages the translation of streaming chunks and tool schemas between the OpenAI format and upstream provider requirements【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L41-L44】.
What models are available when connecting to FreeLLMAPI?
Available models depend on your configured upstream providers. Query the GET /v1/models endpoint to receive a list of supported model identifiers (e.g., openai/gpt-4.1, groq/llama-3). The OpenAICompatProvider maps these identifiers to the correct upstream endpoints and request formats based on your server configuration defined in server/src/providers/openai-compat.ts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →