How to Connect an OpenAI-Compatible Client to FreeLLMAPI: Complete Setup Guide

Point any standard OpenAI SDK or HTTP client to http://localhost:3000/v1 and authenticate with a FreeLLMAPI-generated bearer token to route requests through the self-hosted OpenAI-compatible proxy.

FreeLLMAPI is a self-hosted server that exposes an OpenAI-compatible HTTP API, allowing you to connect an OpenAI-compatible client to FreeLLMAPI and access diverse LLM backends—including Groq, Cloudflare, and OpenRouter—without modifying your application code. The project, available at tashfeenahmed/freellmapi, automatically translates requests and normalizes responses through its provider abstraction layer. This guide details the exact configuration parameters, authentication headers, and code patterns required to integrate any OpenAI SDK with your local instance.


How the OpenAI-Compatible Layer Works

The integration relies on the OpenAICompatProvider class located in server/src/providers/openai-compat.ts. This provider acts as a bidirectional adapter: it accepts standard OpenAI-formatted requests and converts them into upstream-specific payloads, then normalizes responses back into the OpenAI schema.

When you instantiate a client, two critical parameters route traffic through FreeLLMAPI:

  • Base URL: The baseURL parameter (or raw HTTP endpoint) must point to your running FreeLLMAPI instance, typically http://localhost:3000/v1. The provider constructs final request URLs by appending OpenAI-style paths to this.baseUrl inside the OpenAICompatProvider constructor and chatCompletion() method【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L63-L66】【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L41-L44】.
  • API Key: FreeLLMAPI requires a bearer token for authentication. The authHeader() method automatically injects Authorization: Bearer <key> into outgoing requests unless the upstream provider is configured as "keyless"【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L45-L48】. The server validates this key via validateKey() before proxying to upstream services【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L63-L64】.

Server Setup and Base URL Configuration

Before configuring clients, start the FreeLLMAPI server and obtain authentication credentials.

Start the Server

You can launch FreeLLMAPI using Docker (recommended) or Node.js directly. The server listens on port 3000 by default and exposes OpenAI-compatible endpoints under the /v1 path (e.g., http://localhost:3000/v1/chat/completions).

Using Docker:

docker compose up -d

Using npm:

npm install
npm run start

Obtain an API Key

Generate an API key through the FreeLLMAPI web UI, or set the FREELLMAPI_KEY environment variable in a .env file. The server stores keys in the api_keys table and validates them on every request through the validateKey() method in OpenAICompatProvider.


Client Configuration Examples

Any client that supports custom baseURL and apiKey parameters can connect to FreeLLMAPI. The following examples demonstrate Node.js, Python, and cURL configurations.

Node.js / JavaScript

Use the official openai npm package. Set baseURL to your FreeLLMAPI instance and apiKey to your generated key:

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "http://localhost:3000/v1",   // 👈 FreeLLMAPI endpoint
  apiKey: process.env.FREELLMAPI_KEY,    // 👈 Your generated key
});

const result = await client.chat.completions.create({
  model: "openai/gpt-4.1",              // Supported model identifier
  messages: [{ role: "user", content: "Hello, world!" }],
});

console.log(result.choices[0].message.content);

The chatCompletion() method in server/src/providers/openai-compat.ts handles the underlying request transformation and forwarding to the appropriate upstream provider【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L41-L44】.

Python

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3000/v1",
    api_key=os.environ.get("FREELLMAPI_KEY")
)

response = client.chat.completions.create(
    model="openai/gpt-4.1",
    messages=[{"role": "user", "content": "Explain quantum entanglement"}]
)
print(response.choices[0].message.content)

cURL

For direct HTTP testing, include the Authorization header with your bearer token:

curl http://localhost:3000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FREELLMAPI_KEY" \
  -d '{
    "model": "openai/gpt-4.1",
    "messages": [{"role":"user","content":"Hello"}]
  }'

Selecting Models and Advanced Parameters

FreeLLMAPI exposes available models at the GET /v1/models endpoint. Query this to see which identifiers—such as openai/gpt-4.1, groq/llama-3, or openrouter/claude-3—are currently configured:

curl http://localhost:3000/v1/models -H "Authorization: Bearer $FREELLMAPI_KEY"

All standard OpenAI request fields—including temperature, max_tokens, top_p, and tools—are supported. The provider sanitizes platform-specific quirks in the resolveParallelToolCalls() method, which adjusts parameters like forced parallel_tool_calls for NVIDIA NIM compatibility【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L90-L92】.


Streaming Responses and Function Calling

FreeLLMAPI supports streaming and tool use through the same chatCompletion() pipeline.

Streaming Example (Node.js)

const stream = await client.chat.completions.create({
  model: "openai/gpt-4.1",
  messages: [{ role: "user", content: "Write a haiku about AI." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0].delta?.content || "");
}

Tool/Function Calling Example

const response = await client.chat.completions.create({
  model: "openai/gpt-4.1",
  messages: [{ role: "user", content: "What is the weather in London?" }],
  tools: [
    {
      type: "function",
      function: {
        name: "get_weather",
        description: "Fetch current weather for a city",
        parameters: {
          type: "object",
          properties: { city: { type: "string" } },
          required: ["city"],
        },
      },
    },
  ],
  tool_choice: "auto",
});

Both patterns work because the OpenAICompatProvider normalizes the request payload before forwarding it to the selected upstream backend.


Summary

  • Base URL Configuration: Set your client’s baseURL to http://localhost:3000/v1 (or your deployed instance URL) to route requests through FreeLLMAPI.
  • Authentication: Provide a FreeLLMAPI-generated key as the Authorization: Bearer header; the authHeader() method in server/src/providers/openai-compat.ts handles injection.
  • Request Handling: The OpenAICompatProvider class translates OpenAI-formatted requests to upstream provider formats in chatCompletion() and normalizes responses back to the OpenAI schema.
  • Feature Support: Streaming, function calling, and standard parameters work out-of-the-box, with platform-specific adjustments handled automatically in resolveParallelToolCalls().
  • Model Discovery: Query GET /v1/models to see available model identifiers mapped by your FreeLLMAPI instance.

Frequently Asked Questions

Do I need to modify my existing OpenAI integration code to use FreeLLMAPI?

No. Any client that allows overriding the baseURL and apiKey parameters can connect to FreeLLMAPI without code changes. Simply redirect the endpoint to http://localhost:3000/v1 and use your FreeLLMAPI key as the API key. The OpenAICompatProvider handles all translation between OpenAI's request/response format and the upstream provider formats.

Where does FreeLLMAPI validate the API key?

The server validates bearer tokens in the validateKey() method inside server/src/providers/openai-compat.ts【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L63-L64】. This method checks the key against the api_keys table before proxying the request to the selected LLM provider. If validation fails, the request returns a 401 error before reaching any upstream service.

Can I use streaming and function calling with FreeLLMAPI?

Yes. Both streaming responses and tool/function calling are fully supported through the standard OpenAI SDK interfaces. The chatCompletion() method processes these requests and manages the translation of streaming chunks and tool schemas between the OpenAI format and upstream provider requirements【/cache/repos/github.com/tashfeenahmed/freellmapi/main/server/src/providers/openai-compat.ts#L41-L44】.

What models are available when connecting to FreeLLMAPI?

Available models depend on your configured upstream providers. Query the GET /v1/models endpoint to receive a list of supported model identifiers (e.g., openai/gpt-4.1, groq/llama-3). The OpenAICompatProvider maps these identifiers to the correct upstream endpoints and request formats based on your server configuration defined in server/src/providers/openai-compat.ts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →