How to Use FreeLLMAPI with Standard OpenAI Client Libraries

Yes—FreeLLMAPI is fully compatible with standard OpenAI client libraries by simply pointing them to the local /v1 endpoint and using your unified API key.

FreeLLMAPI implements the complete OpenAI-compatible API surface, making it a drop-in replacement for developers who want unified access to multiple LLM providers without rewriting their existing code. This guide walks through the exact configuration needed for Python, Node.js, and LangChain implementations, with references to the underlying source code that powers this compatibility.

How FreeLLMAPI Achieves OpenAI Compatibility

The compatibility layer is built into three core components of the FreeLLMAPI server.

The /v1 Router Registration

In server/src/app.ts, the application mounts the OpenAI-compatible router under the /v1 path:

// Lines 81-92 in server/src/app.ts
app.use('/v1', openaiRouter);

This single line exposes all standard OpenAI endpoints—including /v1/chat/completions, /v1/models, and /v1/embeddings—at the exact URLs that official OpenAI clients expect.

Request Proxying and Validation

The server/src/routes/proxy.ts file handles the core translation logic. When a request arrives at any /v1 endpoint, the proxy:

  1. Validates the unified API key from Authorization, x-api-key, or x-goog-api-key headers (lines 70-84)
  2. Parses the OpenAI request format unchanged (lines 56-62)
  3. Forwards to the selected provider while maintaining the same wire format

This means your existing OpenAI code requires zero structural changes—only the base_url and api_key values differ.

Official Documentation Confirmation

The project documentation explicitly confirms this behavior in docs/clients.md (lines 15-20), stating that any OpenAI SDK—including Python, Node.js, LangChain, and LlamaIndex—works by setting base_url to http://localhost:3001/v1 with the unified dashboard key.

Python OpenAI SDK Configuration (v1.0+)

The modern openai Python package uses a client-based approach. Configure it for FreeLLMAPI as follows:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_UNIFIED_KEY",           # from FreeLLMAPI dashboard

    base_url="http://localhost:3001/v1"   # local FreeLLMAPI server

)

response = client.chat.completions.create(
    model="auto",                          # let FreeLLMAPI route optimally

    messages=[{"role": "user", "content": "Explain quantum computing"}],
    temperature=0.7,
    max_tokens=500
)

print(response.choices[0].message.content)

Key points:

  • Use model="auto" to enable FreeLLMAPI's automatic provider selection
  • The response object maintains identical structure to official OpenAI responses
  • All standard parameters (temperature, top_p, stream, etc.) pass through unchanged

Node.js OpenAI SDK Configuration

The Node.js SDK follows the same pattern with camelCase conventions:

import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.FREELLM_API_KEY,      // your unified key
  baseURL: 'http://localhost:3001/v1',      // FreeLLMAPI endpoint
});

async function generate() {
  const completion = await client.chat.completions.create({
    model: 'auto',
    messages: [{ role: 'user', content: 'Write a haiku about recursion' }],
    stream: false,                            // set true for streaming
  });

  console.log(completion.choices[0].message.content);
}

generate();

For streaming responses, identical to OpenAI's implementation:

const stream = await client.chat.completions.create({
  model: 'auto',
  messages: [{ role: 'user', content: 'Count to 10 slowly' }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || '');
}

LangChain Integration

LangChain's OpenAI wrappers work without modification when environment variables are set:

import os
from langchain.chat_models import ChatOpenAI
from langchain.schema import HumanMessage, SystemMessage

os.environ["OPENAI_API_KEY"] = "YOUR_UNIFIED_KEY"
os.environ["OPENAI_API_BASE"] = "http://localhost:3001/v1"

chat = ChatOpenAI(
    model_name="auto",           # FreeLLMAPI handles routing

    temperature=0.5,
)

messages = [
    SystemMessage(content="You are a helpful coding assistant."),
    HumanMessage(content="How do I handle async/await in Python?")
]

response = chat(messages)
print(response.content)

For LangChain Expression Language (LCEL) chains:

from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

prompt = ChatPromptTemplate.from_template("Explain {topic} in three bullet points")
chain = prompt | chat | StrOutputParser()

result = chain.invoke({"topic": "vector databases"})
print(result)

LlamaIndex and Other Frameworks

Any framework with OpenAI compatibility follows the same pattern. For LlamaIndex:

import os
from llama_index.llms.openai import OpenAI

os.environ["OPENAI_API_KEY"] = "YOUR_UNIFIED_KEY"
os.environ["OPENAI_API_BASE"] = "http://localhost:3001/v1"

llm = OpenAI(model="auto", temperature=0.3)
response = llm.complete("Summarize transformer architecture")
print(response.text)

Environment Variable Setup Best Practices

Rather than hardcoding values, use environment-specific configuration:


# .env file for local development

FREELLM_API_KEY=flm_your_unified_key_here
FREELLM_BASE_URL=http://localhost:3001/v1

Then in application code:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.getenv("FREELLM_API_KEY"),
    base_url=os.getenv("FREELLM_BASE_URL")
)

This pattern enables seamless switching between FreeLLMAPI (local), staging, and direct provider configurations without code changes.

Supported OpenAI-Compatible Endpoints

Based on the router implementation in server/src/app.ts and proxy handlers in server/src/routes/proxy.ts, FreeLLMAPI supports:

Endpoint Purpose OpenAI SDK Method
POST /v1/chat/completions Chat inference chat.completions.create()
GET /v1/models List available models models.list()
POST /v1/embeddings Text embeddings embeddings.create()
POST /v1/completions Legacy completions completions.create()

All endpoints accept and return the exact JSON schemas defined in OpenAI's API specification.

Obtaining Your Unified API Key

  1. Start the FreeLLMAPI server locally: npm run dev (see README.md)
  2. Open the dashboard at http://localhost:3001
  3. Navigate to API Keys → copy your unified key
  4. Configure any OpenAI client with this key and http://localhost:3001/v1 as the base URL

Summary

  • FreeLLMAPI implements the complete OpenAI API surface through its /v1 router in server/src/app.ts
  • Zero code changes required—only base_url and api_key configuration differs from standard OpenAI usage
  • Use model="auto" to leverage FreeLLMAPI's intelligent provider routing (optional but recommended)
  • All major frameworks supported: Python SDK, Node.js SDK, LangChain, LlamaIndex, and any OpenAI-compatible client
  • Request handling is verified in server/src/routes/proxy.ts lines 56-84, ensuring format fidelity with official OpenAI APIs

Frequently Asked Questions

What is the exact base URL for FreeLLMAPI?

Use http://localhost:3001/v1 when running locally. The /v1 path is required—this is where server/src/app.ts mounts the OpenAI-compatible router. For deployed instances, replace localhost:3001 with your server's address.

Do I need to modify my existing OpenAI code?

No. FreeLLMAPI maintains wire-format compatibility with OpenAI's API. The only changes needed are setting base_url (or baseURL) to your FreeLLMAPI endpoint and using your unified API key instead of an OpenAI key.

Can I use streaming responses with FreeLLMAPI?

Yes. The proxy in server/src/routes/proxy.ts passes through streaming requests unchanged. Use stream=True (Python) or stream: true (Node.js) exactly as you would with the official OpenAI API.

What happens when I use model="auto"?

FreeLLMAPI's proxy layer selects the best-available provider based on your configured providers, current rate limits, and cost preferences. You can still specify exact model names (e.g., gpt-4, claude-3-opus) if you need a specific provider.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →