How to Use FreeLLMAPI with the OpenAI Python Client: Complete Integration Guide
Point the OpenAI Python SDK at your FreeLLMAPI router's /v1 endpoint using a unified API key from the dashboard, then invoke standard methods like chat.completions.create() with model="auto" to route requests to the best available free-tier LLM provider.
FreeLLMAPI is a self-hosted aggregation router that exposes a single OpenAI-compatible API surface while distributing requests across dozens of free-tier LLM providers. Because it mirrors the exact request and response schemas used by OpenAI's official endpoints, you can swap it into existing Python applications by changing only the base_url and api_key parameters.
Configuring the OpenAI SDK for FreeLLMAPI
To route requests through FreeLLMAPI, initialize the OpenAI client with the router's local address and a unified API key generated from the web dashboard.
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3001/v1", # FreeLLMAPI router endpoint
api_key="freellmapi-your-unified-key", # From dashboard > Keys
)
The unified API key authorizes your request against the router; individual provider keys are encrypted in a local SQLite database and decrypted only at request time. According to the source code in server/src/routes/proxy.ts, all standard OpenAI endpoints—including /v1/chat/completions, /v1/embeddings, and /v1/audio/*—are implemented and forwarded to the internal routing engine.
Selecting Models and Routing Strategies
FreeLLMAPI supports both automatic and explicit model selection via the model parameter.
model="auto"– Lets the router choose the highest-priority healthy provider based on speed and intelligence scores.model="auto:fast"– Prefers providers with lowest latency.model="auto:smart"– Prefers providers with higher capability scores.- Specific model IDs – Bypass auto-routing and target a specific upstream model directly.
The routing logic resides in server/src/services/router.ts, which evaluates health checks, rate-limit counters, and fall-over policies before selecting an upstream.
Executing Chat Completions
Non-Streaming Requests
Send chat completion requests exactly as you would with the native OpenAI API. The router will select a provider and return the response in the standard format.
response = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Summarize the fall of Rome in one sentence."}]
)
print(response.choices[0].message.content)
print("Routed via:", response.headers.get("x-routed-via")) # Provider that served the request
Streaming Responses
Enable streaming to receive tokens incrementally. The router maintains the connection to the upstream provider and streams chunks back in real-time.
for chunk in client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Write a haiku about clouds."}],
stream=True,
):
print(chunk.choices[0].delta.content or "", end="", flush=True)
Using Tool Calling and Embeddings
The OpenAI-compatible provider abstraction in server/src/providers/openai-compat.ts normalizes tool-calling protocols and embedding formats across heterogeneous upstreams.
Function Calling
Define tools using the standard OpenAI schema; the router translates the request to the appropriate provider and returns the function arguments.
response = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "What is 12 * 8?"}],
tools=[{
"type": "function",
"function": {
"name": "run_math",
"description": "Simple arithmetic",
"parameters": {
"type": "object",
"properties": {"expression": {"type": "string"}}
}
}
}]
)
print(response.choices[0].message.tool_calls)
Text Embeddings
Request embeddings using model="auto" to let FreeLLMAPI route to an embedding-capable provider.
embed = client.embeddings.create(
model="auto",
input="FreeLLMAPI aggregates free tiers from many LLM providers."
)
print(embed.data[0].embedding[:5]) # First five vector components
How Requests Flow Through the Router
When the Python client sends a request to base_url, the following sequence occurs:
- Entry Point –
server/src/routes/proxy.tsreceives the HTTP request and validates the unified API key. - Routing Decision –
server/src/services/router.tschecks provider health, rate limits, and speed/intelligence scores to select the best upstream. - Provider Normalization –
server/src/providers/openai-compat.tsadapts the request for the chosen provider and normalizes the response back to standard OpenAI format. - Response – The router returns the result to the Python client with optional metadata headers like
x-routed-via.
Summary
- Initialization requires only changing
base_urlto your FreeLLMAPI router address (ending in/v1) and providing a unified API key from the dashboard. - Auto-routing via
model="auto"leverages the logic inserver/src/services/router.tsto select healthy, rate-available providers automatically. - Full compatibility with streaming, tool-calling, vision inputs, and embeddings is implemented in
server/src/routes/proxy.tsandserver/src/providers/openai-compat.ts. - Provider security is handled internally; the Python client never sees individual upstream keys, only the encrypted SQLite store managed by the router.
Frequently Asked Questions
What base URL should I use to connect the OpenAI client to FreeLLMAPI?
Use http://<your-router-host>:3001/v1 (or the port you configured when starting the Docker container). The /v1 path is mandatory because it signals the OpenAI-compatible API version that the router implements in server/src/routes/proxy.ts.
How does the automatic model selection decide which provider to use?
The router evaluates configurable priority scores, health check status, and rate-limit counters for each configured provider. According to server/src/services/router.ts, it selects the highest-priority provider that is currently healthy and not rate-limited, preferring faster or smarter profiles when you specify auto:fast or auto:smart.
Can I use local models like llama.cpp with the OpenAI Python client through FreeLLMAPI?
Yes. The server/src/providers/openai-compat.ts abstraction treats any OpenAI-compatible endpoint—whether a remote API or a local llama.cpp server—as a valid upstream. Configure the local endpoint in the FreeLLMAPI dashboard, and the router will include it in the pool of candidates for model="auto" or route to it directly by ID.
Where can I find the complete list of supported endpoints and parameters?
The full API contract, including details on streaming, audio endpoints, and vision inputs, is documented in docs/api.md within the repository. This file specifies that all standard OpenAI parameters—including temperature, max_tokens, top_p, and tools—are supported and forwarded appropriately by the proxy layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →