How to Use FreeLLMAPI with the OpenAI Python Client: Complete Integration Guide

Point the OpenAI Python SDK at your FreeLLMAPI router's /v1 endpoint using a unified API key from the dashboard, then invoke standard methods like chat.completions.create() with model="auto" to route requests to the best available free-tier LLM provider.

FreeLLMAPI is a self-hosted aggregation router that exposes a single OpenAI-compatible API surface while distributing requests across dozens of free-tier LLM providers. Because it mirrors the exact request and response schemas used by OpenAI's official endpoints, you can swap it into existing Python applications by changing only the base_url and api_key parameters.

Configuring the OpenAI SDK for FreeLLMAPI

To route requests through FreeLLMAPI, initialize the OpenAI client with the router's local address and a unified API key generated from the web dashboard.

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",      # FreeLLMAPI router endpoint

    api_key="freellmapi-your-unified-key",    # From dashboard > Keys

)

The unified API key authorizes your request against the router; individual provider keys are encrypted in a local SQLite database and decrypted only at request time. According to the source code in server/src/routes/proxy.ts, all standard OpenAI endpoints—including /v1/chat/completions, /v1/embeddings, and /v1/audio/*—are implemented and forwarded to the internal routing engine.

Selecting Models and Routing Strategies

FreeLLMAPI supports both automatic and explicit model selection via the model parameter.

  • model="auto" – Lets the router choose the highest-priority healthy provider based on speed and intelligence scores.
  • model="auto:fast" – Prefers providers with lowest latency.
  • model="auto:smart" – Prefers providers with higher capability scores.
  • Specific model IDs – Bypass auto-routing and target a specific upstream model directly.

The routing logic resides in server/src/services/router.ts, which evaluates health checks, rate-limit counters, and fall-over policies before selecting an upstream.

Executing Chat Completions

Non-Streaming Requests

Send chat completion requests exactly as you would with the native OpenAI API. The router will select a provider and return the response in the standard format.

response = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Summarize the fall of Rome in one sentence."}]
)

print(response.choices[0].message.content)
print("Routed via:", response.headers.get("x-routed-via"))  # Provider that served the request

Streaming Responses

Enable streaming to receive tokens incrementally. The router maintains the connection to the upstream provider and streams chunks back in real-time.

for chunk in client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Write a haiku about clouds."}],
    stream=True,
):
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Using Tool Calling and Embeddings

The OpenAI-compatible provider abstraction in server/src/providers/openai-compat.ts normalizes tool-calling protocols and embedding formats across heterogeneous upstreams.

Function Calling

Define tools using the standard OpenAI schema; the router translates the request to the appropriate provider and returns the function arguments.

response = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "What is 12 * 8?"}],
    tools=[{
        "type": "function",
        "function": {
            "name": "run_math",
            "description": "Simple arithmetic",
            "parameters": {
                "type": "object",
                "properties": {"expression": {"type": "string"}}
            }
        }
    }]
)

print(response.choices[0].message.tool_calls)

Text Embeddings

Request embeddings using model="auto" to let FreeLLMAPI route to an embedding-capable provider.

embed = client.embeddings.create(
    model="auto",
    input="FreeLLMAPI aggregates free tiers from many LLM providers."
)

print(embed.data[0].embedding[:5])  # First five vector components

How Requests Flow Through the Router

When the Python client sends a request to base_url, the following sequence occurs:

  1. Entry Point – server/src/routes/proxy.ts receives the HTTP request and validates the unified API key.
  2. Routing Decision – server/src/services/router.ts checks provider health, rate limits, and speed/intelligence scores to select the best upstream.
  3. Provider Normalization – server/src/providers/openai-compat.ts adapts the request for the chosen provider and normalizes the response back to standard OpenAI format.
  4. Response – The router returns the result to the Python client with optional metadata headers like x-routed-via.

Summary

  • Initialization requires only changing base_url to your FreeLLMAPI router address (ending in /v1) and providing a unified API key from the dashboard.
  • Auto-routing via model="auto" leverages the logic in server/src/services/router.ts to select healthy, rate-available providers automatically.
  • Full compatibility with streaming, tool-calling, vision inputs, and embeddings is implemented in server/src/routes/proxy.ts and server/src/providers/openai-compat.ts.
  • Provider security is handled internally; the Python client never sees individual upstream keys, only the encrypted SQLite store managed by the router.

Frequently Asked Questions

What base URL should I use to connect the OpenAI client to FreeLLMAPI?

Use http://<your-router-host>:3001/v1 (or the port you configured when starting the Docker container). The /v1 path is mandatory because it signals the OpenAI-compatible API version that the router implements in server/src/routes/proxy.ts.

How does the automatic model selection decide which provider to use?

The router evaluates configurable priority scores, health check status, and rate-limit counters for each configured provider. According to server/src/services/router.ts, it selects the highest-priority provider that is currently healthy and not rate-limited, preferring faster or smarter profiles when you specify auto:fast or auto:smart.

Can I use local models like llama.cpp with the OpenAI Python client through FreeLLMAPI?

Yes. The server/src/providers/openai-compat.ts abstraction treats any OpenAI-compatible endpoint—whether a remote API or a local llama.cpp server—as a valid upstream. Configure the local endpoint in the FreeLLMAPI dashboard, and the router will include it in the pool of candidates for model="auto" or route to it directly by ID.

Where can I find the complete list of supported endpoints and parameters?

The full API contract, including details on streaming, audio endpoints, and vision inputs, is documented in docs/api.md within the repository. This file specifies that all standard OpenAI parameters—including temperature, max_tokens, top_p, and tools—are supported and forwarded appropriately by the proxy layer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →