FreeLLMAPI Endpoints for OpenAI Compatibility: Complete API Reference
FreeLLMAPI exposes a full OpenAI-compatible HTTP surface under the /v1 namespace, routing requests to free-tier LLM providers with automatic failover and unified response normalization.
FreeLLMAPI is an open-source proxy server that implements OpenAI's API specification across multiple free model providers. This article documents every endpoint available for OpenAI compatibility, drawn directly from the tashfeenahmed/freellmaapi source code and documentation.
Core OpenAI-Compatible Endpoints
FreeLLMAPI mirrors the complete OpenAI API surface under the /v1 path prefix. All endpoints accept the same request payloads and return identically-shaped responses, enabling drop-in replacement for OpenAI's base URL.
Chat Completions Endpoint
The primary interface for conversational AI:
POST /v1/chat/completions
Supports streaming, tool calling, vision input, and routing strategies (auto, auto:fast, auto:smart). The router selects the best available free model based on your configured strategy.
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3001/v1",
api_key="freellmapi-your-unified-key",
)
resp = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Explain quantum computing"}],
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))
Legacy Completions Endpoint
Text completion support for editor integrations and legacy clients:
POST /v1/completions
Used primarily for ghost-text autocomplete in IDEs.
Responses Endpoint
Codex-style code generation compatibility:
POST /v1/responses
Handles structured code generation requests in the OpenAI Responses format.
Multimodal and Specialized Endpoints
Embeddings Endpoint
Vector embedding generation with family-aware failover:
POST /v1/embeddings
Failover logic stays within the same model family to preserve embedding comparability. The response includes normalized vector dimensions regardless of which provider handled the request.
Image Generation
POST /v1/images/generations
Routes to free-tier providers exposing image models. Returns standard OpenAI image response format with base64 data or URL references.
Video Generation
POST /v1/videos/generations
Experimental endpoint for short MP4 video generation. Limited to supporting providers only.
Audio Processing
Two distinct audio interfaces:
| Endpoint | Method | Purpose |
|---|---|---|
/v1/audio/speech |
POST | Text-to-speech synthesis |
/v1/audio/transcriptions |
POST | Speech-to-text conversion (limited provider support) |
Utility and Discovery Endpoints
Models List
GET /v1/models
Returns the catalog of available models in OpenAI-shaped format, filtered by your account's enabled providers and routing strategy.
Documentation Endpoints
| Path | Format | Purpose |
|---|---|---|
/v1/docs |
HTML UI | Interactive API documentation |
/v1/openapi.json |
JSON | Raw OpenAPI specification |
These endpoints are enumerated in the repository's README.md (lines 51-53) and fully documented in docs/api.md (lines 9-23).
Extended Compatibility Surfaces
Beyond strict OpenAI compatibility, FreeLLMAPI exposes additional interfaces for broader ecosystem integration.
Anthropic Messages API
Full Claude client compatibility:
POST /v1/messages
Accepts Anthropic's native Messages API format and normalizes responses. Use with official Anthropic SDKs by overriding the base URL:
export ANTHROPIC_BASE_URL=http://localhost:3001
export ANTHROPIC_AUTH_TOKEN=freellmapi-your-unified-key
Native Gemini Interface
Direct Google Gemini API access:
/v1beta/models
/v1beta/models/{model}:generateContent
Preserves Gemini's native request/response formats for applications requiring Gemini-specific features.
Ollama Emulation
Local-model client compatibility via NDJSON surface:
/api/chat
/api/generate
Enables tools like Ollama-compatible UIs and CLI clients to use FreeLLMAPI's pooled providers.
Revocable URL Tokens
Header-less authentication for constrained environments:
/v1/t/{token}/[endpoint]
Useful for browser extensions, webhook integrations, or agents that cannot manipulate HTTP headers directly.
Architectural Implementation
Request Flow Through the System
- Entry point:
server/src/index.ts— Express server registering all/v1/*routes - Authentication: Unified bearer token validation
- Routing: Strategy-based provider selection (
auto,auto:fast,auto:smart) - Translation:
server/src/providers/openai-compat.tstransforms OpenAI schema to provider-native format - Failover: Automatic retry with
X-Fallback-AttemptsandX-Routed-Viaheaders - Normalization: Response unified to OpenAI format regardless of source provider
Key Source Files
| File | Role |
|---|---|
server/src/providers/openai-compat.ts |
Adapter implementing OpenAI-compatible endpoints |
server/src/index.ts |
Main server; route registration and authentication |
docs/api.md |
Human-readable API reference |
README.md |
Feature overview with endpoint list |
docs/architecture.md |
Routing, fallback, and provider integration details |
Summary
- FreeLLMAPI endpoints for OpenAI compatibility cover the complete OpenAI surface: chat completions, legacy completions, responses, embeddings, images, video, audio, and model discovery
- All endpoints live under
/v1with identical request/response shapes to OpenAI's API - Extended surfaces include Anthropic Messages (
/v1/messages), native Gemini (/v1beta/*), and Ollama emulation (/api/*) - Authentication uses a unified bearer token with optional revocable URL tokens for restricted environments
- The router in
server/src/index.tsand adapter inserver/src/providers/openai-compat.tshandle translation, failover, and response normalization automatically
Frequently Asked Questions
Does FreeLLMAPI support streaming responses?
Yes. All streaming-capable OpenAI endpoints including /v1/chat/completions support Server-Sent Events (SSE) streaming. The proxy maintains the stream across provider failovers when possible, or falls back to non-streaming responses if necessary.
How does the auto routing strategy select providers?
The auto strategy evaluates current rate limits, latency, and model capability against your request requirements. Variants like auto:fast prioritize response time while auto:smart optimizes for output quality. The X-Routed-Via response header reveals which provider handled your request.
Can I use my existing OpenAI SDK code without modification?
Yes. Change only the base_url to http://localhost:3001/v1 (or your deployed instance) and replace the api_key with your FreeLLMAPI unified key. All method signatures, parameter names, and response parsing remain identical to OpenAI's official SDK.
What happens when all providers for a request type are rate-limited?
The router returns a structured error response matching OpenAI's error format with HTTP 429 status, including Retry-After guidance when available from provider responses. The X-Fallback-Attempts header indicates how many providers were exhausted.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →