FreeLLMAPI Endpoints for OpenAI Compatibility: Complete API Reference

FreeLLMAPI exposes a full OpenAI-compatible HTTP surface under the /v1 namespace, routing requests to free-tier LLM providers with automatic failover and unified response normalization.

FreeLLMAPI is an open-source proxy server that implements OpenAI's API specification across multiple free model providers. This article documents every endpoint available for OpenAI compatibility, drawn directly from the tashfeenahmed/freellmaapi source code and documentation.

Core OpenAI-Compatible Endpoints

FreeLLMAPI mirrors the complete OpenAI API surface under the /v1 path prefix. All endpoints accept the same request payloads and return identically-shaped responses, enabling drop-in replacement for OpenAI's base URL.

Chat Completions Endpoint

The primary interface for conversational AI:

POST /v1/chat/completions

Supports streaming, tool calling, vision input, and routing strategies (auto, auto:fast, auto:smart). The router selects the best available free model based on your configured strategy.

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="freellmapi-your-unified-key",
)

resp = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Explain quantum computing"}],
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))

Legacy Completions Endpoint

Text completion support for editor integrations and legacy clients:

POST /v1/completions

Used primarily for ghost-text autocomplete in IDEs.

Responses Endpoint

Codex-style code generation compatibility:

POST /v1/responses

Handles structured code generation requests in the OpenAI Responses format.

Multimodal and Specialized Endpoints

Embeddings Endpoint

Vector embedding generation with family-aware failover:

POST /v1/embeddings

Failover logic stays within the same model family to preserve embedding comparability. The response includes normalized vector dimensions regardless of which provider handled the request.

Image Generation

POST /v1/images/generations

Routes to free-tier providers exposing image models. Returns standard OpenAI image response format with base64 data or URL references.

Video Generation

POST /v1/videos/generations

Experimental endpoint for short MP4 video generation. Limited to supporting providers only.

Audio Processing

Two distinct audio interfaces:

Endpoint Method Purpose
/v1/audio/speech POST Text-to-speech synthesis
/v1/audio/transcriptions POST Speech-to-text conversion (limited provider support)

Utility and Discovery Endpoints

Models List

GET /v1/models

Returns the catalog of available models in OpenAI-shaped format, filtered by your account's enabled providers and routing strategy.

Documentation Endpoints

Path Format Purpose
/v1/docs HTML UI Interactive API documentation
/v1/openapi.json JSON Raw OpenAPI specification

These endpoints are enumerated in the repository's README.md (lines 51-53) and fully documented in docs/api.md (lines 9-23).

Extended Compatibility Surfaces

Beyond strict OpenAI compatibility, FreeLLMAPI exposes additional interfaces for broader ecosystem integration.

Anthropic Messages API

Full Claude client compatibility:

POST /v1/messages

Accepts Anthropic's native Messages API format and normalizes responses. Use with official Anthropic SDKs by overriding the base URL:

export ANTHROPIC_BASE_URL=http://localhost:3001
export ANTHROPIC_AUTH_TOKEN=freellmapi-your-unified-key

Native Gemini Interface

Direct Google Gemini API access:

/v1beta/models
/v1beta/models/{model}:generateContent

Preserves Gemini's native request/response formats for applications requiring Gemini-specific features.

Ollama Emulation

Local-model client compatibility via NDJSON surface:

/api/chat
/api/generate

Enables tools like Ollama-compatible UIs and CLI clients to use FreeLLMAPI's pooled providers.

Revocable URL Tokens

Header-less authentication for constrained environments:


/v1/t/{token}/[endpoint]

Useful for browser extensions, webhook integrations, or agents that cannot manipulate HTTP headers directly.

Architectural Implementation

Request Flow Through the System

  1. Entry point: server/src/index.ts — Express server registering all /v1/* routes
  2. Authentication: Unified bearer token validation
  3. Routing: Strategy-based provider selection (auto, auto:fast, auto:smart)
  4. Translation: server/src/providers/openai-compat.ts transforms OpenAI schema to provider-native format
  5. Failover: Automatic retry with X-Fallback-Attempts and X-Routed-Via headers
  6. Normalization: Response unified to OpenAI format regardless of source provider

Key Source Files

File Role
server/src/providers/openai-compat.ts Adapter implementing OpenAI-compatible endpoints
server/src/index.ts Main server; route registration and authentication
docs/api.md Human-readable API reference
README.md Feature overview with endpoint list
docs/architecture.md Routing, fallback, and provider integration details

Summary

  • FreeLLMAPI endpoints for OpenAI compatibility cover the complete OpenAI surface: chat completions, legacy completions, responses, embeddings, images, video, audio, and model discovery
  • All endpoints live under /v1 with identical request/response shapes to OpenAI's API
  • Extended surfaces include Anthropic Messages (/v1/messages), native Gemini (/v1beta/*), and Ollama emulation (/api/*)
  • Authentication uses a unified bearer token with optional revocable URL tokens for restricted environments
  • The router in server/src/index.ts and adapter in server/src/providers/openai-compat.ts handle translation, failover, and response normalization automatically

Frequently Asked Questions

Does FreeLLMAPI support streaming responses?

Yes. All streaming-capable OpenAI endpoints including /v1/chat/completions support Server-Sent Events (SSE) streaming. The proxy maintains the stream across provider failovers when possible, or falls back to non-streaming responses if necessary.

How does the auto routing strategy select providers?

The auto strategy evaluates current rate limits, latency, and model capability against your request requirements. Variants like auto:fast prioritize response time while auto:smart optimizes for output quality. The X-Routed-Via response header reveals which provider handled your request.

Can I use my existing OpenAI SDK code without modification?

Yes. Change only the base_url to http://localhost:3001/v1 (or your deployed instance) and replace the api_key with your FreeLLMAPI unified key. All method signatures, parameter names, and response parsing remain identical to OpenAI's official SDK.

What happens when all providers for a request type are rate-limited?

The router returns a structured error response matching OpenAI's error format with HTTP 429 status, including Retry-After guidance when available from provider responses. The X-Fallback-Attempts header indicates how many providers were exhausted.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →