FreeLLMAPI API Surface: Complete OpenAI-Compatible Endpoint Reference

FreeLLMAPI implements a unified OpenAI-compatible REST API under the /v1/ namespace, exposing twenty distinct endpoints for inference, embeddings, audio processing, model discovery, and administrative control through an Express-based meta-router.

FreeLLMAPI is a self-hosted meta-router that aggregates multiple LLM providers behind a single OpenAI-compatible HTTP interface. All public API surfaces reside under the /v1/ path and are implemented as Express routes that forward requests to an internal routing engine. This architecture allows clients to use standard OpenAI client libraries while the server automatically selects optimal providers based on quota, latency, and model availability.

Core Inference Endpoints

The primary function of FreeLLMAPI is to route generation requests to the best available upstream provider. These endpoints mirror the OpenAI API specification and are handled by the central routing logic in server/src/services/router.ts.

Chat Completions and Legacy Completions

The /v1/chat/completions endpoint accepts standard OpenAI chat format requests and supports both streaming and non-streaming responses. According to the source code in server/src/routes/proxy.ts, this endpoint inspects the incoming payload, applies rate limiting and quota checks, then invokes the routeRequest function to select the optimal provider. The legacy /v1/completions endpoint for text completion requests is implemented in the same file and follows identical routing logic.

await fetch('https://my-freellmapi.local/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
  },
  body: JSON.stringify({
    model: 'auto',
    messages: [{ role: 'user', content: 'Explain quantum tunneling.' }],
    temperature: 0.7,
  }),
});

Embeddings

Vector embedding requests are handled by /v1/embeddings as implemented in server/src/routes/embeddings.ts. The router selects embedding-capable providers such as OpenRouter, NVIDIA, or Cloudflare based on the requested model and current provider health. The endpoint accepts an array of input strings and returns the corresponding vector representations.

await fetch('https://my-freellmapi.local/v1/embeddings', {
  method: 'POST',
  headers: { 
    'Content-Type': 'application/json', 
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}` 
  },
  body: JSON.stringify({ 
    model: 'auto', 
    input: ['hello world'] 
  }),
});

Audio and Speech Processing

The media endpoints in server/src/routes/media.ts handle multimodal requests. The /v1/audio/transcriptions endpoint performs speech-to-text by forwarding audio files to providers supporting transcription, such as OpenAI, Groq, or NVIDIA. Conversely, /v1/audio/speech handles text-to-synthesis requests, routing to TTS-capable providers and returning generated audio streams.

await fetch('https://my-freellmapi.local/v1/audio/transcriptions', {
  method: 'POST',
  headers: { Authorization: `Bearer ${process.env.FREELLMAPI_KEY}` },
  body: new FormData().append('file', fs.createReadStream('voice.wav')),
});

Image Generation and Vision

Image generation requests sent to /v1/images/generations are routed to providers like Stability AI, OpenRouter, or NVIDIA based on availability. The internal /v1/vision endpoint, also handled within server/src/routes/proxy.ts, processes image-augmented chat requests when payloads contain visual inputs, forwarding them to vision-capable models.

Provider Compatibility Layer

FreeLLMAPI extends beyond OpenAI compatibility to support alternative API formats, enabling drop-in replacement for diverse client ecosystems.

Anthropic Messages API

The /v1/messages endpoint, implemented in server/src/routes/anthropic.ts, provides Anthropic-compatible message formatting. The router maps incoming Anthropic-style message structures onto the internal OpenAI chat format before applying the standard routing logic. This allows Claude-specific clients to function without modification while leveraging FreeLLMAPI's provider aggregation.

Legacy Responses Endpoint

The /v1/responses endpoint serves as a thin shim for historical OpenAI response patterns. As implemented in server/src/routes/responses.ts, this route translates requests into standard chat completion calls and returns simplified JSON payloads for backward compatibility with older integrations.

Discovery and Model Management

Understanding available capabilities is essential for client applications dynamically selecting models.

Listing Available Models

The /v1/models GET endpoint, defined in server/src/routes/models.ts, returns a comprehensive catalog of all routable models, including "auto" entries and provider-specific identifiers. Clients use this endpoint to discover which model IDs are valid for the current configuration without hardcoding provider-specific strings.

const res = await fetch('https://my-freellmapi.local/v1/models', {
  headers: { Authorization: `Bearer ${process.env.FREELLMAPI_KEY}` },
});
const data = await res.json();
console.log(data);

Provider Health and Status

System visibility is provided through /v1/providers and /v1/status, both implemented in server/src/routes/status.ts. The providers endpoint returns quota, limits, and error states for each upstream connection, while the status endpoint reports general service health including database connectivity and process uptime. These endpoints power the administrative dashboard and enable programmatic health monitoring.

const health = await fetch('https://my-freellmapi.local/v1/providers');
console.log(await health.json());

Administrative and System Endpoints

FreeLLMAPI exposes management interfaces for configuration, authentication, and licensing.

API Key Management

The /v1/keys endpoint supports full CRUD operations for API keys and client profiles, as implemented in server/src/routes/keys.ts. Keys can be scoped to specific profiles and optionally carry system prompts, allowing fine-grained access control and request pre-configuration.

Configuration and Settings

Runtime router behavior is adjustable through /v1/settings, defined in server/src/routes/settings.ts. This endpoint accepts GET and PATCH methods to read or modify configuration parameters including headroom thresholds, custom provider weights, and routing strategies without requiring server restarts.

Health Checks and Monitoring

Load balancers and orchestration systems interact with /v1/health from server/src/routes/health.ts, which returns lightweight ready/live status indicators. For detailed operational metrics, /v1/analytics in server/src/routes/analytics.ts exposes token consumption statistics per provider, model, and API key.

Advanced Routing and Caching

The internal architecture exposes specialized endpoints for optimization and debugging.

Request Routing Architecture

All inference endpoints ultimately invoke the routeRequest function in server/src/services/router.ts. This central engine evaluates each request against quota limits, rate limiting rules, and sticky-session constraints, then scores every enabled model-provider combination to select the optimal candidate. When primary providers exhaust their quotas, the system executes configurable fallback strategies.

Caching and Analytics

The /v1/cache endpoint in server/src/routes/cache.ts allows inspection and priming of the deterministic request cache, enabling clients to pre-warm responses for specific temperature-controlled queries. The /v1/fallback endpoint from server/src/routes/fallback.ts exposes current fallback-loop configurations, revealing the hierarchy of provider failover chains. The /v1/license endpoint in server/src/routes/premium.ts enforces free-tier quotas and handles activation flows.

Summary

FreeLLMAPI exposes a comprehensive OpenAI-compatible API surface through twenty distinct endpoints under the /v1/ namespace. Key capabilities include:

  • Unified inference routing for chat, completion, embedding, audio, and image generation requests through /v1/chat/completions, /v1/embeddings, /v1/audio/*, and /v1/images/generations
  • Provider abstraction via the routeRequest function in server/src/services/router.ts, which selects optimal upstream providers based on real-time quota and latency heuristics
  • Multi-format compatibility supporting both OpenAI and Anthropic API conventions through dedicated route handlers
  • Operational visibility through /v1/models, /v1/providers, and /v1/analytics endpoints that expose system status and usage statistics
  • Administrative control via /v1/keys, /v1/settings, and /v1/cache for runtime configuration and access management

Frequently Asked Questions

Is FreeLLMAPI fully compatible with OpenAI's REST API?

Yes, FreeLLMAPI implements the complete OpenAI API specification under the /v1/ namespace, including chat completions, embeddings, audio transcription, text-to-speech, and image generation. Clients using standard OpenAI SDKs can point their base URL to a FreeLLMAPI instance without code changes, as the endpoints accept identical request payloads and return compatible response formats.

How does FreeLLMAPI decide which provider handles a request?

The routing decision occurs in server/src/services/router.ts within the routeRequest function. This engine evaluates enabled providers against current quota availability, rate limits, and configured weights, then scores each model-provider pair to identify the best candidate. If the primary provider fails or exhausts its limits, the system automatically executes fallback logic defined in server/src/routes/fallback.ts.

Can I restrict which models or providers specific API keys can access?

Yes, the /v1/keys endpoint implemented in server/src/routes/keys.ts supports creating scoped API keys with associated profiles. These profiles can carry system prompts and provider restrictions, allowing administrators to limit specific clients to particular model families or upstream providers while maintaining a single unified API endpoint.

Does FreeLLMAPI support vision models and multimodal inputs?

Yes, the /v1/vision endpoint handled within server/src/routes/proxy.ts processes image-augmented chat requests, routing them to vision-capable providers. Additionally, /v1/images/generations handles image synthesis requests. The router automatically detects multimodal content in payloads and forwards them to appropriate upstream services that support the requested modalities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →