# FreeLLMAPI API Surface: Complete OpenAI-Compatible Endpoint Reference

> Explore the FreeLLMAPI API surface with a complete OpenAI-compatible endpoint reference. Discover twenty distinct endpoints for inference, embeddings, audio, and more.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: api-reference
- Published: 2026-09-01

---

**FreeLLMAPI implements a unified OpenAI-compatible REST API under the `/v1/` namespace, exposing twenty distinct endpoints for inference, embeddings, audio processing, model discovery, and administrative control through an Express-based meta-router.**

FreeLLMAPI is a self-hosted meta-router that aggregates multiple LLM providers behind a single OpenAI-compatible HTTP interface. All public API surfaces reside under the `/v1/` path and are implemented as Express routes that forward requests to an internal routing engine. This architecture allows clients to use standard OpenAI client libraries while the server automatically selects optimal providers based on quota, latency, and model availability.

## Core Inference Endpoints

The primary function of FreeLLMAPI is to route generation requests to the best available upstream provider. These endpoints mirror the OpenAI API specification and are handled by the central routing logic in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts).

### Chat Completions and Legacy Completions

The `/v1/chat/completions` endpoint accepts standard OpenAI chat format requests and supports both streaming and non-streaming responses. According to the source code in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts), this endpoint inspects the incoming payload, applies rate limiting and quota checks, then invokes the `routeRequest` function to select the optimal provider. The legacy `/v1/completions` endpoint for text completion requests is implemented in the same file and follows identical routing logic.

```javascript
await fetch('https://my-freellmapi.local/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
  },
  body: JSON.stringify({
    model: 'auto',
    messages: [{ role: 'user', content: 'Explain quantum tunneling.' }],
    temperature: 0.7,
  }),
});

```

### Embeddings

Vector embedding requests are handled by `/v1/embeddings` as implemented in [`server/src/routes/embeddings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/embeddings.ts). The router selects embedding-capable providers such as OpenRouter, NVIDIA, or Cloudflare based on the requested model and current provider health. The endpoint accepts an array of input strings and returns the corresponding vector representations.

```javascript
await fetch('https://my-freellmapi.local/v1/embeddings', {
  method: 'POST',
  headers: { 
    'Content-Type': 'application/json', 
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}` 
  },
  body: JSON.stringify({ 
    model: 'auto', 
    input: ['hello world'] 
  }),
});

```

### Audio and Speech Processing

The media endpoints in [`server/src/routes/media.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/media.ts) handle multimodal requests. The `/v1/audio/transcriptions` endpoint performs speech-to-text by forwarding audio files to providers supporting transcription, such as OpenAI, Groq, or NVIDIA. Conversely, `/v1/audio/speech` handles text-to-synthesis requests, routing to TTS-capable providers and returning generated audio streams.

```javascript
await fetch('https://my-freellmapi.local/v1/audio/transcriptions', {
  method: 'POST',
  headers: { Authorization: `Bearer ${process.env.FREELLMAPI_KEY}` },
  body: new FormData().append('file', fs.createReadStream('voice.wav')),
});

```

### Image Generation and Vision

Image generation requests sent to `/v1/images/generations` are routed to providers like Stability AI, OpenRouter, or NVIDIA based on availability. The internal `/v1/vision` endpoint, also handled within [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts), processes image-augmented chat requests when payloads contain visual inputs, forwarding them to vision-capable models.

## Provider Compatibility Layer

FreeLLMAPI extends beyond OpenAI compatibility to support alternative API formats, enabling drop-in replacement for diverse client ecosystems.

### Anthropic Messages API

The `/v1/messages` endpoint, implemented in [`server/src/routes/anthropic.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/anthropic.ts), provides Anthropic-compatible message formatting. The router maps incoming Anthropic-style message structures onto the internal OpenAI chat format before applying the standard routing logic. This allows Claude-specific clients to function without modification while leveraging FreeLLMAPI's provider aggregation.

### Legacy Responses Endpoint

The `/v1/responses` endpoint serves as a thin shim for historical OpenAI response patterns. As implemented in [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts), this route translates requests into standard chat completion calls and returns simplified JSON payloads for backward compatibility with older integrations.

## Discovery and Model Management

Understanding available capabilities is essential for client applications dynamically selecting models.

### Listing Available Models

The `/v1/models` GET endpoint, defined in [`server/src/routes/models.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/models.ts), returns a comprehensive catalog of all routable models, including "auto" entries and provider-specific identifiers. Clients use this endpoint to discover which model IDs are valid for the current configuration without hardcoding provider-specific strings.

```javascript
const res = await fetch('https://my-freellmapi.local/v1/models', {
  headers: { Authorization: `Bearer ${process.env.FREELLMAPI_KEY}` },
});
const data = await res.json();
console.log(data);

```

### Provider Health and Status

System visibility is provided through `/v1/providers` and `/v1/status`, both implemented in [`server/src/routes/status.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/status.ts). The providers endpoint returns quota, limits, and error states for each upstream connection, while the status endpoint reports general service health including database connectivity and process uptime. These endpoints power the administrative dashboard and enable programmatic health monitoring.

```javascript
const health = await fetch('https://my-freellmapi.local/v1/providers');
console.log(await health.json());

```

## Administrative and System Endpoints

FreeLLMAPI exposes management interfaces for configuration, authentication, and licensing.

### API Key Management

The `/v1/keys` endpoint supports full CRUD operations for API keys and client profiles, as implemented in [`server/src/routes/keys.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/keys.ts). Keys can be scoped to specific profiles and optionally carry system prompts, allowing fine-grained access control and request pre-configuration.

### Configuration and Settings

Runtime router behavior is adjustable through `/v1/settings`, defined in [`server/src/routes/settings.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/settings.ts). This endpoint accepts GET and PATCH methods to read or modify configuration parameters including headroom thresholds, custom provider weights, and routing strategies without requiring server restarts.

### Health Checks and Monitoring

Load balancers and orchestration systems interact with `/v1/health` from [`server/src/routes/health.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/health.ts), which returns lightweight ready/live status indicators. For detailed operational metrics, `/v1/analytics` in [`server/src/routes/analytics.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/analytics.ts) exposes token consumption statistics per provider, model, and API key.

## Advanced Routing and Caching

The internal architecture exposes specialized endpoints for optimization and debugging.

### Request Routing Architecture

All inference endpoints ultimately invoke the `routeRequest` function in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts). This central engine evaluates each request against quota limits, rate limiting rules, and sticky-session constraints, then scores every enabled model-provider combination to select the optimal candidate. When primary providers exhaust their quotas, the system executes configurable fallback strategies.

### Caching and Analytics

The `/v1/cache` endpoint in [`server/src/routes/cache.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/cache.ts) allows inspection and priming of the deterministic request cache, enabling clients to pre-warm responses for specific temperature-controlled queries. The `/v1/fallback` endpoint from [`server/src/routes/fallback.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/fallback.ts) exposes current fallback-loop configurations, revealing the hierarchy of provider failover chains. The `/v1/license` endpoint in [`server/src/routes/premium.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/premium.ts) enforces free-tier quotas and handles activation flows.

## Summary

FreeLLMAPI exposes a comprehensive OpenAI-compatible API surface through twenty distinct endpoints under the `/v1/` namespace. Key capabilities include:

- **Unified inference routing** for chat, completion, embedding, audio, and image generation requests through `/v1/chat/completions`, `/v1/embeddings`, `/v1/audio/*`, and `/v1/images/generations`
- **Provider abstraction** via the `routeRequest` function in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts), which selects optimal upstream providers based on real-time quota and latency heuristics
- **Multi-format compatibility** supporting both OpenAI and Anthropic API conventions through dedicated route handlers
- **Operational visibility** through `/v1/models`, `/v1/providers`, and `/v1/analytics` endpoints that expose system status and usage statistics
- **Administrative control** via `/v1/keys`, `/v1/settings`, and `/v1/cache` for runtime configuration and access management

## Frequently Asked Questions

### Is FreeLLMAPI fully compatible with OpenAI's REST API?

Yes, FreeLLMAPI implements the complete OpenAI API specification under the `/v1/` namespace, including chat completions, embeddings, audio transcription, text-to-speech, and image generation. Clients using standard OpenAI SDKs can point their base URL to a FreeLLMAPI instance without code changes, as the endpoints accept identical request payloads and return compatible response formats.

### How does FreeLLMAPI decide which provider handles a request?

The routing decision occurs in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts) within the `routeRequest` function. This engine evaluates enabled providers against current quota availability, rate limits, and configured weights, then scores each model-provider pair to identify the best candidate. If the primary provider fails or exhausts its limits, the system automatically executes fallback logic defined in [`server/src/routes/fallback.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/fallback.ts).

### Can I restrict which models or providers specific API keys can access?

Yes, the `/v1/keys` endpoint implemented in [`server/src/routes/keys.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/keys.ts) supports creating scoped API keys with associated profiles. These profiles can carry system prompts and provider restrictions, allowing administrators to limit specific clients to particular model families or upstream providers while maintaining a single unified API endpoint.

### Does FreeLLMAPI support vision models and multimodal inputs?

Yes, the `/v1/vision` endpoint handled within [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) processes image-augmented chat requests, routing them to vision-capable providers. Additionally, `/v1/images/generations` handles image synthesis requests. The router automatically detects multimodal content in payloads and forwards them to appropriate upstream services that support the requested modalities.