# How to Use the FreeLLMAPI Chat Completions Endpoint: A Complete API Guide

> Learn how to use the FreeLLMAPI chat completions endpoint. This guide covers OpenAI-compatible requests, intelligent proxy routing, streaming, and quota management.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: api-guide
- Published: 2026-08-29

---

**The FreeLLMAPI chat completions endpoint at `/v1/chat/completions` accepts standard OpenAI-compatible requests, automatically routes them to available providers through an intelligent proxy system, and supports both streaming and non-streaming responses with built-in quota management.**

The FreeLLMAPI service provides a drop-in replacement for OpenAI's chat API, enabling developers to access multiple large language model providers through a single unified endpoint. According to the `tashfeenahmed/freellmapi` source code, the system implements intelligent request routing, automatic model selection, and transparent fallback mechanisms while maintaining full compatibility with the OpenAI SDK format. This guide covers the complete request lifecycle and implementation details derived directly from the repository's architecture.

## Endpoint Architecture and Request Flow

The chat completions endpoint follows a six-stage processing pipeline implemented in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts). When your request hits the `/v1/chat/completions` route, the system executes the following workflow:

- **Request Validation** – The incoming JSON body is validated against a strict Zod schema that mirrors the OpenAI chat-completion payload structure【proxy.ts†L1381-L1394】. This ensures all required fields (`messages`, `model`) and parameters (`temperature`, `max_tokens`) conform to expected types before processing begins.

- **Automatic Model Selection** – If you omit the `model` parameter, the system enters "auto" mode and selects the best-available chat model from the healthy provider pool【proxy.ts†L656-L658】. This logic queries the model discovery service to identify currently operational endpoints.

- **Quota Verification** – Before forwarding, the quota subsystem checks your API key's usage limits and attaches a **quota context** to the request for real-time accounting【proxy.ts†L1269-L1274】. Exceeded quotas return immediate 429 responses without wasting provider resources.

- **Provider Dispatch** – Validated requests are handed to the **fusion service** ([`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts)), which routes to the selected provider's `chatCompletion` method (OpenAI, Anthropic, Google Gemini, etc.)【fusion.ts†L241-L242】.

- **Response Streaming** – Depending on your `stream` parameter, the system either aggregates a single `chat.completion` object or returns server-sent events containing `chat.completion.chunk` objects【proxy.ts†L1381-L1403】【proxy.ts†L1640-L1660】.

- **Error Normalization** – Upstream provider errors are captured and reshaped into standard OpenAI error formats, ensuring your client code handles failures consistently regardless of the backing infrastructure.

## Authentication and Request Format

All requests to the FreeLLMAPI chat completions endpoint require bearer token authentication. The API accepts standard OpenAI chat completion parameters, making migration from existing OpenAI implementations straightforward.

### Required Headers

```bash
Content-Type: application/json
Authorization: Bearer YOUR_API_KEY

```

### Request Body Structure

The endpoint accepts a JSON payload with the following key fields:

- **model** – String identifier (e.g., `"gpt-3.5-turbo"`, `"gpt-4"`, or `"auto"` for automatic selection)
- **messages** – Array of message objects with `role` (`system`, `user`, `assistant`) and `content` fields
- **stream** – Boolean indicating whether to stream partial progress (default: `false`)
- **max_tokens** – Integer limiting response length
- **temperature** – Float between 0 and 2 controlling randomness

## Making Your First Request

You can interact with the FreeLLMAPI chat completions endpoint using any HTTP client. The following examples demonstrate basic usage patterns.

### Using cURL

```bash
curl https://api.freellmapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
        "model": "gpt-3.5-turbo",
        "messages": [{"role":"user","content":"Hello, who are you?"}],
        "max_tokens": 50,
        "temperature": 0.7
      }'

```

### Using Node.js (fetch)

```javascript
import fetch from 'node-fetch';

const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
  },
  body: JSON.stringify({
    model: 'gpt-3.5-turbo',
    messages: [{role: 'user', content: 'Explain quantum tunnelling in simple terms.'}],
    max_tokens: 150,
  }),
});

const data = await resp.json();
console.log(data.choices[0].message.content);

```

## Handling Streaming Responses

When you set `stream: true` in your request, the FreeLLMAPI chat completions endpoint returns a text/event-stream containing incremental JSON objects rather than a single complete response. This reduces time-to-first-token and enables real-time display in user interfaces.

### Streaming Implementation (Node.js 18+)

```javascript
const response = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
  },
  body: JSON.stringify({
    model: 'gpt-4',
    messages: [{role: 'user', content: 'Write a haiku about sunrise.'}],
    stream: true,
  }),
});

for await (const chunk of response.body) {
  // Each chunk is a JSON-encoded chat.completion.chunk
  const parsed = JSON.parse(chunk.toString());
  const content = parsed.choices[0]?.delta?.content;
  if (content) process.stdout.write(content);
}

```

Each chunk follows the OpenAI delta format, containing `id`, `object: "chat.completion.chunk"`, `created` timestamp, `model` name, and a `choices` array with `delta` objects that accumulate to form the final message.

## Automatic Model Selection and Provider Fallbacks

One of the distinguishing features of the FreeLLMAPI chat completions endpoint is its transparent **provider abstraction**. When you specify `"model": "auto"` or omit the model parameter entirely, the system queries [`server/src/services/model-discovery.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-discovery.ts) to identify healthy providers and selects the optimal model based on availability, latency, and cost constraints.

If your selected provider returns an error or times out, the fusion service ([`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts)) can transparently retry with alternative backends, normalizing all responses into the standard OpenAPI schema regardless of the underlying provider's native format.

## Quota Management and Rate Limiting

Every request passing through the FreeLLMAPI chat completions endpoint is subject to quota enforcement defined in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts). The system tracks both request counts and token usage (input + output) against your API key's limits.

The quota context attached during request processing【proxy.ts†L1269-L1274】 enables granular usage analytics and prevents bill shock by rejecting requests that would exceed your configured thresholds. Token usage statistics appear in the `usage` field of non-streaming responses and are aggregated server-side for streaming requests upon completion.

## Key Implementation Files

The following source files define the core functionality of the chat completions endpoint:

- **[`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts)** – Implements the HTTP route handler, request validation, model auto-selection, and streaming response logic【proxy.ts†L1381 onwards】.
- **[`shared/types.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/shared/types.ts)** – Contains TypeScript interfaces for `ChatCompletion` and `ChatCompletionChunk` objects, ensuring type safety across the request lifecycle.
- **[`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts)** – Handles provider routing and method dispatch to underlying LLM services【fusion.ts†L241-L242】.
- **[`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts)** – Enforces per-key rate limits and tracks token consumption for usage reporting and quota enforcement.
- **[`server/src/services/model-discovery.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-discovery.ts)** – Maintains the health-checked pool of available models and powers the automatic selection algorithm.

## Summary

- The **FreeLLMAPI chat completions endpoint** at `/v1/chat/completions` provides OpenAI-compatible request and response formats, enabling direct SDK substitution.
- **Automatic model selection** occurs when you omit the model parameter or specify `"auto"`, routing to the healthiest available provider based on real-time status checks.
- **Streaming responses** use server-sent events with `chat.completion.chunk` objects, while non-streaming requests return complete `chat.completion` JSON.
- **Quota enforcement** happens before provider dispatch, attaching usage context that tracks token consumption against your API key limits.
- **Provider abstraction** in [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts) normalizes errors and responses across multiple backends (OpenAI, Anthropic, Google) into a consistent interface.

## Frequently Asked Questions

### Is the FreeLLMAPI chat completions endpoint compatible with the OpenAI Python SDK?

Yes. Because the endpoint mirrors OpenAI's API specification exactly—including request schemas, authentication headers, and response formats—you can point the official OpenAI Python client to `https://api.freellmapi.com` by changing the `base_url` parameter. All methods including `chat.completions.create()` and streaming iterators function identically.

### How does automatic model selection work when I don't specify a model?

When you omit the `model` field or set it to `"auto"`, the system queries [`server/src/services/model-discovery.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/model-discovery.ts) to identify currently healthy providers, then selects the optimal model based on availability, latency metrics, and cost parameters【proxy.ts†L656-L658】. This ensures requests succeed even if specific providers are experiencing outages.

### What happens if a provider returns an error or rate limit?

The error handling middleware in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) captures upstream provider errors and normalizes them into standard OpenAI error objects with consistent status codes and message formats. For transient failures, the fusion service may attempt fallback to alternative providers before returning an error to your client, though this behavior depends on your specific configuration and quota settings.

### How do I monitor my quota usage for chat completion requests?

Non-streaming responses include a `usage` object containing `prompt_tokens`, `completion_tokens`, and `total_tokens` fields. For streaming requests, token counts are aggregated server-side and logged to your quota context【proxy.ts†L1269-L1274】. You can query your current usage statistics through the separate quota endpoints defined in [`server/src/services/ratelimit.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/ratelimit.ts), which track consumption against your API key's configured limits in real-time.