How to Use the FreeLLMAPI Chat Completions Endpoint: A Complete API Guide

The FreeLLMAPI chat completions endpoint at /v1/chat/completions accepts standard OpenAI-compatible requests, automatically routes them to available providers through an intelligent proxy system, and supports both streaming and non-streaming responses with built-in quota management.

The FreeLLMAPI service provides a drop-in replacement for OpenAI's chat API, enabling developers to access multiple large language model providers through a single unified endpoint. According to the tashfeenahmed/freellmapi source code, the system implements intelligent request routing, automatic model selection, and transparent fallback mechanisms while maintaining full compatibility with the OpenAI SDK format. This guide covers the complete request lifecycle and implementation details derived directly from the repository's architecture.

Endpoint Architecture and Request Flow

The chat completions endpoint follows a six-stage processing pipeline implemented in server/src/routes/proxy.ts. When your request hits the /v1/chat/completions route, the system executes the following workflow:

  • Request Validation – The incoming JSON body is validated against a strict Zod schema that mirrors the OpenAI chat-completion payload structure【proxy.ts†L1381-L1394】. This ensures all required fields (messages, model) and parameters (temperature, max_tokens) conform to expected types before processing begins.

  • Automatic Model Selection – If you omit the model parameter, the system enters "auto" mode and selects the best-available chat model from the healthy provider pool【proxy.ts†L656-L658】. This logic queries the model discovery service to identify currently operational endpoints.

  • Quota Verification – Before forwarding, the quota subsystem checks your API key's usage limits and attaches a quota context to the request for real-time accounting【proxy.ts†L1269-L1274】. Exceeded quotas return immediate 429 responses without wasting provider resources.

  • Provider Dispatch – Validated requests are handed to the fusion service (server/src/services/fusion.ts), which routes to the selected provider's chatCompletion method (OpenAI, Anthropic, Google Gemini, etc.)【fusion.ts†L241-L242】.

  • Response Streaming – Depending on your stream parameter, the system either aggregates a single chat.completion object or returns server-sent events containing chat.completion.chunk objects【proxy.ts†L1381-L1403】【proxy.ts†L1640-L1660】.

  • Error Normalization – Upstream provider errors are captured and reshaped into standard OpenAI error formats, ensuring your client code handles failures consistently regardless of the backing infrastructure.

Authentication and Request Format

All requests to the FreeLLMAPI chat completions endpoint require bearer token authentication. The API accepts standard OpenAI chat completion parameters, making migration from existing OpenAI implementations straightforward.

Required Headers

Content-Type: application/json
Authorization: Bearer YOUR_API_KEY

Request Body Structure

The endpoint accepts a JSON payload with the following key fields:

  • model – String identifier (e.g., "gpt-3.5-turbo", "gpt-4", or "auto" for automatic selection)
  • messages – Array of message objects with role (system, user, assistant) and content fields
  • stream – Boolean indicating whether to stream partial progress (default: false)
  • max_tokens – Integer limiting response length
  • temperature – Float between 0 and 2 controlling randomness

Making Your First Request

You can interact with the FreeLLMAPI chat completions endpoint using any HTTP client. The following examples demonstrate basic usage patterns.

Using cURL

curl https://api.freellmapi.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
        "model": "gpt-3.5-turbo",
        "messages": [{"role":"user","content":"Hello, who are you?"}],
        "max_tokens": 50,
        "temperature": 0.7
      }'

Using Node.js (fetch)

import fetch from 'node-fetch';

const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
  },
  body: JSON.stringify({
    model: 'gpt-3.5-turbo',
    messages: [{role: 'user', content: 'Explain quantum tunnelling in simple terms.'}],
    max_tokens: 150,
  }),
});

const data = await resp.json();
console.log(data.choices[0].message.content);

Handling Streaming Responses

When you set stream: true in your request, the FreeLLMAPI chat completions endpoint returns a text/event-stream containing incremental JSON objects rather than a single complete response. This reduces time-to-first-token and enables real-time display in user interfaces.

Streaming Implementation (Node.js 18+)

const response = await fetch('https://api.freellmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
  },
  body: JSON.stringify({
    model: 'gpt-4',
    messages: [{role: 'user', content: 'Write a haiku about sunrise.'}],
    stream: true,
  }),
});

for await (const chunk of response.body) {
  // Each chunk is a JSON-encoded chat.completion.chunk
  const parsed = JSON.parse(chunk.toString());
  const content = parsed.choices[0]?.delta?.content;
  if (content) process.stdout.write(content);
}

Each chunk follows the OpenAI delta format, containing id, object: "chat.completion.chunk", created timestamp, model name, and a choices array with delta objects that accumulate to form the final message.

Automatic Model Selection and Provider Fallbacks

One of the distinguishing features of the FreeLLMAPI chat completions endpoint is its transparent provider abstraction. When you specify "model": "auto" or omit the model parameter entirely, the system queries server/src/services/model-discovery.ts to identify healthy providers and selects the optimal model based on availability, latency, and cost constraints.

If your selected provider returns an error or times out, the fusion service (server/src/services/fusion.ts) can transparently retry with alternative backends, normalizing all responses into the standard OpenAPI schema regardless of the underlying provider's native format.

Quota Management and Rate Limiting

Every request passing through the FreeLLMAPI chat completions endpoint is subject to quota enforcement defined in server/src/services/ratelimit.ts. The system tracks both request counts and token usage (input + output) against your API key's limits.

The quota context attached during request processing【proxy.ts†L1269-L1274】 enables granular usage analytics and prevents bill shock by rejecting requests that would exceed your configured thresholds. Token usage statistics appear in the usage field of non-streaming responses and are aggregated server-side for streaming requests upon completion.

Key Implementation Files

The following source files define the core functionality of the chat completions endpoint:

  • server/src/routes/proxy.ts – Implements the HTTP route handler, request validation, model auto-selection, and streaming response logic【proxy.ts†L1381 onwards】.
  • shared/types.ts – Contains TypeScript interfaces for ChatCompletion and ChatCompletionChunk objects, ensuring type safety across the request lifecycle.
  • server/src/services/fusion.ts – Handles provider routing and method dispatch to underlying LLM services【fusion.ts†L241-L242】.
  • server/src/services/ratelimit.ts – Enforces per-key rate limits and tracks token consumption for usage reporting and quota enforcement.
  • server/src/services/model-discovery.ts – Maintains the health-checked pool of available models and powers the automatic selection algorithm.

Summary

  • The FreeLLMAPI chat completions endpoint at /v1/chat/completions provides OpenAI-compatible request and response formats, enabling direct SDK substitution.
  • Automatic model selection occurs when you omit the model parameter or specify "auto", routing to the healthiest available provider based on real-time status checks.
  • Streaming responses use server-sent events with chat.completion.chunk objects, while non-streaming requests return complete chat.completion JSON.
  • Quota enforcement happens before provider dispatch, attaching usage context that tracks token consumption against your API key limits.
  • Provider abstraction in server/src/services/fusion.ts normalizes errors and responses across multiple backends (OpenAI, Anthropic, Google) into a consistent interface.

Frequently Asked Questions

Is the FreeLLMAPI chat completions endpoint compatible with the OpenAI Python SDK?

Yes. Because the endpoint mirrors OpenAI's API specification exactly—including request schemas, authentication headers, and response formats—you can point the official OpenAI Python client to https://api.freellmapi.com by changing the base_url parameter. All methods including chat.completions.create() and streaming iterators function identically.

How does automatic model selection work when I don't specify a model?

When you omit the model field or set it to "auto", the system queries server/src/services/model-discovery.ts to identify currently healthy providers, then selects the optimal model based on availability, latency metrics, and cost parameters【proxy.ts†L656-L658】. This ensures requests succeed even if specific providers are experiencing outages.

What happens if a provider returns an error or rate limit?

The error handling middleware in server/src/routes/proxy.ts captures upstream provider errors and normalizes them into standard OpenAI error objects with consistent status codes and message formats. For transient failures, the fusion service may attempt fallback to alternative providers before returning an error to your client, though this behavior depends on your specific configuration and quota settings.

How do I monitor my quota usage for chat completion requests?

Non-streaming responses include a usage object containing prompt_tokens, completion_tokens, and total_tokens fields. For streaming requests, token counts are aggregated server-side and logged to your quota context【proxy.ts†L1269-L1274】. You can query your current usage statistics through the separate quota endpoints defined in server/src/services/ratelimit.ts, which track consumption against your API key's configured limits in real-time.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →