How to Use the FreeLLMAPI Chat Completions Endpoint: A Complete API Guide
The FreeLLMAPI chat completions endpoint at /v1/chat/completions accepts standard OpenAI-compatible requests, automatically routes them to available providers through an intelligent proxy system, and supports both streaming and non-streaming responses with built-in quota management.
The FreeLLMAPI service provides a drop-in replacement for OpenAI's chat API, enabling developers to access multiple large language model providers through a single unified endpoint. According to the tashfeenahmed/freellmapi source code, the system implements intelligent request routing, automatic model selection, and transparent fallback mechanisms while maintaining full compatibility with the OpenAI SDK format. This guide covers the complete request lifecycle and implementation details derived directly from the repository's architecture.
Endpoint Architecture and Request Flow
The chat completions endpoint follows a six-stage processing pipeline implemented in server/src/routes/proxy.ts. When your request hits the /v1/chat/completions route, the system executes the following workflow:
-
Request Validation – The incoming JSON body is validated against a strict Zod schema that mirrors the OpenAI chat-completion payload structure【proxy.ts†L1381-L1394】. This ensures all required fields (
messages,model) and parameters (temperature,max_tokens) conform to expected types before processing begins. -
Automatic Model Selection – If you omit the
modelparameter, the system enters "auto" mode and selects the best-available chat model from the healthy provider pool【proxy.ts†L656-L658】. This logic queries the model discovery service to identify currently operational endpoints. -
Quota Verification – Before forwarding, the quota subsystem checks your API key's usage limits and attaches a quota context to the request for real-time accounting【proxy.ts†L1269-L1274】. Exceeded quotas return immediate 429 responses without wasting provider resources.
-
Provider Dispatch – Validated requests are handed to the fusion service (
server/src/services/fusion.ts), which routes to the selected provider'schatCompletionmethod (OpenAI, Anthropic, Google Gemini, etc.)【fusion.ts†L241-L242】. -
Response Streaming – Depending on your
streamparameter, the system either aggregates a singlechat.completionobject or returns server-sent events containingchat.completion.chunkobjects【proxy.ts†L1381-L1403】【proxy.ts†L1640-L1660】. -
Error Normalization – Upstream provider errors are captured and reshaped into standard OpenAI error formats, ensuring your client code handles failures consistently regardless of the backing infrastructure.
Authentication and Request Format
All requests to the FreeLLMAPI chat completions endpoint require bearer token authentication. The API accepts standard OpenAI chat completion parameters, making migration from existing OpenAI implementations straightforward.
Required Headers
Content-Type: application/json
Authorization: Bearer YOUR_API_KEY
Request Body Structure
The endpoint accepts a JSON payload with the following key fields:
- model – String identifier (e.g.,
"gpt-3.5-turbo","gpt-4", or"auto"for automatic selection) - messages – Array of message objects with
role(system,user,assistant) andcontentfields - stream – Boolean indicating whether to stream partial progress (default:
false) - max_tokens – Integer limiting response length
- temperature – Float between 0 and 2 controlling randomness
Making Your First Request
You can interact with the FreeLLMAPI chat completions endpoint using any HTTP client. The following examples demonstrate basic usage patterns.
Using cURL
curl https://api.freellmapi.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-3.5-turbo",
"messages": [{"role":"user","content":"Hello, who are you?"}],
"max_tokens": 50,
"temperature": 0.7
}'
Using Node.js (fetch)
import fetch from 'node-fetch';
const resp = await fetch('https://api.freellmapi.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
},
body: JSON.stringify({
model: 'gpt-3.5-turbo',
messages: [{role: 'user', content: 'Explain quantum tunnelling in simple terms.'}],
max_tokens: 150,
}),
});
const data = await resp.json();
console.log(data.choices[0].message.content);
Handling Streaming Responses
When you set stream: true in your request, the FreeLLMAPI chat completions endpoint returns a text/event-stream containing incremental JSON objects rather than a single complete response. This reduces time-to-first-token and enables real-time display in user interfaces.
Streaming Implementation (Node.js 18+)
const response = await fetch('https://api.freellmapi.com/v1/chat/completions', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: `Bearer ${process.env.FREELLMAPI_KEY}`,
},
body: JSON.stringify({
model: 'gpt-4',
messages: [{role: 'user', content: 'Write a haiku about sunrise.'}],
stream: true,
}),
});
for await (const chunk of response.body) {
// Each chunk is a JSON-encoded chat.completion.chunk
const parsed = JSON.parse(chunk.toString());
const content = parsed.choices[0]?.delta?.content;
if (content) process.stdout.write(content);
}
Each chunk follows the OpenAI delta format, containing id, object: "chat.completion.chunk", created timestamp, model name, and a choices array with delta objects that accumulate to form the final message.
Automatic Model Selection and Provider Fallbacks
One of the distinguishing features of the FreeLLMAPI chat completions endpoint is its transparent provider abstraction. When you specify "model": "auto" or omit the model parameter entirely, the system queries server/src/services/model-discovery.ts to identify healthy providers and selects the optimal model based on availability, latency, and cost constraints.
If your selected provider returns an error or times out, the fusion service (server/src/services/fusion.ts) can transparently retry with alternative backends, normalizing all responses into the standard OpenAPI schema regardless of the underlying provider's native format.
Quota Management and Rate Limiting
Every request passing through the FreeLLMAPI chat completions endpoint is subject to quota enforcement defined in server/src/services/ratelimit.ts. The system tracks both request counts and token usage (input + output) against your API key's limits.
The quota context attached during request processing【proxy.ts†L1269-L1274】 enables granular usage analytics and prevents bill shock by rejecting requests that would exceed your configured thresholds. Token usage statistics appear in the usage field of non-streaming responses and are aggregated server-side for streaming requests upon completion.
Key Implementation Files
The following source files define the core functionality of the chat completions endpoint:
server/src/routes/proxy.ts– Implements the HTTP route handler, request validation, model auto-selection, and streaming response logic【proxy.ts†L1381 onwards】.shared/types.ts– Contains TypeScript interfaces forChatCompletionandChatCompletionChunkobjects, ensuring type safety across the request lifecycle.server/src/services/fusion.ts– Handles provider routing and method dispatch to underlying LLM services【fusion.ts†L241-L242】.server/src/services/ratelimit.ts– Enforces per-key rate limits and tracks token consumption for usage reporting and quota enforcement.server/src/services/model-discovery.ts– Maintains the health-checked pool of available models and powers the automatic selection algorithm.
Summary
- The FreeLLMAPI chat completions endpoint at
/v1/chat/completionsprovides OpenAI-compatible request and response formats, enabling direct SDK substitution. - Automatic model selection occurs when you omit the model parameter or specify
"auto", routing to the healthiest available provider based on real-time status checks. - Streaming responses use server-sent events with
chat.completion.chunkobjects, while non-streaming requests return completechat.completionJSON. - Quota enforcement happens before provider dispatch, attaching usage context that tracks token consumption against your API key limits.
- Provider abstraction in
server/src/services/fusion.tsnormalizes errors and responses across multiple backends (OpenAI, Anthropic, Google) into a consistent interface.
Frequently Asked Questions
Is the FreeLLMAPI chat completions endpoint compatible with the OpenAI Python SDK?
Yes. Because the endpoint mirrors OpenAI's API specification exactly—including request schemas, authentication headers, and response formats—you can point the official OpenAI Python client to https://api.freellmapi.com by changing the base_url parameter. All methods including chat.completions.create() and streaming iterators function identically.
How does automatic model selection work when I don't specify a model?
When you omit the model field or set it to "auto", the system queries server/src/services/model-discovery.ts to identify currently healthy providers, then selects the optimal model based on availability, latency metrics, and cost parameters【proxy.ts†L656-L658】. This ensures requests succeed even if specific providers are experiencing outages.
What happens if a provider returns an error or rate limit?
The error handling middleware in server/src/routes/proxy.ts captures upstream provider errors and normalizes them into standard OpenAI error objects with consistent status codes and message formats. For transient failures, the fusion service may attempt fallback to alternative providers before returning an error to your client, though this behavior depends on your specific configuration and quota settings.
How do I monitor my quota usage for chat completion requests?
Non-streaming responses include a usage object containing prompt_tokens, completion_tokens, and total_tokens fields. For streaming requests, token counts are aggregated server-side and logged to your quota context【proxy.ts†L1269-L1274】. You can query your current usage statistics through the separate quota endpoints defined in server/src/services/ratelimit.ts, which track consumption against your API key's configured limits in real-time.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →