How to Use Streaming with FreeLLMAPI: A Complete Developer Guide
FreeLLMAPI supports true token-by-token streaming across all supported providers—including Groq, Cerebras, Gemini, and OpenAI-compatible endpoints—by setting stream=True in your request and iterating over SSE chunks containing delta fields.
FreeLLMAPI is an open-source unified gateway that proxies requests to multiple large language model providers while preserving native streaming capabilities. According to the tashfeenahmed/freellmapi source code, the platform implements a provider-agnostic streaming architecture that forwards Server-Sent Events (SSE) directly from upstream services to your client without buffering. This guide explains how to activate streaming, consume token deltas, and handle provider-specific implementations using the repository's abstract provider pattern.
Architecture of FreeLLMAPI Streaming
The streaming implementation in FreeLLMAPI is designed to be provider-agnostic, allowing uniform consumption across disparate LLM backends.
The Provider Abstraction Layer
At the core of the system lies the abstract Provider class defined in server/src/providers/base.ts. This base class declares the streamChatCompletion method, which all concrete providers must implement to handle streaming requests. The method signature enforces a consistent interface for initiating token-by-token responses regardless of the upstream service.
Provider-Specific Implementations
Concrete provider implementations handle the protocol translation for their respective services:
-
server/src/providers/openai-compat.ts: Handles streaming for OpenAI-compatible providers including Groq, Cerebras, Mistral, OpenRouter, GitHub Models, HuggingFace, Cloudflare, and Cohere. This implementation forwards thestream=Trueparameter directly to the upstream API and yields chunks as they arrive. -
server/src/providers/google.ts: Implements Gemini-specific streaming logic using thestreamGenerateContentendpoint. Unlike standard OpenAI-compatible streams, Gemini requires the?alt=ssequery parameter to enable Server-Sent Events.
The HTTP router in server/src/routes/proxy.ts pipes these chunks straight back to the client, preserving the SSE semantics defined in the OpenAI specification.
Enabling Streaming in Your Requests
To activate streaming with FreeLLMAPI, you must explicitly set the streaming parameter in your request payload.
For OpenAI-compatible providers, include "stream": true in your JSON payload:
{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}
For Google Gemini, append ?alt=sse to your request URL or use the provider-specific streaming flag, as implemented in server/src/providers/google.ts.
Consuming Streamed Responses
When streaming is enabled, FreeLLMAPI returns a series of JSON objects rather than a single complete response. Each chunk follows the OpenAI delta format:
delta: Contains the next token string or tool-call fragment. Access this viachoices[0].delta.contentin client libraries.finish_reason: Signals stream termination. Values include"stop","length", or"tool_calls".- Tool-call deltas: Initial chunks may contain
delta.tool_callswith partial function arguments. The stream ends withfinish_reason: "tool_calls"once the function call is complete.
Implementation Examples for FreeLLMAPI Streaming
You can consume FreeLLMAPI streams using any HTTP client that supports Server-Sent Events.
Python with OpenAI Client
import openai
client = openai.OpenAI(
base_url="https://api.freellmapi.com/v1",
api_key="freellmapi-<your-token>"
)
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Explain streaming in FreeLLMAPI"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
Node.js with OpenAI SDK
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.freellmapi.com/v1",
apiKey: "freellmapi-<your-token>",
});
const stream = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: "Show streaming usage" }],
stream: true,
});
for await (const chunk of stream) {
const txt = chunk.choices[0].delta?.content ?? "";
process.stdout.write(txt);
}
Raw HTTP with cURL
curl -N -X POST https://api.freellmapi.com/v1/chat/completions \
-H "Authorization: Bearer freellmapi-<your-token>" \
-H "Content-Type: application/json" \
-d '{
"model":"gpt-4o-mini",
"messages":[{"role":"user","content":"Streaming demo"}],
"stream":true
}'
The -N flag disables buffering so each SSE chunk prints immediately upon arrival.
Summary
- FreeLLMAPI implements streaming through an abstract provider pattern defined in
server/src/providers/base.ts, ensuring consistent behavior across Groq, Cerebras, Gemini, and other providers. - Enable streaming by setting
stream=Truefor OpenAI-compatible endpoints or?alt=ssefor Google Gemini. - Process responses by iterating over chunks and extracting
delta.contentfields, watching forfinish_reasonto detect stream completion. - Tool-call deltas are supported within the stream, with function arguments arriving incrementally across multiple chunks.
- The proxy layer in
server/src/routes/proxy.tsforwards upstream SSE events without buffering, minimizing time-to-first-token latency.
Frequently Asked Questions
Does FreeLLMAPI support streaming for all LLM providers?
Yes, FreeLLMAPI supports streaming for every provider in its ecosystem, including OpenAI-compatible services (Groq, Cerebras, Mistral, OpenRouter, GitHub Models, HuggingFace, Cloudflare, and Cohere) and Google Gemini. According to the source code in server/src/providers/, each provider implements the abstract streamChatCompletion method to handle protocol-specific streaming logic.
How do I handle tool-call deltas in streaming responses?
Tool calls arrive incrementally across multiple SSE chunks. The first chunks contain delta.tool_calls with partial function names and arguments, while subsequent chunks stream the remaining JSON. Your client should accumulate these fragments until the chunk contains finish_reason: "tool_calls", indicating the complete function call is ready for execution.
What is the difference between OpenAI-compatible and Gemini streaming?
OpenAI-compatible providers use the standard stream=True JSON parameter and return deltas in the choices[0].delta format. Google Gemini requires the ?alt=sse query parameter and uses a different endpoint structure (streamGenerateContent), though FreeLLMAPI normalizes both formats to the OpenAI delta schema in server/src/providers/google.ts.
How can I measure streaming latency with FreeLLMAPI?
FreeLLMAPI logs the latency of the first streamed token for analytics purposes, as documented in docs/architecture.md. To measure this in your own client, record the timestamp when sending the request and compare it to the timestamp when receiving the first SSE chunk containing non-null delta.content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →