# How to Use Streaming with FreeLLMAPI: A Complete Developer Guide

> Learn how to use streaming with FreeLLMAPI for efficient token-by-token responses from LLM providers. A complete developer guide.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-08-29

---

**FreeLLMAPI supports true token-by-token streaming across all supported providers—including Groq, Cerebras, Gemini, and OpenAI-compatible endpoints—by setting `stream=True` in your request and iterating over SSE chunks containing `delta` fields.**

FreeLLMAPI is an open-source unified gateway that proxies requests to multiple large language model providers while preserving native streaming capabilities. According to the tashfeenahmed/freellmapi source code, the platform implements a provider-agnostic streaming architecture that forwards Server-Sent Events (SSE) directly from upstream services to your client without buffering. This guide explains how to activate streaming, consume token deltas, and handle provider-specific implementations using the repository's abstract provider pattern.

## Architecture of FreeLLMAPI Streaming

The streaming implementation in FreeLLMAPI is designed to be provider-agnostic, allowing uniform consumption across disparate LLM backends.

### The Provider Abstraction Layer

At the core of the system lies the abstract `Provider` class defined in [`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts). This base class declares the `streamChatCompletion` method, which all concrete providers must implement to handle streaming requests. The method signature enforces a consistent interface for initiating token-by-token responses regardless of the upstream service.

### Provider-Specific Implementations

Concrete provider implementations handle the protocol translation for their respective services:

- **[`server/src/providers/openai-compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/openai-compat.ts)**: Handles streaming for OpenAI-compatible providers including Groq, Cerebras, Mistral, OpenRouter, GitHub Models, HuggingFace, Cloudflare, and Cohere. This implementation forwards the `stream=True` parameter directly to the upstream API and yields chunks as they arrive.

- **[`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts)**: Implements Gemini-specific streaming logic using the `streamGenerateContent` endpoint. Unlike standard OpenAI-compatible streams, Gemini requires the `?alt=sse` query parameter to enable Server-Sent Events.

The HTTP router in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) pipes these chunks straight back to the client, preserving the SSE semantics defined in the OpenAI specification.

## Enabling Streaming in Your Requests

To activate streaming with FreeLLMAPI, you must explicitly set the streaming parameter in your request payload.

For **OpenAI-compatible providers**, include `"stream": true` in your JSON payload:

```json
{
  "model": "gpt-4o-mini",
  "messages": [{"role": "user", "content": "Hello"}],
  "stream": true
}

```

For **Google Gemini**, append `?alt=sse` to your request URL or use the provider-specific streaming flag, as implemented in [`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts).

## Consuming Streamed Responses

When streaming is enabled, FreeLLMAPI returns a series of JSON objects rather than a single complete response. Each chunk follows the OpenAI delta format:

- **`delta`**: Contains the next token string or tool-call fragment. Access this via `choices[0].delta.content` in client libraries.
- **`finish_reason`**: Signals stream termination. Values include `"stop"`, `"length"`, or `"tool_calls"`.
- **Tool-call deltas**: Initial chunks may contain `delta.tool_calls` with partial function arguments. The stream ends with `finish_reason: "tool_calls"` once the function call is complete.

## Implementation Examples for FreeLLMAPI Streaming

You can consume FreeLLMAPI streams using any HTTP client that supports Server-Sent Events.

### Python with OpenAI Client

```python
import openai

client = openai.OpenAI(
    base_url="https://api.freellmapi.com/v1",
    api_key="freellmapi-<your-token>"
)

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain streaming in FreeLLMAPI"}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

```

### Node.js with OpenAI SDK

```ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.freellmapi.com/v1",
  apiKey: "freellmapi-<your-token>",
});

const stream = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Show streaming usage" }],
  stream: true,
});

for await (const chunk of stream) {
  const txt = chunk.choices[0].delta?.content ?? "";
  process.stdout.write(txt);
}

```

### Raw HTTP with cURL

```bash
curl -N -X POST https://api.freellmapi.com/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-<your-token>" \
  -H "Content-Type: application/json" \
  -d '{
        "model":"gpt-4o-mini",
        "messages":[{"role":"user","content":"Streaming demo"}],
        "stream":true
      }'

```

The `-N` flag disables buffering so each SSE chunk prints immediately upon arrival.

## Summary

- **FreeLLMAPI** implements streaming through an abstract provider pattern defined in [`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts), ensuring consistent behavior across Groq, Cerebras, Gemini, and other providers.
- Enable streaming by setting `stream=True` for OpenAI-compatible endpoints or `?alt=sse` for Google Gemini.
- Process responses by iterating over chunks and extracting `delta.content` fields, watching for `finish_reason` to detect stream completion.
- Tool-call deltas are supported within the stream, with function arguments arriving incrementally across multiple chunks.
- The proxy layer in [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) forwards upstream SSE events without buffering, minimizing time-to-first-token latency.

## Frequently Asked Questions

### Does FreeLLMAPI support streaming for all LLM providers?

Yes, FreeLLMAPI supports streaming for every provider in its ecosystem, including OpenAI-compatible services (Groq, Cerebras, Mistral, OpenRouter, GitHub Models, HuggingFace, Cloudflare, and Cohere) and Google Gemini. According to the source code in `server/src/providers/`, each provider implements the abstract `streamChatCompletion` method to handle protocol-specific streaming logic.

### How do I handle tool-call deltas in streaming responses?

Tool calls arrive incrementally across multiple SSE chunks. The first chunks contain `delta.tool_calls` with partial function names and arguments, while subsequent chunks stream the remaining JSON. Your client should accumulate these fragments until the chunk contains `finish_reason: "tool_calls"`, indicating the complete function call is ready for execution.

### What is the difference between OpenAI-compatible and Gemini streaming?

OpenAI-compatible providers use the standard `stream=True` JSON parameter and return deltas in the `choices[0].delta` format. Google Gemini requires the `?alt=sse` query parameter and uses a different endpoint structure (`streamGenerateContent`), though FreeLLMAPI normalizes both formats to the OpenAI delta schema in [`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts).

### How can I measure streaming latency with FreeLLMAPI?

FreeLLMAPI logs the latency of the first streamed token for analytics purposes, as documented in [`docs/architecture.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/architecture.md). To measure this in your own client, record the timestamp when sending the request and compare it to the timestamp when receiving the first SSE chunk containing non-null `delta.content`.