# How to Stream Responses Using FreeLLMAPI: Complete SSE Implementation Guide

> Learn to stream responses using FreeLLMAPI with this SSE implementation guide. Get real-time token delivery and automatic failover across endpoints.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: how-to-guide
- Published: 2026-09-02

---

**FreeLLMAPI implements a pure Server-Sent Events (SSE) streaming pipeline that normalizes upstream LLM outputs through provider-agnostic async generators, enabling real-time token delivery with automatic mid-stream failover across all compatible endpoints.**

FreeLLMAPI is an open-source unified gateway that aggregates multiple LLM providers behind a single API surface. When you stream responses using FreeLLMAPI, the system leverages a sophisticated SSE architecture to deliver tokens incrementally while handling upstream failures transparently. This guide walks through the core mechanisms defined in the `tashfeenahmed/freellmapi` repository, including the provider abstraction layer and intelligent routing logic.

## Core Streaming Architecture

The streaming system revolves around three fundamental components that bridge upstream providers to your client application.

### streamChatCompletion Generator

Every provider implementation exposes an async generator method named `streamChatCompletion` that yields raw SSE chunks from the upstream service. These generators live in provider-specific files such as [`server/src/providers/openai-compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/openai-compat.ts), [`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts), and [`server/src/providers/cohere.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/cohere.ts). Each generator normalizes proprietary streaming formats into OpenAI-style JSON deltas before yielding to the router, ensuring consistent `delta.content` objects regardless of the upstream source.

### Streaming Router

The router layer in [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) (handling `/v1/chat/completions`) and [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) (handling `/v1/responses`) consumes the generator and manages SSE framing. It performs **turn-integrity validation** to ensure clients receive complete responses or explicit error frames. The router also implements **mid-stream failover**—if an upstream fails before delivering the first payload byte, it transparently retries the next available provider key using the bandit router logic.

### Pipeline Documentation

The architectural specification in [`docs/architecture/03-streaming-pipeline.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/architecture/03-streaming-pipeline.md) defines the exact event ordering: dispatch → upstream fetch → text extraction via `streamChunkText()` and `streamReasoningText()` → final `[DONE]` frame. This document also specifies how tool-call deltas and reasoning content are handled during the stream.

## End-to-End Streaming Flow

Understanding how to stream responses using FreeLLMAPI requires following the request lifecycle through the codebase.

1. **Client Request**: The client sends a POST request to `/v1/chat/completions` with `stream: true` (or `?alt=sse` for Gemini endpoints). The request hits the dispatch logic in the router.

2. **Provider Dispatch**: The `dispatch()` function selects a provider key using quota and cooldown logic, then invokes `provider.streamChatCompletion()`.

3. **Upstream Generation**: The provider streams raw SSE chunks from the underlying LLM service, yielding them unchanged to maintain low latency.

4. **Normalization**: The router extracts payload content using `streamChunkText(chunk)` for `delta.content` and `streamReasoningText(chunk)` for reasoning blocks. The first chunk containing a `model` field is captured as `upstreamModel` for metadata tracking.

5. **SSE Framing**: Each yielded chunk is written to the HTTP response as a JSON frame prefixed with `data: `. Upon completion, the router appends `data: [DONE]`.

6. **Failover Logic**: If an error frame arrives before any payload is flushed, the router retries the next provider key. After the first payload is sent, errors surface as `stream_error` frames to the client.

7. **Observability**: On successful completion, the router calls `recordUpstreamSuccess()` to log usage metrics and model information.

## Runtime Configuration Variables

FreeLLMAPI exposes several environment variables to tune streaming behavior, documented in [`docs/env/01-variables.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/env/01-variables.md):

- **`PROVIDER_TIMEOUT_<PLATFORM>`**: Controls the first-byte deadline (default 60,000ms for most OpenAI-compatible providers, 180,000ms for NVIDIA).
- **`PROVIDER_STREAM_STALL_TIMEOUT_MS`**: Mid-stream watchdog timer (default 90,000ms) that aborts inactive connections and triggers failover.
- **`VALIDATE_TOOL_ARGUMENTS`**: Enables strict validation of tool-call arguments, surfacing failures as `invalid_tool_arguments`.
- **`MAX_CONSECUTIVE_UPSTREAM_FAILS`**: Circuit-breaker threshold; after N consecutive failures, the router returns HTTP 503 instead of continuing failover.

## Client Implementation Examples

You can consume the streaming API using standard HTTP clients or official SDKs.

### cURL with Server-Sent Events

Use the `-N` flag to disable buffering and see real-time token delivery:

```bash
curl -N -H "Content-Type: application/json" \
     -H "Authorization: Bearer $FREE_LLMAPI_KEY" \
     -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain quantum tunnelling"}],"stream":true}' \
     https://api.freellmapi.com/v1/chat/completions

```

The response streams lines starting with `data: ` followed by JSON delta objects, terminating with `data: [DONE]`.

### Node.js with OpenAI SDK

The OpenAI SDK works unchanged against FreeLLMAPI's compatible endpoints:

```javascript
import { OpenAI } from "openai";

const client = new OpenAI({
  apiKey: process.env.FREE_LLMAPI_KEY,
  baseURL: "https://api.freellmapi.com/v1",
});

const stream = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Explain quantum tunnelling" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

```

### Python with httpx

For manual SSE handling without heavy dependencies:

```python
import httpx
import json

api_key = "YOUR_FREE_LLMAPI_KEY"
url = "https://api.freellmapi.com/v1/chat/completions"

payload = {
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Explain quantum tunnelling"}],
    "stream": True,
}

with httpx.stream(
    "POST",
    url,
    json=payload,
    headers={"Authorization": f"Bearer {api_key}"},
    timeout=None,
) as response:
    for line in response.iter_lines():
        if line.startswith(b"data: "):
            data = json.loads(line[6:])
            if "choices" in data:
                print(data["choices"][0]["delta"].get("content", ""), end="", flush=True)
            elif data == "[DONE]":
                break

```

## Key Source Files

Understanding the implementation requires examining these specific files in the `tashfeenahmed/freellmapi` repository:

- **[`server/src/providers/base.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts)**: Defines the abstract `streamChatCompletion` method that all providers must implement.
- **[`server/src/providers/openai-compat.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/openai-compat.ts)**: Pass-through streaming implementation for OpenAI-compatible providers.
- **[`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts)**: Gemini provider that translates Google's streaming format into OpenAI-style frames.
- **[`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts)**: Core routing logic for `/v1/chat/completions` with failover support.
- **[`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts)**: Generic `/v1/responses` endpoint implementation.
- **[`docs/architecture/03-streaming-pipeline.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/architecture/03-streaming-pipeline.md)**: Comprehensive documentation of the SSE flow, error frames, and metadata extraction.
- **[`docs/env/01-variables.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/env/01-variables.md)**: Reference for all streaming-related environment variables.

## Summary

- FreeLLMAPI uses **pure Server-Sent Events (SSE)** for all streaming endpoints, supporting `/v1/chat/completions`, `/v1/responses`, and Gemini-compatible routes.
- The **provider abstraction layer** in `server/src/providers/` normalizes upstream formats through the `streamChatCompletion` async generator.
- **Mid-stream failover** occurs automatically in [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) when errors appear before the first payload byte.
- **Runtime tuning** is available via environment variables like `PROVIDER_STREAM_STALL_TIMEOUT_MS` and `MAX_CONSECUTIVE_UPSTREAM_FAILS`.
- Clients can consume streams using standard HTTP tools, the OpenAI SDK, or custom implementations that parse `data:` prefixed JSON lines.

## Frequently Asked Questions

### How does FreeLLMAPI handle streaming errors from upstream providers?

The streaming router in [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) implements turn-integrity validation. If an upstream error occurs before any content is flushed to the client, the router automatically retries the next available provider key. Once the first payload byte is sent, subsequent errors are surfaced as `stream_error` frames in the SSE stream rather than triggering silent retries.

### What is the difference between the `/v1/chat/completions` and `/v1/responses` streaming endpoints?

Both endpoints support SSE streaming, but `/v1/chat/completions` (handled in [`responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/responses.ts)) is the OpenAI-compatible surface, while `/v1/responses` (handled in [`proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/proxy.ts)) provides a generic proxy surface. Both use the same underlying `streamChatCompletion` generator and failover logic, differing only in request/response schema validation.

### Which environment variable controls streaming timeouts?

Two variables govern timeouts: `PROVIDER_TIMEOUT_<PLATFORM>` sets the first-byte deadline (default 60-180 seconds depending on provider), while `PROVIDER_STREAM_STALL_TIMEOUT_MS` (default 90,000ms) acts as a mid-stream watchdog that aborts inactive connections. These are documented in [`docs/env/01-variables.md`](https://github.com/tashfeenahmed/freellmapi/blob/main/docs/env/01-variables.md).

### Can I stream responses from Google Gemini models through FreeLLMAPI?

Yes, the Google provider in [`server/src/providers/google.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts) translates Gemini's streaming format into OpenAI-compatible SSE frames. For Gemini-specific endpoints, append `?alt=sse` to the request, or use the standard `/v1/chat/completions` endpoint with a Gemini model identifier to receive normalized streaming deltas.