How to Stream Responses Using FreeLLMAPI: Complete SSE Implementation Guide

FreeLLMAPI implements a pure Server-Sent Events (SSE) streaming pipeline that normalizes upstream LLM outputs through provider-agnostic async generators, enabling real-time token delivery with automatic mid-stream failover across all compatible endpoints.

FreeLLMAPI is an open-source unified gateway that aggregates multiple LLM providers behind a single API surface. When you stream responses using FreeLLMAPI, the system leverages a sophisticated SSE architecture to deliver tokens incrementally while handling upstream failures transparently. This guide walks through the core mechanisms defined in the tashfeenahmed/freellmapi repository, including the provider abstraction layer and intelligent routing logic.

Core Streaming Architecture

The streaming system revolves around three fundamental components that bridge upstream providers to your client application.

streamChatCompletion Generator

Every provider implementation exposes an async generator method named streamChatCompletion that yields raw SSE chunks from the upstream service. These generators live in provider-specific files such as server/src/providers/openai-compat.ts, server/src/providers/google.ts, and server/src/providers/cohere.ts. Each generator normalizes proprietary streaming formats into OpenAI-style JSON deltas before yielding to the router, ensuring consistent delta.content objects regardless of the upstream source.

Streaming Router

The router layer in server/src/routes/responses.ts (handling /v1/chat/completions) and server/src/routes/proxy.ts (handling /v1/responses) consumes the generator and manages SSE framing. It performs turn-integrity validation to ensure clients receive complete responses or explicit error frames. The router also implements mid-stream failover—if an upstream fails before delivering the first payload byte, it transparently retries the next available provider key using the bandit router logic.

Pipeline Documentation

The architectural specification in docs/architecture/03-streaming-pipeline.md defines the exact event ordering: dispatch → upstream fetch → text extraction via streamChunkText() and streamReasoningText() → final [DONE] frame. This document also specifies how tool-call deltas and reasoning content are handled during the stream.

End-to-End Streaming Flow

Understanding how to stream responses using FreeLLMAPI requires following the request lifecycle through the codebase.

  1. Client Request: The client sends a POST request to /v1/chat/completions with stream: true (or ?alt=sse for Gemini endpoints). The request hits the dispatch logic in the router.

  2. Provider Dispatch: The dispatch() function selects a provider key using quota and cooldown logic, then invokes provider.streamChatCompletion().

  3. Upstream Generation: The provider streams raw SSE chunks from the underlying LLM service, yielding them unchanged to maintain low latency.

  4. Normalization: The router extracts payload content using streamChunkText(chunk) for delta.content and streamReasoningText(chunk) for reasoning blocks. The first chunk containing a model field is captured as upstreamModel for metadata tracking.

  5. SSE Framing: Each yielded chunk is written to the HTTP response as a JSON frame prefixed with data: . Upon completion, the router appends data: [DONE].

  6. Failover Logic: If an error frame arrives before any payload is flushed, the router retries the next provider key. After the first payload is sent, errors surface as stream_error frames to the client.

  7. Observability: On successful completion, the router calls recordUpstreamSuccess() to log usage metrics and model information.

Runtime Configuration Variables

FreeLLMAPI exposes several environment variables to tune streaming behavior, documented in docs/env/01-variables.md:

  • PROVIDER_TIMEOUT_<PLATFORM>: Controls the first-byte deadline (default 60,000ms for most OpenAI-compatible providers, 180,000ms for NVIDIA).
  • PROVIDER_STREAM_STALL_TIMEOUT_MS: Mid-stream watchdog timer (default 90,000ms) that aborts inactive connections and triggers failover.
  • VALIDATE_TOOL_ARGUMENTS: Enables strict validation of tool-call arguments, surfacing failures as invalid_tool_arguments.
  • MAX_CONSECUTIVE_UPSTREAM_FAILS: Circuit-breaker threshold; after N consecutive failures, the router returns HTTP 503 instead of continuing failover.

Client Implementation Examples

You can consume the streaming API using standard HTTP clients or official SDKs.

cURL with Server-Sent Events

Use the -N flag to disable buffering and see real-time token delivery:

curl -N -H "Content-Type: application/json" \
     -H "Authorization: Bearer $FREE_LLMAPI_KEY" \
     -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Explain quantum tunnelling"}],"stream":true}' \
     https://api.freellmapi.com/v1/chat/completions

The response streams lines starting with data: followed by JSON delta objects, terminating with data: [DONE].

Node.js with OpenAI SDK

The OpenAI SDK works unchanged against FreeLLMAPI's compatible endpoints:

import { OpenAI } from "openai";

const client = new OpenAI({
  apiKey: process.env.FREE_LLMAPI_KEY,
  baseURL: "https://api.freellmapi.com/v1",
});

const stream = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Explain quantum tunnelling" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Python with httpx

For manual SSE handling without heavy dependencies:

import httpx
import json

api_key = "YOUR_FREE_LLMAPI_KEY"
url = "https://api.freellmapi.com/v1/chat/completions"

payload = {
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Explain quantum tunnelling"}],
    "stream": True,
}

with httpx.stream(
    "POST",
    url,
    json=payload,
    headers={"Authorization": f"Bearer {api_key}"},
    timeout=None,
) as response:
    for line in response.iter_lines():
        if line.startswith(b"data: "):
            data = json.loads(line[6:])
            if "choices" in data:
                print(data["choices"][0]["delta"].get("content", ""), end="", flush=True)
            elif data == "[DONE]":
                break

Key Source Files

Understanding the implementation requires examining these specific files in the tashfeenahmed/freellmapi repository:

Summary

  • FreeLLMAPI uses pure Server-Sent Events (SSE) for all streaming endpoints, supporting /v1/chat/completions, /v1/responses, and Gemini-compatible routes.
  • The provider abstraction layer in server/src/providers/ normalizes upstream formats through the streamChatCompletion async generator.
  • Mid-stream failover occurs automatically in server/src/routes/responses.ts when errors appear before the first payload byte.
  • Runtime tuning is available via environment variables like PROVIDER_STREAM_STALL_TIMEOUT_MS and MAX_CONSECUTIVE_UPSTREAM_FAILS.
  • Clients can consume streams using standard HTTP tools, the OpenAI SDK, or custom implementations that parse data: prefixed JSON lines.

Frequently Asked Questions

How does FreeLLMAPI handle streaming errors from upstream providers?

The streaming router in server/src/routes/responses.ts implements turn-integrity validation. If an upstream error occurs before any content is flushed to the client, the router automatically retries the next available provider key. Once the first payload byte is sent, subsequent errors are surfaced as stream_error frames in the SSE stream rather than triggering silent retries.

What is the difference between the /v1/chat/completions and /v1/responses streaming endpoints?

Both endpoints support SSE streaming, but /v1/chat/completions (handled in responses.ts) is the OpenAI-compatible surface, while /v1/responses (handled in proxy.ts) provides a generic proxy surface. Both use the same underlying streamChatCompletion generator and failover logic, differing only in request/response schema validation.

Which environment variable controls streaming timeouts?

Two variables govern timeouts: PROVIDER_TIMEOUT_<PLATFORM> sets the first-byte deadline (default 60-180 seconds depending on provider), while PROVIDER_STREAM_STALL_TIMEOUT_MS (default 90,000ms) acts as a mid-stream watchdog that aborts inactive connections. These are documented in docs/env/01-variables.md.

Can I stream responses from Google Gemini models through FreeLLMAPI?

Yes, the Google provider in server/src/providers/google.ts translates Gemini's streaming format into OpenAI-compatible SSE frames. For Gemini-specific endpoints, append ?alt=sse to the request, or use the standard /v1/chat/completions endpoint with a Gemini model identifier to receive normalized streaming deltas.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →