How FreeLLMAPI Handles Streaming Responses: Architecture and Implementation
FreeLLMAPI streams LLM responses by delegating to each provider's streamChatCompletion async generator, which parses upstream Server-Sent Events (SSE) into standardized chunks and forwards them to the client through a persistent HTTP connection.
The tashfeenahmed/freellmapi repository implements real-time token streaming across heterogeneous LLM backends using a unified asynchronous generator pattern. When a client request includes "stream": true, the server routes execution to a provider-specific implementation that reads the upstream response as a raw byte stream, yielding parsed chunks incrementally. This design decouples protocol handling from business logic, ensuring consistent SSE delivery regardless of whether the backend is OpenAI, Google, Cohere, or Cloudflare.
The Provider Abstraction Layer
All streaming capabilities in FreeLLMAPI originate from the abstract base class defined in [server/src/providers/base.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/base.ts). Each provider must implement the streamChatCompletion method, which returns an AsyncGenerator<ChatChunk> yielding standardized response fragments.
// server/src/providers/base.ts
abstract streamChatCompletion(
apiKey: string,
messages: ChatMessage[],
modelId: string,
options?: ProviderOptions
): AsyncGenerator<ChatChunk>;
This abstraction allows the routing layer in [server/src/routes/responses.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) to consume streams polymorphically. When a request arrives with streaming enabled, the handler invokes the generator without knowing the underlying provider implementation details.
Concrete Provider Implementations
Each supported LLM provider implements its own byte-parsing logic within the async generator. For example, the OpenAI-compatible provider in [server/src/providers/openai-compat.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/openai-compat.ts) reads the upstream response body as a stream, buffering partial lines until complete SSE events are received before yielding parsed JSON chunks.
// server/src/providers/openai-compat.ts (simplified)
async *streamChatCompletion(apiKey, messages, modelId, options) {
const response = await fetch(this.endpoint, {
method: 'POST',
headers: this.authHeaders(apiKey),
body: JSON.stringify({ model: modelId, messages, stream: true, ...options })
});
const reader = response.body?.getReader();
let buffer = '';
while (reader) {
const { value, done } = await reader.read();
if (done) break;
buffer += new TextDecoder().decode(value);
const lines = buffer.split('\n');
buffer = lines.pop() ?? '';
for (const line of lines) {
if (line.startsWith('data:')) {
const payload = JSON.parse(line.slice(5).trim());
yield payload; // yields a ChatChunk
}
}
}
}
Similar generator implementations exist for Google ([server/src/providers/google.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/google.ts)), Cohere ([server/src/providers/cohere.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/cohere.ts)), Cloudflare ([server/src/providers/cloudflare.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/providers/cloudflare.ts)), and AI-Horde, each handling provider-specific SSE formatting while conforming to the same AsyncGenerator interface.
Request Routing and SSE Delivery
The HTTP handlers in [server/src/routes/responses.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) and [server/src/routes/proxy.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) act as the entry points for streaming requests. These routes verify the stream: true parameter, invoke the appropriate provider generator, and forward each yielded chunk to the client using Server-Sent Events format.
The utility module [server/src/lib/inbound-chat.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/inbound-chat.ts) orchestrates this transmission loop, consuming the async generator and writing properly formatted SSE data frames to the response stream.
// server/src/routes/responses.ts (conceptual flow)
const gen = route.provider.streamChatCompletion(
route.apiKey, messages, route.modelId, options
);
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache');
res.setHeader('Connection', 'keep-alive');
for await (const chunk of gen) {
res.write(`data: ${JSON.stringify(chunk)}\n\n`);
}
res.end();
Resilience and Fusion Streaming
FreeLLMAPI includes fault-tolerant streaming through the fusion service in [server/src/services/fusion.ts](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts). If a provider's stream terminates unexpectedly (timeout, 5xx error, or connection reset), the system records a failure penalty and can seamlessly transition to an alternative provider's stream without closing the client connection.
The fusion implementation aggregates multiple provider generators, yielding chunks from the fastest or most reliable upstream while masking partial failures. This ensures that once the client receives the initial data: event, the stream continues to completion even if individual backends fail mid-generation.
Client Integration Examples
Clients consume FreeLLMAPI streams by including "stream": true in the request payload and processing the response as a ReadableStream or EventSource.
Node.js Client Example:
const resp = await fetch('http://localhost:3000/v1/chat/completions', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
model: 'gpt-4',
messages: [{ role: 'user', content: 'Explain streaming' }],
stream: true
})
});
const reader = resp.body.getReader();
while (true) {
const { value, done } = await reader.read();
if (done) break;
const text = new TextDecoder().decode(value);
console.log(text); // Each SSE chunk: data: {...}\n\n
}
Browser EventSource Example:
const eventSource = new EventSource('/v1/chat/completions?stream=true');
eventSource.onmessage = (event) => {
const chunk = JSON.parse(event.data);
console.log(chunk.choices[0].delta.content);
};
Summary
- AsyncGenerator Abstraction: FreeLLMAPI standardizes streaming through the
streamChatCompletionmethod inserver/src/providers/base.ts, returningAsyncGenerator<ChatChunk>regardless of backend provider. - Provider-Specific Parsing: Each provider (OpenAI, Google, Cohere, etc.) implements custom SSE parsing logic in its respective file under
server/src/providers/, handling upstream protocol differences internally. - Unified SSE Output: The routing layer in
server/src/routes/responses.tsand utility functions inserver/src/lib/inbound-chat.tsconsume these generators and transmit chunks as standard Server-Sent Events to the client. - Resilient Streaming: The fusion service in
server/src/services/fusion.tsenables failover between providers mid-stream, ensuring continuity even when individual backends fail. - Simple Client Contract: Clients activate streaming with
"stream": trueand consume the response via standard HTTP streaming mechanisms (ReadableStream or EventSource).
Frequently Asked Questions
How do I enable streaming in a FreeLLMAPI request?
Add "stream": true to your JSON request body when calling the /v1/chat/completions endpoint. The server will switch from a synchronous JSON response to a streaming SSE response, yielding tokens as they are generated by the underlying LLM provider.
What response format does FreeLLMAPI use for streaming?
FreeLLMAPI uses Server-Sent Events (SSE) with the MIME type text/event-stream. Each generated token or chunk is sent as a line prefixed with data: followed by a JSON object containing the delta content, matching the OpenAI streaming specification for compatibility with existing client libraries.
Which LLM providers support streaming in FreeLLMAPI?
All major providers implemented in the repository support streaming, including OpenAI-compatible endpoints, Google Gemini, Cohere, Cloudflare Workers AI, and AI-Horde. Each provider implements the abstract streamChatCompletion method in server/src/providers/base.ts to handle provider-specific SSE parsing while exposing a uniform interface to the router.
How does FreeLLMAPI handle network errors during an active stream?
If a provider connection drops mid-stream, the fusion logic in server/src/services/fusion.ts captures the error, penalizes the failing provider, and attempts to continue generation using an alternative provider's stream if configured. The client connection remains open, and the stream continues with minimal interruption, though the content may switch context if a fallback provider generates different tokens.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →