How FreeLLMAPI Handles Streaming Responses and X‑Routed‑Via Headers

FreeLLMAPI streams data directly from upstream LLM providers to clients using async‑iterator pipelines, while injecting an X‑Routed‑Via header to identify exactly which platform and model handled the request.

FreeLLMAPI acts as a unified gateway for multiple LLM providers, implementing real‑time streaming and transparent routing metadata. This article examines the streaming pipeline architecture and the X‑Routed‑Via header injection mechanism based on the source code in the tashfeenahmed/freellmapi repository.

The Streaming Architecture

FreeLLMAPI implements true streaming by piping data directly from the upstream provider to the client without buffering the complete response. This occurs in the Fusion service, which coordinates all LLM calls across the system.

The Fusion Service Pipeline

In server/src/services/fusion.ts, the router invokes the provider’s streamChatCompletion method and immediately forwards every chunk to the HTTP response:

for await (const chunk of route.provider.streamChatCompletion(
    route.apiKey,
    messages,
    route.modelId,
    options,
)) {
  stream.write(chunk);
}

The for await...of loop processes the async iterator returned by the provider, writing each chunk to the response stream as it arrives. This pattern applies consistently across all streaming endpoints, including /v1/chat/completions and /v1/completions.

Error Handling During Streaming

If a chunk throws an error—such as a 5xx upstream failure or an empty‑stream response—the loop aborts immediately. The request is marked as failed, and the error propagates to the client. This prevents streams from hanging and ensures clients receive immediate feedback when upstream services fail.

Setting the X‑Routed‑Via Header

Every response includes an X‑Routed‑Via header that reveals which provider and model served the request. This enables client‑side observability and debugging across the multi‑provider architecture.

Header Construction in header‑value.ts

The routedViaValue function in server/src/lib/header-value.ts sanitizes the platform and model identifier:

export function routedViaValue(platform: string, modelId: string): string {
  const raw = `${platform}/${modelId}`;
  return safeHeaderValue(raw);
}

This produces a URL‑encoded string in the format <platform>/<modelId> that is safe for HTTP transmission. The safeHeaderValue helper ensures no invalid characters reach the response headers.

Response Integration in Route Handlers

Before the first byte streams to the client, route handlers set the header using the sanitized value. In server/src/routes/proxy.ts:

res.setHeader('X-Routed-Via', routedViaValue(route.platform, route.modelId));

The same pattern appears in server/src/routes/responses.ts for standard API responses. If the request is served from cache, the header value is set to "cache" instead of the provider identifier.

Practical Implementation Examples

Node.js Client Implementation

To consume the streaming endpoint and inspect the routing header:

import fetch from 'node-fetch';

const resp = await fetch('https://api.free.llmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'Authorization': 'Bearer <your-api-key>',
  },
  body: JSON.stringify({
    model: 'openai/gpt-4',
    stream: true,
    messages: [{ role: 'user', content: 'Explain quantum tunneling.' }],
  }),
});

console.log('Routed via:', resp.headers.get('X-Routed-Via'));

const decoder = new TextDecoder();
for await (const chunk of resp.body) {
  process.stdout.write(decoder.decode(chunk));
}

Verifying Headers with cURL

Use the -N flag to disable buffering and view the header in real‑time:

curl -N -H "Authorization: Bearer $API_KEY" \
     -H "Content-Type: application/json" \
     -d '{"model":"openai/gpt-4","stream":true,"messages":[{"role":"user","content":"Hello"}]}' \
     https://api.free.llmapi.com/v1/chat/completions \
  -i | grep X-Routed-Via

This outputs the provider routing information (e.g., X-Routed-Via: openai/gpt-4) alongside the streamed content.

Summary

  • Async Iterator Pipeline: server/src/services/fusion.ts uses for await...of loops to stream chunks directly from the provider’s streamChatCompletion method to the client without intermediate buffering.
  • X‑Routed‑Via Construction: The routedViaValue function in server/src/lib/header-value.ts sanitizes provider and model names into a safe <platform>/<model> format.
  • Route‑Level Injection: Both server/src/routes/proxy.ts and server/src/routes/responses.ts set the header before transmitting the first chunk to the client.
  • Error Propagation: Stream failures abort the loop immediately, propagating errors to the client rather than leaving connections hanging.
  • Observability Values: The header supports "cache" for cached responses or "platform/model" for live upstream calls, enabling precise request tracing.

Frequently Asked Questions

How does FreeLLMAPI handle upstream failures during streaming?

When the for await...of loop in server/src/services/fusion.ts encounters an error from the provider’s async iterator—such as a 5xx error or empty stream—it aborts the loop immediately. The error propagates to the client, and the request is marked as failed in the logs. This ensures clients receive immediate notification rather than waiting for a connection timeout.

What format does the X‑Routed‑Via header use?

The header uses a <platform>/<model> format (e.g., openai/gpt-4), constructed by the routedViaValue function in server/src/lib/header-value.ts. The function URL‑encodes the value to ensure HTTP safety. For cached responses, the header value is explicitly set to "cache".

Can I disable streaming and still receive the X‑Routed‑Via header?

Yes. The X‑Routed‑Via header is set in both streaming and non‑streaming responses. The server/src/routes/responses.ts handler sets the header for standard synchronous responses, while server/src/routes/proxy.ts handles it for streaming and cached scenarios.

Which providers support streaming in FreeLLMAPI?

Any provider that implements the streamChatCompletion method in their adapter supports streaming. The router checks for this method’s availability at runtime in server/src/services/router.ts. If a provider lacks streaming support, the system falls back to non‑streaming responses or alternative providers based on the routing configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →