# How FreeLLMAPI Handles Streaming Responses and X‑Routed‑Via Headers

> Discover how FreeLLMAPI streams responses and uses X-Routed-Via headers to trace requests through LLM providers, ensuring transparency and efficiency.

- Repository: [Tashfeen/freellmapi](https://github.com/tashfeenahmed/freellmapi)
- Tags: internals
- Published: 2026-08-31

---

**FreeLLMAPI streams data directly from upstream LLM providers to clients using async‑iterator pipelines, while injecting an `X‑Routed‑Via` header to identify exactly which platform and model handled the request.**

FreeLLMAPI acts as a unified gateway for multiple LLM providers, implementing real‑time streaming and transparent routing metadata. This article examines the streaming pipeline architecture and the `X‑Routed‑Via` header injection mechanism based on the source code in the `tashfeenahmed/freellmapi` repository.

## The Streaming Architecture

FreeLLMAPI implements true streaming by piping data directly from the upstream provider to the client without buffering the complete response. This occurs in the **Fusion** service, which coordinates all LLM calls across the system.

### The Fusion Service Pipeline

In [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts), the router invokes the provider’s `streamChatCompletion` method and immediately forwards every chunk to the HTTP response:

```ts
for await (const chunk of route.provider.streamChatCompletion(
    route.apiKey,
    messages,
    route.modelId,
    options,
)) {
  stream.write(chunk);
}

```

The `for await...of` loop processes the async iterator returned by the provider, writing each chunk to the response stream as it arrives. This pattern applies consistently across all streaming endpoints, including `/v1/chat/completions` and `/v1/completions`.

### Error Handling During Streaming

If a chunk throws an error—such as a 5xx upstream failure or an empty‑stream response—the loop aborts immediately. The request is marked as failed, and the error propagates to the client. This prevents streams from hanging and ensures clients receive immediate feedback when upstream services fail.

## Setting the X‑Routed‑Via Header

Every response includes an `X‑Routed‑Via` header that reveals which provider and model served the request. This enables client‑side observability and debugging across the multi‑provider architecture.

### Header Construction in header‑value.ts

The `routedViaValue` function in [`server/src/lib/header-value.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/header-value.ts) sanitizes the platform and model identifier:

```ts
export function routedViaValue(platform: string, modelId: string): string {
  const raw = `${platform}/${modelId}`;
  return safeHeaderValue(raw);
}

```

This produces a URL‑encoded string in the format `<platform>/<modelId>` that is safe for HTTP transmission. The `safeHeaderValue` helper ensures no invalid characters reach the response headers.

### Response Integration in Route Handlers

Before the first byte streams to the client, route handlers set the header using the sanitized value. In [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts):

```ts
res.setHeader('X-Routed-Via', routedViaValue(route.platform, route.modelId));

```

The same pattern appears in [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) for standard API responses. If the request is served from cache, the header value is set to `"cache"` instead of the provider identifier.

## Practical Implementation Examples

### Node.js Client Implementation

To consume the streaming endpoint and inspect the routing header:

```js
import fetch from 'node-fetch';

const resp = await fetch('https://api.free.llmapi.com/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'Authorization': 'Bearer <your-api-key>',
  },
  body: JSON.stringify({
    model: 'openai/gpt-4',
    stream: true,
    messages: [{ role: 'user', content: 'Explain quantum tunneling.' }],
  }),
});

console.log('Routed via:', resp.headers.get('X-Routed-Via'));

const decoder = new TextDecoder();
for await (const chunk of resp.body) {
  process.stdout.write(decoder.decode(chunk));
}

```

### Verifying Headers with cURL

Use the `-N` flag to disable buffering and view the header in real‑time:

```bash
curl -N -H "Authorization: Bearer $API_KEY" \
     -H "Content-Type: application/json" \
     -d '{"model":"openai/gpt-4","stream":true,"messages":[{"role":"user","content":"Hello"}]}' \
     https://api.free.llmapi.com/v1/chat/completions \
  -i | grep X-Routed-Via

```

This outputs the provider routing information (e.g., `X-Routed-Via: openai/gpt-4`) alongside the streamed content.

## Summary

*   **Async Iterator Pipeline**: [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts) uses `for await...of` loops to stream chunks directly from the provider’s `streamChatCompletion` method to the client without intermediate buffering.
*   **X‑Routed‑Via Construction**: The `routedViaValue` function in [`server/src/lib/header-value.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/header-value.ts) sanitizes provider and model names into a safe `<platform>/<model>` format.
*   **Route‑Level Injection**: Both [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) and [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) set the header before transmitting the first chunk to the client.
*   **Error Propagation**: Stream failures abort the loop immediately, propagating errors to the client rather than leaving connections hanging.
*   **Observability Values**: The header supports `"cache"` for cached responses or `"platform/model"` for live upstream calls, enabling precise request tracing.

## Frequently Asked Questions

### How does FreeLLMAPI handle upstream failures during streaming?

When the `for await...of` loop in [`server/src/services/fusion.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/fusion.ts) encounters an error from the provider’s async iterator—such as a 5xx error or empty stream—it aborts the loop immediately. The error propagates to the client, and the request is marked as failed in the logs. This ensures clients receive immediate notification rather than waiting for a connection timeout.

### What format does the X‑Routed‑Via header use?

The header uses a `<platform>/<model>` format (e.g., `openai/gpt-4`), constructed by the `routedViaValue` function in [`server/src/lib/header-value.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/lib/header-value.ts). The function URL‑encodes the value to ensure HTTP safety. For cached responses, the header value is explicitly set to `"cache"`.

### Can I disable streaming and still receive the X‑Routed‑Via header?

Yes. The `X‑Routed‑Via` header is set in both streaming and non‑streaming responses. The [`server/src/routes/responses.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/responses.ts) handler sets the header for standard synchronous responses, while [`server/src/routes/proxy.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/routes/proxy.ts) handles it for streaming and cached scenarios.

### Which providers support streaming in FreeLLMAPI?

Any provider that implements the `streamChatCompletion` method in their adapter supports streaming. The router checks for this method’s availability at runtime in [`server/src/services/router.ts`](https://github.com/tashfeenahmed/freellmapi/blob/main/server/src/services/router.ts). If a provider lacks streaming support, the system falls back to non‑streaming responses or alternative providers based on the routing configuration.