# How Streaming Works in the Chat Completions API in G0DM0D3

> Learn how streaming works in the Chat Completions API with G0DM0D3. Discover how SSE and OpenRouter enable real-time responses and optional post-processing for efficient AI interaction.

- Repository: [pliny/G0DM0D3](https://github.com/elder-plinius/G0DM0D3)
- Tags: internals
- Published: 2026-07-19

---

**The G0DM0D3 proxy implements Server-Sent Events (SSE) streaming by detecting the `stream: true` flag in [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts), forwarding the request to OpenRouter with forced streaming enabled, and relaying incremental chunks back to the client while applying optional Short-Term Memory (STM) post-processing before emitting the final termination signal.**

The Chat Completions endpoint in elder-plinius/G0DM0D3 supports real-time token streaming through a robust Server-Sent Events implementation. When clients set `stream: true` in their request payload, the system initiates a persistent connection to OpenRouter, processes the incoming byte stream via the **ReadableStream** API, and forwards incremental content chunks while preserving the full G0DM0D3 processing pipeline including post-generation corrections and metadata logging.

## Detecting Streaming Mode and Configuring SSE Headers

### Parsing the Stream Flag

The route handler in [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts) inspects the incoming request body to determine whether to enable streaming mode. By default, the system operates in standard blocking mode unless explicitly instructed otherwise.

```typescript
const { …, stream = false, … } = req.body

```

This extraction occurs at lines 13-15 of the chat route, where the boolean `stream` parameter defaults to `false` to maintain backward compatibility with non-streaming clients.

### Establishing the SSE Connection

When `stream` evaluates to `true`, the server immediately configures the response headers to establish a persistent HTTP connection compatible with Server-Sent Events. The implementation disables buffering at both the Node.js and reverse-proxy levels to ensure real-time delivery.

```typescript
res.setHeader('Content-Type', 'text/event-stream')
res.setHeader('Cache-Control', 'no-cache')
res.setHeader('Connection', 'keep-alive')
res.setHeader('X-Accel-Buffering', 'no')
res.flushHeaders()

```

This header configuration, found at lines 33-38 of [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts), prepares the client to receive a continuous stream of `data:` prefixed events rather than a single JSON payload.

## Relaying Chunks from the Upstream Provider

### Forcing Stream Mode on OpenRouter

Before dispatching the request to OpenRouter, the system constructs a payload that mirrors the OpenAI Chat Completions format but explicitly forces streaming regardless of the original client intent. This ensures the upstream provider returns incremental chunks that the proxy can process and forward.

```typescript
const streamBody = {
  model,
  messages: pipeline.processedMessages,
  temperature: pipeline.finalParams.temperature,
  max_tokens,
  stream: true,
}

```

This payload construction at lines 44-52 ensures that the G0DM0D3 pipeline—including preprocessing modules like **Parseltongue** and **Autotune**—executes before streaming begins. The request is then sent to `https://openrouter.ai/api/v1/chat/completions` with automatic retry logic that implements exponential backoff on `429` rate-limit responses (lines 60-78).

### Consuming the ReadableStream

The low-level stream consumption logic resides in [`src/lib/openrouter.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/openrouter.ts), where the response body is treated as a **ReadableStream**. The implementation uses a `TextDecoder` to handle byte-to-string conversion while maintaining an internal buffer to accommodate partial JSON objects split across network packets.

```typescript
const reader = upstreamRes!.body?.getReader()
const decoder = new TextDecoder()
let buffer = ''
while (true) {
  const { done, value } = await reader.read()
  if (done) break
  buffer += decoder.decode(value, { stream: true })
  // split on newlines, keep any incomplete line in `buffer`
}

```

This streaming loop at lines 31-38 processes incoming bytes incrementally, splitting on newline characters to isolate individual SSE events from OpenRouter.

### Formatting and Forwarding Chunks

Each complete line starting with `data: ` is parsed as JSON, and the incremental `content` field is extracted and repackaged into a standard OpenAI-compatible chunk format. The server then writes this chunk to the client connection.

```typescript
const chunk = {
  id: completionId,
  object: 'chat.completion.chunk',
  created,
  model,
  choices: [{ index: 0, delta: { content }, finish_reason: null }],
}
res.write(`data: ${JSON.stringify(chunk)}\n\n`)

```

This forwarding logic at lines 88-100 of [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts) maintains the SSE protocol by double-newline termination while preserving metadata like the completion ID and model name for SDK compatibility.

## Finalizing the Stream with Post-Processing

### STM Corrections

When the upstream stream signals completion via `data: [DONE]`, G0DM0D3 performs optional post-processing through the Short-Term Memory (STM) system. The accumulated `fullContent` string is passed to `applySTMPost`, which may apply transformations such as hedge reduction or direct mode adjustments.

If STM modifies the text, the server emits an additional correction chunk marked with `stm_applied: true` before the final termination signal (lines 124-136). This allows clients to receive the fully processed content even when corrections occur after generation completes.

### Emitting Termination Signals

Regardless of whether STM corrections are applied, the server emits a final chunk with `finish_reason: 'stop'` followed by the mandatory SSE terminator `data: [DONE]`. In the event of upstream errors, the system sends a single error chunk with `finish_reason: 'error'` and then terminates the connection (lines 138-152).

## Metadata Logging and Error Handling

Every streaming request is logged via `recordEvent` in [`src/lib/metadata.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/metadata.ts), capturing timestamps, pipeline configuration, model usage, and response length (lines 158-170 of [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts)). This metadata appears under the `x_g0dm0d3` field in the response and provides analytics capabilities without breaking compatibility with standard OpenAI SDKs.

## Reusable Streaming Generator

For internal services requiring programmatic access to streaming content, the library exposes `streamMessage` in [`src/lib/openrouter.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/openrouter.ts). This generator function encapsulates the same low-level streaming logic but yields plain text fragments rather than formatted SSE events.

```typescript
import { streamMessage } from './src/lib/openrouter'

async function demo() {
  const gen = streamMessage({
    messages: [{ role: 'user', content: 'Explain streaming.' }],
    model: 'nousresearch/hermes-4-70b',
    apiKey: process.env.OPENROUTER_API_KEY!,
  })

  for await (const txt of gen) {
    console.log(txt) // incremental text fragments
  }
}
demo()

```

Located at lines 71-85, this utility demonstrates the reusable streaming core used by the public endpoint and allows other services to consume token streams without managing SSE formatting.

## Client Integration Examples

### Node.js OpenAI SDK

```typescript
import OpenAI from 'openai'

const client = new OpenAI({ 
  baseURL: 'https://your-api.com/v1', 
  apiKey: 'sk-dummy' 
})

for await (const chunk of client.chat.completions.create({
  model: 'nousresearch/hermes-4-70b',
  messages: [{ role: 'user', content: 'Explain streaming.' }],
  stream: true,
})) {
  process.stdout.write(chunk.choices[0].delta.content || '')
}

```

### Direct HTTP with cURL

```bash
curl -X POST https://your-api.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model":"nousresearch/hermes-4-70b",
        "messages":[{"role":"user","content":"Explain streaming."}],
        "stream":true
      }'

```

The SSE response follows the standard format:

```json
{
  "id":"chatcmpl-…",
  "object":"chat.completion.chunk",
  "created":1721451234,
  "model":"nousresearch/hermes-4-70b",
  "choices":[{"index":0,"delta":{"content":"…"},"finish_reason":null}]
}

```

## Summary

- **Streaming detection** occurs in [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts) by reading the `stream` flag from the request body, defaulting to `false` for standard blocking responses.
- **SSE headers** including `text/event-stream` and `X-Accel-Buffering: no` are set before calling `res.flushHeaders()` to establish a persistent connection.
- **Upstream requests** to OpenRouter explicitly set `stream: true` to force incremental responses, with exponential backoff retry logic for rate limits.
- **Chunk processing** uses the **ReadableStream** API in [`src/lib/openrouter.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/openrouter.ts) to parse `data:` prefixed lines and extract incremental content.
- **Post-processing** via STM modules may emit correction chunks before the final `finish_reason: 'stop'` signal and `data: [DONE]` terminator.
- **Metadata recording** via `recordEvent` captures analytics for every stream while maintaining OpenAI SDK compatibility.

## Frequently Asked Questions

### How does G0DM0D3 handle network interruptions during streaming?

If the connection to OpenRouter fails or returns a non-2xx status code after retries, the server sends a single SSE chunk with `finish_reason: 'error'` and immediately terminates the stream with `data: [DONE]`. The error is also logged via `recordEvent` for debugging purposes.

### Can I use standard OpenAI SDKs with the G0DM0D3 streaming endpoint?

Yes. The response format mirrors the OpenAI Chat Completions API exactly, including the `chat.completion.chunk` object type and `delta.content` field structure. The additional `x_g0dm0d3` metadata field is ignored by standard SDKs but available for custom clients.

### What is the purpose of the `streamMessage` generator function?

The `streamMessage` function in [`src/lib/openrouter.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/openrouter.ts) provides a reusable, programmatic interface to streaming that yields raw text fragments rather than formatted SSE events. It is designed for internal services that need to consume streams without HTTP request/response handling.

### Does enabling streaming bypass the G0DM0D3 preprocessing pipeline?

No. The system applies all preprocessing modules—including **Parseltongue** transformations and **Autotune** parameter adjustments—before sending the request to OpenRouter. Streaming only affects how the response is delivered, not how the input is processed.