How Streaming Works in the Chat Completions API in G0DM0D3

The G0DM0D3 proxy implements Server-Sent Events (SSE) streaming by detecting the stream: true flag in api/routes/chat.ts, forwarding the request to OpenRouter with forced streaming enabled, and relaying incremental chunks back to the client while applying optional Short-Term Memory (STM) post-processing before emitting the final termination signal.

The Chat Completions endpoint in elder-plinius/G0DM0D3 supports real-time token streaming through a robust Server-Sent Events implementation. When clients set stream: true in their request payload, the system initiates a persistent connection to OpenRouter, processes the incoming byte stream via the ReadableStream API, and forwards incremental content chunks while preserving the full G0DM0D3 processing pipeline including post-generation corrections and metadata logging.

Detecting Streaming Mode and Configuring SSE Headers

Parsing the Stream Flag

The route handler in api/routes/chat.ts inspects the incoming request body to determine whether to enable streaming mode. By default, the system operates in standard blocking mode unless explicitly instructed otherwise.

const { …, stream = false, … } = req.body

This extraction occurs at lines 13-15 of the chat route, where the boolean stream parameter defaults to false to maintain backward compatibility with non-streaming clients.

Establishing the SSE Connection

When stream evaluates to true, the server immediately configures the response headers to establish a persistent HTTP connection compatible with Server-Sent Events. The implementation disables buffering at both the Node.js and reverse-proxy levels to ensure real-time delivery.

res.setHeader('Content-Type', 'text/event-stream')
res.setHeader('Cache-Control', 'no-cache')
res.setHeader('Connection', 'keep-alive')
res.setHeader('X-Accel-Buffering', 'no')
res.flushHeaders()

This header configuration, found at lines 33-38 of api/routes/chat.ts, prepares the client to receive a continuous stream of data: prefixed events rather than a single JSON payload.

Relaying Chunks from the Upstream Provider

Forcing Stream Mode on OpenRouter

Before dispatching the request to OpenRouter, the system constructs a payload that mirrors the OpenAI Chat Completions format but explicitly forces streaming regardless of the original client intent. This ensures the upstream provider returns incremental chunks that the proxy can process and forward.

const streamBody = {
  model,
  messages: pipeline.processedMessages,
  temperature: pipeline.finalParams.temperature,
  max_tokens,
  stream: true,
}

This payload construction at lines 44-52 ensures that the G0DM0D3 pipeline—including preprocessing modules like Parseltongue and Autotune—executes before streaming begins. The request is then sent to https://openrouter.ai/api/v1/chat/completions with automatic retry logic that implements exponential backoff on 429 rate-limit responses (lines 60-78).

Consuming the ReadableStream

The low-level stream consumption logic resides in src/lib/openrouter.ts, where the response body is treated as a ReadableStream. The implementation uses a TextDecoder to handle byte-to-string conversion while maintaining an internal buffer to accommodate partial JSON objects split across network packets.

const reader = upstreamRes!.body?.getReader()
const decoder = new TextDecoder()
let buffer = ''
while (true) {
  const { done, value } = await reader.read()
  if (done) break
  buffer += decoder.decode(value, { stream: true })
  // split on newlines, keep any incomplete line in `buffer`
}

This streaming loop at lines 31-38 processes incoming bytes incrementally, splitting on newline characters to isolate individual SSE events from OpenRouter.

Formatting and Forwarding Chunks

Each complete line starting with data: is parsed as JSON, and the incremental content field is extracted and repackaged into a standard OpenAI-compatible chunk format. The server then writes this chunk to the client connection.

const chunk = {
  id: completionId,
  object: 'chat.completion.chunk',
  created,
  model,
  choices: [{ index: 0, delta: { content }, finish_reason: null }],
}
res.write(`data: ${JSON.stringify(chunk)}\n\n`)

This forwarding logic at lines 88-100 of api/routes/chat.ts maintains the SSE protocol by double-newline termination while preserving metadata like the completion ID and model name for SDK compatibility.

Finalizing the Stream with Post-Processing

STM Corrections

When the upstream stream signals completion via data: [DONE], G0DM0D3 performs optional post-processing through the Short-Term Memory (STM) system. The accumulated fullContent string is passed to applySTMPost, which may apply transformations such as hedge reduction or direct mode adjustments.

If STM modifies the text, the server emits an additional correction chunk marked with stm_applied: true before the final termination signal (lines 124-136). This allows clients to receive the fully processed content even when corrections occur after generation completes.

Emitting Termination Signals

Regardless of whether STM corrections are applied, the server emits a final chunk with finish_reason: 'stop' followed by the mandatory SSE terminator data: [DONE]. In the event of upstream errors, the system sends a single error chunk with finish_reason: 'error' and then terminates the connection (lines 138-152).

Metadata Logging and Error Handling

Every streaming request is logged via recordEvent in src/lib/metadata.ts, capturing timestamps, pipeline configuration, model usage, and response length (lines 158-170 of api/routes/chat.ts). This metadata appears under the x_g0dm0d3 field in the response and provides analytics capabilities without breaking compatibility with standard OpenAI SDKs.

Reusable Streaming Generator

For internal services requiring programmatic access to streaming content, the library exposes streamMessage in src/lib/openrouter.ts. This generator function encapsulates the same low-level streaming logic but yields plain text fragments rather than formatted SSE events.

import { streamMessage } from './src/lib/openrouter'

async function demo() {
  const gen = streamMessage({
    messages: [{ role: 'user', content: 'Explain streaming.' }],
    model: 'nousresearch/hermes-4-70b',
    apiKey: process.env.OPENROUTER_API_KEY!,
  })

  for await (const txt of gen) {
    console.log(txt) // incremental text fragments
  }
}
demo()

Located at lines 71-85, this utility demonstrates the reusable streaming core used by the public endpoint and allows other services to consume token streams without managing SSE formatting.

Client Integration Examples

Node.js OpenAI SDK

import OpenAI from 'openai'

const client = new OpenAI({ 
  baseURL: 'https://your-api.com/v1', 
  apiKey: 'sk-dummy' 
})

for await (const chunk of client.chat.completions.create({
  model: 'nousresearch/hermes-4-70b',
  messages: [{ role: 'user', content: 'Explain streaming.' }],
  stream: true,
})) {
  process.stdout.write(chunk.choices[0].delta.content || '')
}

Direct HTTP with cURL

curl -X POST https://your-api.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model":"nousresearch/hermes-4-70b",
        "messages":[{"role":"user","content":"Explain streaming."}],
        "stream":true
      }'

The SSE response follows the standard format:

{
  "id":"chatcmpl-…",
  "object":"chat.completion.chunk",
  "created":1721451234,
  "model":"nousresearch/hermes-4-70b",
  "choices":[{"index":0,"delta":{"content":"…"},"finish_reason":null}]
}

Summary

  • Streaming detection occurs in api/routes/chat.ts by reading the stream flag from the request body, defaulting to false for standard blocking responses.
  • SSE headers including text/event-stream and X-Accel-Buffering: no are set before calling res.flushHeaders() to establish a persistent connection.
  • Upstream requests to OpenRouter explicitly set stream: true to force incremental responses, with exponential backoff retry logic for rate limits.
  • Chunk processing uses the ReadableStream API in src/lib/openrouter.ts to parse data: prefixed lines and extract incremental content.
  • Post-processing via STM modules may emit correction chunks before the final finish_reason: 'stop' signal and data: [DONE] terminator.
  • Metadata recording via recordEvent captures analytics for every stream while maintaining OpenAI SDK compatibility.

Frequently Asked Questions

How does G0DM0D3 handle network interruptions during streaming?

If the connection to OpenRouter fails or returns a non-2xx status code after retries, the server sends a single SSE chunk with finish_reason: 'error' and immediately terminates the stream with data: [DONE]. The error is also logged via recordEvent for debugging purposes.

Can I use standard OpenAI SDKs with the G0DM0D3 streaming endpoint?

Yes. The response format mirrors the OpenAI Chat Completions API exactly, including the chat.completion.chunk object type and delta.content field structure. The additional x_g0dm0d3 metadata field is ignored by standard SDKs but available for custom clients.

What is the purpose of the streamMessage generator function?

The streamMessage function in src/lib/openrouter.ts provides a reusable, programmatic interface to streaming that yields raw text fragments rather than formatted SSE events. It is designed for internal services that need to consume streams without HTTP request/response handling.

Does enabling streaming bypass the G0DM0D3 preprocessing pipeline?

No. The system applies all preprocessing modules—including Parseltongue transformations and Autotune parameter adjustments—before sending the request to OpenRouter. Streaming only affects how the response is delivered, not how the input is processed.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →