How to Stream Agent Chat Using the LLM Wiki API: A Complete Implementation Guide

The LLM Wiki API provides a streamChat function in src/lib/llm-client.ts that enables real-time token streaming through Server-Sent Events (SSE), utilizing three callbacks—onToken, onError, and onDone—to handle live responses from providers like OpenAI and Anthropic.

The LLM Wiki repository offers a robust TypeScript implementation for streaming conversational AI responses. This guide explains how to stream agent chat using the LLM Wiki API by leveraging the internal client architecture, which abstracts provider-specific SSE handling into a unified interface callable from any frontend component.

Understanding the Streaming Architecture

The streaming pipeline in LLM Wiki consists of five coordinated layers that handle everything from state management to content filtering.

The Entry Point: streamChat Function

The public API surface resides in src/lib/llm-client.ts. The streamChat function orchestrates the entire streaming workflow, accepting a configuration object, an array of ChatMessage objects (containing role and content pairs), and a callbacks object with three methods:

  • onToken(token: string) – Invoked for each streamed token
  • onError(error: Error) – Invoked when network or API errors occur
  • onDone() – Invoked when the stream completes cleanly

State Management with the Chat Store

The chat-store.ts file in src/stores/ maintains conversation state using Zustand. It tracks a mutable streamingContent string that accumulates tokens as they arrive. Critical for UX, the store clears stale content when a new stream initiates to prevent cross-conversation bleed-through via setStreaming and appendStreaming actions.

Provider Abstraction Layer

Located in src/lib/llm-providers.ts, this layer constructs provider-specific request payloads for OpenAI, Anthropic, and other supported services. The stream boolean flag in the configuration toggles whether the provider returns an SSE stream versus a complete JSON response.

Network and SSE Parsing

The tauri-fetch.ts module in src/lib/ performs the actual HTTP request using Tauri's fetch API. For streaming responses, it reads the response body line-by-line, parsing each SSE line through parseOpenAiLine or the equivalent Anthropic parser, then forwards extracted tokens to the caller's onToken callback.

Content Filtering

Before tokens reach the UI, src/lib/reasoning-detector.ts filters out non-answer "reasoning" data that some models emit. This ensures only user-visible content populates the streamingContent state.

Implementing streamChat in Your Application

Because streamChat is a pure TypeScript function, you can invoke it from any frontend framework.

Basic TypeScript Implementation

Import the client and type definitions, then invoke the stream with your message history:

import { streamChat } from '@/lib/llm-client'
import type { ChatMessage } from '@/lib/chat-agent-types'

// Build the message history – the most recent user query is last
const messages: ChatMessage[] = [
  { role: 'system', content: 'You are an AI assistant.' },
  { role: 'user',   content: 'Explain the theory of relativity.' },
]

// Call the streaming API
await streamChat(
  {
    model: 'gpt-4o-mini',
    temperature: 0.7,
    maxTokens: 1024,
    streaming: true, // Required for live token delivery
  },
  messages,
  {
    onToken: (token) => {
      // Append token to UI or state store
      console.log('⧖', token)
    },
    onError: (err) => {
      console.error('Streaming error:', err)
    },
    onDone: () => {
      console.log('✅ Stream finished')
    },
  },
)

React Integration Example

For React applications, connect the callbacks to local state:

import { useEffect, useState } from 'react'
import { streamChat } from '@/lib/llm-client'

export default function Chat() {
  const [output, setOutput] = useState('')

  const send = async (prompt: string) => {
    await streamChat(
      { model: 'claude-3-haiku-20240307', streaming: true },
      [{ role: 'user', content: prompt }],
      {
        onToken: (t) => setOutput((prev) => prev + t),
        onError: (e) => console.error(e),
        onDone: () => console.log('finished'),
      },
    )
  }

  useEffect(() => {
    send('What are the main benefits of streaming LLM responses?')
  }, [])

  return <pre>{output}</pre>
}

Configuration and Callback Reference

The Configuration Object

The first parameter to streamChat accepts provider-specific settings:

  • model: String identifier (e.g., 'gpt-4o-mini', 'claude-3-haiku-20240307')
  • temperature: Number between 0 and 1 controlling randomness
  • maxTokens: Integer limiting response length
  • streaming: Boolean that must be true to enable SSE mode

Event Callbacks

The third parameter is a callbacks object implementing the streaming interface:

  1. onToken(token: string): Receives each decoded token string as the LLM generates it
  2. onError(error: Error): Handles HTTP failures, DNS issues, or API quota errors
  3. onDone(): Signals completion after the final SSE line is parsed

Summary

  • The streamChat function in src/lib/llm-client.ts serves as the primary interface for streaming agent chat using the LLM Wiki API
  • Enable streaming by setting streaming: true in the configuration object passed to the function
  • The architecture separates concerns across five files: chat-store.ts for state, llm-providers.ts for payloads, tauri-fetch.ts for network SSE parsing, and reasoning-detector.ts for content filtering
  • Implement three callbacks—onToken, onError, and onDone—to handle real-time data, errors, and completion states
  • The system supports multiple providers including OpenAI and Anthropic through unified TypeScript interfaces

Frequently Asked Questions

How do I enable streaming mode in the LLM Wiki API?

Set the streaming property to true in the configuration object passed to streamChat. According to the source code in src/lib/llm-providers.ts, this flag instructs the provider to return an SSE stream rather than a complete JSON response.

What file handles the SSE parsing for streaming responses?

The tauri-fetch.ts file in src/lib/ manages low-level HTTP requests and SSE parsing. It reads the response body line-by-line and uses provider-specific parsers like parseOpenAiLine to extract tokens from Server-Sent Events.

How does the chat store prevent content bleed between conversations?

The chat-store.ts implementation clears the streamingContent state when initiating a new stream. This ensures that residual tokens from previous conversations do not appear in new chat sessions, preventing cross-conversation contamination.

Can I use streamChat with providers other than OpenAI?

Yes. The llm-providers.ts abstraction supports multiple LLM providers including Anthropic. The streamChat function accepts a model identifier string that routes to the appropriate provider implementation while maintaining the same callback interface.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →