How to Stream Agent Chat Using the LLM Wiki API: A Complete Implementation Guide
The LLM Wiki API provides a streamChat function in src/lib/llm-client.ts that enables real-time token streaming through Server-Sent Events (SSE), utilizing three callbacks—onToken, onError, and onDone—to handle live responses from providers like OpenAI and Anthropic.
The LLM Wiki repository offers a robust TypeScript implementation for streaming conversational AI responses. This guide explains how to stream agent chat using the LLM Wiki API by leveraging the internal client architecture, which abstracts provider-specific SSE handling into a unified interface callable from any frontend component.
Understanding the Streaming Architecture
The streaming pipeline in LLM Wiki consists of five coordinated layers that handle everything from state management to content filtering.
The Entry Point: streamChat Function
The public API surface resides in src/lib/llm-client.ts. The streamChat function orchestrates the entire streaming workflow, accepting a configuration object, an array of ChatMessage objects (containing role and content pairs), and a callbacks object with three methods:
onToken(token: string)– Invoked for each streamed tokenonError(error: Error)– Invoked when network or API errors occuronDone()– Invoked when the stream completes cleanly
State Management with the Chat Store
The chat-store.ts file in src/stores/ maintains conversation state using Zustand. It tracks a mutable streamingContent string that accumulates tokens as they arrive. Critical for UX, the store clears stale content when a new stream initiates to prevent cross-conversation bleed-through via setStreaming and appendStreaming actions.
Provider Abstraction Layer
Located in src/lib/llm-providers.ts, this layer constructs provider-specific request payloads for OpenAI, Anthropic, and other supported services. The stream boolean flag in the configuration toggles whether the provider returns an SSE stream versus a complete JSON response.
Network and SSE Parsing
The tauri-fetch.ts module in src/lib/ performs the actual HTTP request using Tauri's fetch API. For streaming responses, it reads the response body line-by-line, parsing each SSE line through parseOpenAiLine or the equivalent Anthropic parser, then forwards extracted tokens to the caller's onToken callback.
Content Filtering
Before tokens reach the UI, src/lib/reasoning-detector.ts filters out non-answer "reasoning" data that some models emit. This ensures only user-visible content populates the streamingContent state.
Implementing streamChat in Your Application
Because streamChat is a pure TypeScript function, you can invoke it from any frontend framework.
Basic TypeScript Implementation
Import the client and type definitions, then invoke the stream with your message history:
import { streamChat } from '@/lib/llm-client'
import type { ChatMessage } from '@/lib/chat-agent-types'
// Build the message history – the most recent user query is last
const messages: ChatMessage[] = [
{ role: 'system', content: 'You are an AI assistant.' },
{ role: 'user', content: 'Explain the theory of relativity.' },
]
// Call the streaming API
await streamChat(
{
model: 'gpt-4o-mini',
temperature: 0.7,
maxTokens: 1024,
streaming: true, // Required for live token delivery
},
messages,
{
onToken: (token) => {
// Append token to UI or state store
console.log('⧖', token)
},
onError: (err) => {
console.error('Streaming error:', err)
},
onDone: () => {
console.log('✅ Stream finished')
},
},
)
React Integration Example
For React applications, connect the callbacks to local state:
import { useEffect, useState } from 'react'
import { streamChat } from '@/lib/llm-client'
export default function Chat() {
const [output, setOutput] = useState('')
const send = async (prompt: string) => {
await streamChat(
{ model: 'claude-3-haiku-20240307', streaming: true },
[{ role: 'user', content: prompt }],
{
onToken: (t) => setOutput((prev) => prev + t),
onError: (e) => console.error(e),
onDone: () => console.log('finished'),
},
)
}
useEffect(() => {
send('What are the main benefits of streaming LLM responses?')
}, [])
return <pre>{output}</pre>
}
Configuration and Callback Reference
The Configuration Object
The first parameter to streamChat accepts provider-specific settings:
model: String identifier (e.g.,'gpt-4o-mini','claude-3-haiku-20240307')temperature: Number between 0 and 1 controlling randomnessmaxTokens: Integer limiting response lengthstreaming: Boolean that must betrueto enable SSE mode
Event Callbacks
The third parameter is a callbacks object implementing the streaming interface:
onToken(token: string): Receives each decoded token string as the LLM generates itonError(error: Error): Handles HTTP failures, DNS issues, or API quota errorsonDone(): Signals completion after the final SSE line is parsed
Summary
- The
streamChatfunction insrc/lib/llm-client.tsserves as the primary interface for streaming agent chat using the LLM Wiki API - Enable streaming by setting
streaming: truein the configuration object passed to the function - The architecture separates concerns across five files:
chat-store.tsfor state,llm-providers.tsfor payloads,tauri-fetch.tsfor network SSE parsing, andreasoning-detector.tsfor content filtering - Implement three callbacks—
onToken,onError, andonDone—to handle real-time data, errors, and completion states - The system supports multiple providers including OpenAI and Anthropic through unified TypeScript interfaces
Frequently Asked Questions
How do I enable streaming mode in the LLM Wiki API?
Set the streaming property to true in the configuration object passed to streamChat. According to the source code in src/lib/llm-providers.ts, this flag instructs the provider to return an SSE stream rather than a complete JSON response.
What file handles the SSE parsing for streaming responses?
The tauri-fetch.ts file in src/lib/ manages low-level HTTP requests and SSE parsing. It reads the response body line-by-line and uses provider-specific parsers like parseOpenAiLine to extract tokens from Server-Sent Events.
How does the chat store prevent content bleed between conversations?
The chat-store.ts implementation clears the streamingContent state when initiating a new stream. This ensures that residual tokens from previous conversations do not appear in new chat sessions, preventing cross-conversation contamination.
Can I use streamChat with providers other than OpenAI?
Yes. The llm-providers.ts abstraction supports multiple LLM providers including Anthropic. The streamChat function accepts a model identifier string that routes to the appropriate provider implementation while maintaining the same callback interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →