How Provider-Specific Chat Services in 5ire Handle Streaming Responses from Different AI Models

The 5ire application abstracts streaming interactions with diverse LLM providers through a dual-layer architecture where provider-specific services construct streaming requests and specialized readers normalize the chunked responses into a uniform format for the UI.

The open-source 5ire project (nanbingxyz/5ire) implements a clean abstraction layer that allows users to switch between AI providers like OpenAI, Google Gemini, and Anthropic without changing the chat interface. This article examines how provider-specific chat services handle the technical complexity of streaming responses from different AI models, converting each provider's unique server-sent events format into a standardized message stream.

The Streaming Architecture in 5ire

Service Layer: Provider-Specific Request Construction

Each provider implements the IChatService interface defined in src/intellichat/services/IChatService.ts. The common contract includes a stream?: boolean flag defined in src/intellichat/types.ts (lines 24, 174, 224, 257) that controls whether the model returns a streaming response.

Provider-specific services set this flag according to their capabilities:

Reader Layer: Normalizing Stream Formats

After the HTTP response arrives, provider-specific readers parse the binary chunks into text. All readers extend BaseReader (src/intellichat/readers/BaseReader.ts), which owns a ReadableStreamDefaultReader<Uint8Array> (this.streamReader) and implements the generic readStream() algorithm (lines 170-284).

The base class handles low-level stream consumption: reading Uint8Array chunks, decoding them to strings, splitting by newlines, and invoking processChunk() for each line. Provider-specific subclasses override processChunk() to handle unique JSON schemas:

  • OpenAIReader extracts incremental content from choices[0].delta.content fields in OpenAI's SSE-style JSON lines.
  • GoogleReader maintains an internal buffer because Gemini may split JSON objects across network chunks; it parses only complete objects before emitting messages.
  • AnthropicReader parses Anthropic's type: "completion" JSON lines to extract text deltas.

Event Propagation to the UI

Processed message fragments emit through standard callbacks (onMessage, onError, onDone) that the chat store subscribes to. For NextChatService, which uses WebSocket communication rather than HTTP streaming, the service manually constructs a ReadableStream and forwards events via Electron IPC channels ('stream-data', 'stream-end', 'stream-error') as implemented in lines 193-241 of src/intellichat/services/NextChatService.ts.

The chat store (src/stores/useChatStore.ts) consumes the generic IChatService interface, remaining agnostic to the underlying provider. It receives standardized Message objects regardless of whether the source was OpenAI's SSE format or Gemini's chunked JSON.

Implementation Examples

The following examples demonstrate initializing streaming chat services, extending the architecture for custom providers, and consuming services through the chat store.

// Initialize OpenAI streaming chat
import { OpenAIChatService } from '@/intellichat/services/OpenAIChatService';
import { OpenAIReader } from '@/intellichat/readers/OpenAIReader';

const service = new OpenAIChatService({ model: 'gpt-4o', apiKey: '<YOUR_KEY>' });
const reader = new OpenAIReader(service.createRequest());

reader.onMessage = (msg) => console.log('▐', msg.content);
reader.onDone = () => console.log('\n--- done');
reader.onError = (e) => console.error('error', e);

reader.start();
// Adding a new provider (MyAI)
import { IChatService } from '@/intellichat/services/IChatService';
import { BaseReader } from '@/intellichat/readers/BaseReader';

class MyAIChatService implements IChatService {
  async request(messages: Message[]) {
    const payload = { messages, stream: true };
    const resp = await fetch('https://api.myai.com/v1/chat', {
      method: 'POST',
      headers: { Authorization: `Bearer ${this.token}` },
      body: JSON.stringify(payload),
    });
    return resp.body!.getReader();
  }
}

class MyAIReader extends BaseReader {
  protected async processChunk(chunk: string) {
    const data = JSON.parse(chunk);
    this.emitMessage({ role: 'assistant', content: data.content });
  }
}
// Consuming any service from the chat store
import { useChatStore } from '@/stores/useChatStore';
import { GoogleChatService } from '@/intellichat/services/GoogleChatService';
import { GoogleReader } from '@/intellichat/readers/GoogleReader';

const chatStore = useChatStore();
chatStore.startConversation(
  new GoogleChatService({ model: 'gemini-1.5-pro' }),
  new GoogleReader(...)
);

Summary

  • Dual-layer abstraction: Provider-specific services construct requests while dedicated readers parse responses, separating transport logic from UI consumption.
  • Standardized interfaces: The IChatService contract and BaseReader class allow the chat store to remain provider-agnostic.
  • Flexible streaming control: The stream boolean flag in src/intellichat/types.ts enables feature-flagged streaming for models that do or do not support real-time generation.
  • Extensible design: Adding support for new LLM providers requires only implementing a service class that sets payload.stream = true and a reader that overrides processChunk() for the provider's JSON format.

Frequently Asked Questions

How does 5ire handle providers that do not support streaming?

Services check model capabilities before setting the stream flag. For example, OpenAIChatService uses stream: !model.noStreaming to disable streaming for specific models, falling back to synchronous request-response cycles while maintaining the same interface.

What is the role of BaseReader in the streaming pipeline?

BaseReader (src/intellichat/readers/BaseReader.ts) manages the low-level binary stream consumption, decoding Uint8Array chunks and line-splitting logic (lines 170-284). Provider-specific subclasses only implement processChunk() to handle JSON parsing, eliminating duplicate stream management code across OpenAIReader, GoogleReader, and AnthropicReader.

How does NextChatService differ from HTTP-based services?

Unlike services that use standard HTTP SSE, NextChatService communicates via WebSocket and Electron IPC. It constructs a manual ReadableStream and emits events through IPC channels ('stream-data', 'stream-end', 'stream-error') as shown in lines 193-241 of src/intellichat/services/NextChatService.ts, yet still conforms to the IChatService interface consumed by the chat store.

Where is the streaming flag defined in the type system?

The optional stream?: boolean flag appears in the core type definitions at src/intellichat/types.ts on lines 24, 174, 224, and 257. This flag propagates through the service layer to control whether the backend returns chunked streaming responses or complete JSON objects.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →