Architectural Design of the Chat Completions API in y-gui: Multi-Provider Support Explained

The y-gui chat completions API implements a four-layer streaming architecture that decouples HTTP handling from provider-specific implementations via a factory pattern, enabling seamless multi-provider support through a unified BaseProvider interface.

The y-gui repository (luohy15/y-gui) provides a modular backend for AI-driven conversations built on Cloudflare Workers. Its chat completions endpoint (/api/chat-completions) handles real-time streaming, intent-based routing, and extensible provider integration through a clean separation of concerns.

Four-Layer Architecture Overview

The system organizes functionality into distinct layers that isolate transport logic from AI provider specifics:

Layer Responsibility Key File
HTTP Handler Parses requests, initializes SSE streams, wires services backend/src/api/chat-completions.ts
Service Layer Orchestrates chat state, intent analysis, and tool integration backend/src/serivce/chat.ts
Provider Factory Abstracts provider instantiation based on bot configuration backend/src/providers/provider-factory.ts
Provider Implementations Executes provider-specific API calls and response parsing backend/src/providers/openai-format-provider.ts

This stratification allows the system to support new LLM vendors without modifying the core chat logic.

Request-to-Response Flow

The architectural design of the chat completions API in y-gui follows an eight-stage pipeline that transforms HTTP requests into streaming SSE responses.

1. HTTP Handler & Stream Initialization

The handleChatCompletions function in backend/src/api/chat-completions.ts receives JSON payloads containing content, botName, chatId, and optional parameters. It immediately constructs a TransformStream to enable Server-Sent Events (SSE) streaming, ensuring low-latency UI updates for large responses.

2. Configuration Resolution

The handler retrieves chat history via ChatD1Repository and bot configuration via BotD1Repository. Missing API credentials automatically fallback to environment variables (OPENROUTER_FREE_KEY, OPENROUTER_BASE_URL), allowing zero-configuration deployments.

3. Provider Factory Instantiation

Multi-provider support materializes through ProviderFactory.createProvider(botConfig):

static createProvider(botConfig: BotConfig): BaseProvider {
  // Future: switch on botConfig.api_type or other flags
  return new OpenAIFormatProvider(botConfig);
}

Currently defaulting to OpenAIFormatProvider, the factory pattern enables one-line registration of alternative providers (e.g., native Anthropic or Gemini implementations) without touching consumer code.

4. Intent Analysis & MCP Integration

Before invoking the LLM, the system queries IntentAnalyzer to generate a RoutingDecision—determining whether to enable web-search, think-mode, or standard completion. Simultaneously, McpManager retrieves system prompts and tool definitions from McpServerD1Repository, preparing optional tool-calling capabilities.

5. Chat Service Execution

ChatService.initializeChat loads conversation state and assembles the system prompt. The service then calls ChatService.processUserMessage, which formats message history and delegates to the provider's callChatCompletions method.

6. Streaming Response Processing

OpenAIFormatProvider.callChatCompletions submits POST requests with stream:true and parses SSE chunks in real-time. The implementation extracts content, reasoning_content, and URL citations, yielding standardized ProviderResponseChunk objects:

export interface BaseProvider {
  callChatCompletions(
    messages: Message[],
    systemPrompt?: string,
    decision?: RoutingDecision
  ): AsyncGenerator<ProviderResponseChunk, void, unknown>;
}

7. Client Delivery

The HTTP handler consumes the async generator, writing each chunk to the TransformStream writer. When the provider signals [DONE], the stream closes gracefully.

8. Unified Error Propagation

Both the provider layer and service layer wrap exceptions into a consistent error object containing type, status, and message fields, streamed back to clients as SSE events for uniform front-end handling.

Multi-Provider Architecture Deep Dive

The y-gui multi-provider support relies on two critical abstractions that separate interface from implementation.

BaseProvider Interface Contract

Defined in backend/src/providers/provider-interface.ts, this contract mandates that all providers implement an async generator pattern for streaming:

export interface BaseProvider {
  callChatCompletions(
    messages: Message[],
    systemPrompt?: string,
    decision?: RoutingDecision
  ): AsyncGenerator<ProviderResponseChunk, void, unknown>;
}

This abstraction allows the service layer to remain agnostic to whether the underlying API uses OpenAI's format, Anthropic's Claude API, or Google Gemini's interface.

OpenAIFormatProvider Implementation

The concrete provider in backend/src/providers/openai-format-provider.ts handles OpenAI-compatible endpoints (including Anthropic Claude via OpenRouter). Key responsibilities include:

  • Message preparation: Injecting system prompts and cache_control headers for Claude models
  • Request building: Appending model suffixes (:online for web search), reasoning flags, and provider overrides
  • Stream parsing: Transforming SSE deltas into structured chunks containing content, reasoning traces, and citations

Extending the Factory

Adding a new provider requires implementing BaseProvider and registering it in the factory:

static createProvider(botConfig: BotConfig): BaseProvider {
  switch (botConfig.api_type) {
    case 'anthropic':
      return new ClaudeProvider(botConfig);
    case 'openai':
    default:
      return new OpenAIFormatProvider(botConfig);
  }
}

Update the bot configuration schema in shared/types/index.ts to include the api_type discriminator, and the system automatically routes requests to the appropriate implementation.

Implementation Examples

Consuming the Streaming API (Client-Side)

Clients interact with the architectural design of the chat completions API in y-gui through standard fetch requests and SSE parsing:

async function sendMessage(
  content: string,
  botName: string,
  chatId: string
) {
  const resp = await fetch('/api/chat-completions', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ content, botName, chatId })
  });

  const reader = resp.body?.getReader();
  const decoder = new TextDecoder();
  let buffer = '';

  while (reader) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const events = buffer.split('\n\n');
    
    for (let i = 0; i < events.length - 1; i++) {
      const line = events[i];
      if (line.startsWith('data:')) {
        const data = JSON.parse(line.replace('data:', '').trim());
        console.log('Chunk:', data);
        // Handle content, reasoning, or citations
      }
    }
    buffer = events[events.length - 1];
  }
}

Adding a Custom Provider

To extend multi-provider support with a native Anthropic implementation:

  1. Create backend/src/providers/claude-provider.ts implementing BaseProvider
  2. Register the provider in backend/src/providers/provider-factory.ts:
case 'anthropic':
  return new ClaudeProvider(botConfig);
  1. Add api_type: 'anthropic' to the bot configuration schema. The factory automatically instantiates the correct provider based on the bot's configuration flags.

Summary

  • Four-layer architecture separates HTTP transport, business logic, provider abstraction, and vendor-specific implementations in the y-gui chat completions API.
  • Provider factory pattern enables multi-provider support by decoupling bot configuration from concrete LLM client implementations.
  • Streaming SSE architecture delivers real-time responses through TransformStream and async generators, minimizing latency for large language model outputs.
  • Unified interfaces (BaseProvider, ProviderResponseChunk) ensure that adding new AI vendors requires changes only in the provider layer, leaving the chat service and HTTP handler untouched.
  • Intent analysis and MCP integration provide smart routing and tool execution without complicating the core request flow.

Frequently Asked Questions

How does y-gui handle multiple LLM providers without code duplication?

y-gui implements the provider factory pattern in backend/src/providers/provider-factory.ts, which instantiates concrete provider classes based on bot configuration. All providers implement the BaseProvider interface defined in provider-interface.ts, ensuring the service layer calls a unified callChatCompletions method regardless of whether the backend uses OpenAI, Anthropic, or other formats.

What enables real-time streaming in the y-gui chat completions API?

The API creates a TransformStream in backend/src/api/chat-completions.ts that converts the provider's async generator output into Server-Sent Events (SSE). The OpenAIFormatProvider parses incoming SSE chunks from the upstream LLM, extracts content and reasoning data, and yields standardized chunks that the HTTP handler writes immediately to the client stream.

Can I add support for a non-OpenAI-compatible LLM service?

Yes. Implement the BaseProvider interface in a new class (e.g., GeminiProvider), add a case to ProviderFactory.createProvider that returns your class based on a configuration flag like api_type, and update the bot schema in shared/types/index.ts. The existing chat service in backend/src/serivce/chat.ts will automatically use your implementation without further modifications.

How does the system decide when to use web search or reasoning modes?

The IntentAnalyzer evaluates incoming messages and returns a RoutingDecision that specifies whether to enable web-search, think-mode, or standard completion. This decision passes to the provider's callChatCompletions method, which adjusts API parameters (such as adding the :online model suffix for OpenRouter) without requiring changes to the client-side code or the core chat logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →