# Architectural Design of the Chat Completions API in y-gui: Multi-Provider Support Explained

> Explore the y-gui chat completions API architecture. Learn how its four-layer streaming design and factory pattern enable multi-provider support through a unified BaseProvider interface.

- Repository: [luohy15/y-gui](https://github.com/luohy15/y-gui)
- Tags: architecture
- Published: 2026-03-06

---

**The y-gui chat completions API implements a four-layer streaming architecture that decouples HTTP handling from provider-specific implementations via a factory pattern, enabling seamless multi-provider support through a unified `BaseProvider` interface.**

The **y-gui** repository (`luohy15/y-gui`) provides a modular backend for AI-driven conversations built on Cloudflare Workers. Its chat completions endpoint (`/api/chat-completions`) handles real-time streaming, intent-based routing, and extensible provider integration through a clean separation of concerns.

## Four-Layer Architecture Overview

The system organizes functionality into distinct layers that isolate transport logic from AI provider specifics:

| Layer | Responsibility | Key File |
|-------|----------------|----------|
| **HTTP Handler** | Parses requests, initializes SSE streams, wires services | [`backend/src/api/chat-completions.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/api/chat-completions.ts) |
| **Service Layer** | Orchestrates chat state, intent analysis, and tool integration | [`backend/src/serivce/chat.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/serivce/chat.ts) |
| **Provider Factory** | Abstracts provider instantiation based on bot configuration | [`backend/src/providers/provider-factory.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/providers/provider-factory.ts) |
| **Provider Implementations** | Executes provider-specific API calls and response parsing | [`backend/src/providers/openai-format-provider.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/providers/openai-format-provider.ts) |

This stratification allows the system to support new LLM vendors without modifying the core chat logic.

## Request-to-Response Flow

The architectural design of the chat completions API in y-gui follows an eight-stage pipeline that transforms HTTP requests into streaming SSE responses.

### 1. HTTP Handler & Stream Initialization

The `handleChatCompletions` function in [`backend/src/api/chat-completions.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/api/chat-completions.ts) receives JSON payloads containing `content`, `botName`, `chatId`, and optional parameters. It immediately constructs a **TransformStream** to enable Server-Sent Events (SSE) streaming, ensuring low-latency UI updates for large responses.

### 2. Configuration Resolution

The handler retrieves chat history via `ChatD1Repository` and bot configuration via `BotD1Repository`. Missing API credentials automatically fallback to environment variables (`OPENROUTER_FREE_KEY`, `OPENROUTER_BASE_URL`), allowing zero-configuration deployments.

### 3. Provider Factory Instantiation

Multi-provider support materializes through `ProviderFactory.createProvider(botConfig)`:

```typescript
static createProvider(botConfig: BotConfig): BaseProvider {
  // Future: switch on botConfig.api_type or other flags
  return new OpenAIFormatProvider(botConfig);
}

```

Currently defaulting to `OpenAIFormatProvider`, the factory pattern enables one-line registration of alternative providers (e.g., native Anthropic or Gemini implementations) without touching consumer code.

### 4. Intent Analysis & MCP Integration

Before invoking the LLM, the system queries `IntentAnalyzer` to generate a `RoutingDecision`—determining whether to enable web-search, think-mode, or standard completion. Simultaneously, `McpManager` retrieves system prompts and tool definitions from `McpServerD1Repository`, preparing optional tool-calling capabilities.

### 5. Chat Service Execution

`ChatService.initializeChat` loads conversation state and assembles the system prompt. The service then calls `ChatService.processUserMessage`, which formats message history and delegates to the provider's `callChatCompletions` method.

### 6. Streaming Response Processing

`OpenAIFormatProvider.callChatCompletions` submits POST requests with `stream:true` and parses SSE chunks in real-time. The implementation extracts `content`, `reasoning_content`, and URL citations, yielding standardized `ProviderResponseChunk` objects:

```typescript
export interface BaseProvider {
  callChatCompletions(
    messages: Message[],
    systemPrompt?: string,
    decision?: RoutingDecision
  ): AsyncGenerator<ProviderResponseChunk, void, unknown>;
}

```

### 7. Client Delivery

The HTTP handler consumes the async generator, writing each chunk to the TransformStream writer. When the provider signals `[DONE]`, the stream closes gracefully.

### 8. Unified Error Propagation

Both the provider layer and service layer wrap exceptions into a consistent error object containing `type`, `status`, and `message` fields, streamed back to clients as SSE events for uniform front-end handling.

## Multi-Provider Architecture Deep Dive

The y-gui multi-provider support relies on two critical abstractions that separate interface from implementation.

### BaseProvider Interface Contract

Defined in [`backend/src/providers/provider-interface.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/providers/provider-interface.ts), this contract mandates that all providers implement an async generator pattern for streaming:

```typescript
export interface BaseProvider {
  callChatCompletions(
    messages: Message[],
    systemPrompt?: string,
    decision?: RoutingDecision
  ): AsyncGenerator<ProviderResponseChunk, void, unknown>;
}

```

This abstraction allows the service layer to remain agnostic to whether the underlying API uses OpenAI's format, Anthropic's Claude API, or Google Gemini's interface.

### OpenAIFormatProvider Implementation

The concrete provider in [`backend/src/providers/openai-format-provider.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/providers/openai-format-provider.ts) handles OpenAI-compatible endpoints (including Anthropic Claude via OpenRouter). Key responsibilities include:

- **Message preparation**: Injecting system prompts and `cache_control` headers for Claude models
- **Request building**: Appending model suffixes (`:online` for web search), reasoning flags, and provider overrides
- **Stream parsing**: Transforming SSE deltas into structured chunks containing content, reasoning traces, and citations

### Extending the Factory

Adding a new provider requires implementing `BaseProvider` and registering it in the factory:

```typescript
static createProvider(botConfig: BotConfig): BaseProvider {
  switch (botConfig.api_type) {
    case 'anthropic':
      return new ClaudeProvider(botConfig);
    case 'openai':
    default:
      return new OpenAIFormatProvider(botConfig);
  }
}

```

Update the bot configuration schema in [`shared/types/index.ts`](https://github.com/luohy15/y-gui/blob/main/shared/types/index.ts) to include the `api_type` discriminator, and the system automatically routes requests to the appropriate implementation.

## Implementation Examples

### Consuming the Streaming API (Client-Side)

Clients interact with the architectural design of the chat completions API in y-gui through standard fetch requests and SSE parsing:

```typescript
async function sendMessage(
  content: string,
  botName: string,
  chatId: string
) {
  const resp = await fetch('/api/chat-completions', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ content, botName, chatId })
  });

  const reader = resp.body?.getReader();
  const decoder = new TextDecoder();
  let buffer = '';

  while (reader) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const events = buffer.split('\n\n');
    
    for (let i = 0; i < events.length - 1; i++) {
      const line = events[i];
      if (line.startsWith('data:')) {
        const data = JSON.parse(line.replace('data:', '').trim());
        console.log('Chunk:', data);
        // Handle content, reasoning, or citations
      }
    }
    buffer = events[events.length - 1];
  }
}

```

### Adding a Custom Provider

To extend multi-provider support with a native Anthropic implementation:

1. Create [`backend/src/providers/claude-provider.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/providers/claude-provider.ts) implementing `BaseProvider`
2. Register the provider in [`backend/src/providers/provider-factory.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/providers/provider-factory.ts):

```typescript
case 'anthropic':
  return new ClaudeProvider(botConfig);

```

3. Add `api_type: 'anthropic'` to the bot configuration schema. The factory automatically instantiates the correct provider based on the bot's configuration flags.

## Summary

- **Four-layer architecture** separates HTTP transport, business logic, provider abstraction, and vendor-specific implementations in the y-gui chat completions API.
- **Provider factory pattern** enables multi-provider support by decoupling bot configuration from concrete LLM client implementations.
- **Streaming SSE architecture** delivers real-time responses through `TransformStream` and async generators, minimizing latency for large language model outputs.
- **Unified interfaces** (`BaseProvider`, `ProviderResponseChunk`) ensure that adding new AI vendors requires changes only in the provider layer, leaving the chat service and HTTP handler untouched.
- **Intent analysis and MCP integration** provide smart routing and tool execution without complicating the core request flow.

## Frequently Asked Questions

### How does y-gui handle multiple LLM providers without code duplication?

y-gui implements the **provider factory pattern** in [`backend/src/providers/provider-factory.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/providers/provider-factory.ts), which instantiates concrete provider classes based on bot configuration. All providers implement the `BaseProvider` interface defined in [`provider-interface.ts`](https://github.com/luohy15/y-gui/blob/main/provider-interface.ts), ensuring the service layer calls a unified `callChatCompletions` method regardless of whether the backend uses OpenAI, Anthropic, or other formats.

### What enables real-time streaming in the y-gui chat completions API?

The API creates a **TransformStream** in [`backend/src/api/chat-completions.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/api/chat-completions.ts) that converts the provider's async generator output into Server-Sent Events (SSE). The `OpenAIFormatProvider` parses incoming SSE chunks from the upstream LLM, extracts content and reasoning data, and yields standardized chunks that the HTTP handler writes immediately to the client stream.

### Can I add support for a non-OpenAI-compatible LLM service?

Yes. Implement the `BaseProvider` interface in a new class (e.g., `GeminiProvider`), add a case to `ProviderFactory.createProvider` that returns your class based on a configuration flag like `api_type`, and update the bot schema in [`shared/types/index.ts`](https://github.com/luohy15/y-gui/blob/main/shared/types/index.ts). The existing chat service in [`backend/src/serivce/chat.ts`](https://github.com/luohy15/y-gui/blob/main/backend/src/serivce/chat.ts) will automatically use your implementation without further modifications.

### How does the system decide when to use web search or reasoning modes?

The `IntentAnalyzer` evaluates incoming messages and returns a `RoutingDecision` that specifies whether to enable web-search, think-mode, or standard completion. This decision passes to the provider's `callChatCompletions` method, which adjusts API parameters (such as adding the `:online` model suffix for OpenRouter) without requiring changes to the client-side code or the core chat logic.