# How Bella OpenAPI Handles Streaming Responses for Chat Completions and TTS

> Bella OpenAPI expertly manages streaming responses for chat completions and TTS using SSE and chunked HTTP. Discover its unified callback architecture `HttpUtils.streamRequest` for real-time audio and text.

- Repository: [Ke Technologies/bella-openapi](https://github.com/lianjiatech/bella-openapi)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Bella OpenAPI delivers real-time chat completions via Server-Sent Events (SSE) and streams Text-to-Speech audio through chunked HTTP responses, using a unified callback architecture anchored by `HttpUtils.streamRequest` and protocol-specific adaptors.**

Bella OpenAPI supports real-time streaming for both conversational AI and voice synthesis workloads. According to the `lianjiatech/bella-openapi` source code, the platform routes streaming requests through protocol adaptors that bridge provider-specific APIs to standardized SSE or raw byte delivery mechanisms. This article explains the exact code paths in [`ChatController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/ChatController.java) and [`AudioController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/AudioController.java) that enable low-latency streaming for chat completions and TTS.

## Chat Completions Streaming Architecture

When a client sends a chat request with `stream: true`, Bella OpenAPI initiates a Server-Sent Events (SSE) connection that remains open for up to 30 minutes. The controller creates an `SseEmitter`, delegates the provider request to a `CompletionAdaptor`, and chains multiple callback processors to transform and safety-check each chunk before forwarding it to the client.

### SSE Implementation in ChatController

In [`api/server/src/main/java/com/ke/bella/openapi/endpoints/ChatController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/server/src/main/java/com/ke/bella/openapi/endpoints/ChatController.java), the `processCompletionRequest` method detects streaming mode through `request.isStream()`. When true, it instantiates an emitter via `SseHelper.createSse()` and invokes the adaptor's streaming method:

```java
if (request.isStream()) {
    SseEmitter sse = SseHelper.createSse(30 * 60 * 1000L, processData.getRequestId());
    
    adaptor.streamCompletion(
        request,
        ctx.url,
        property,
        StreamCallbackProvider.provide(sse, processData, EndpointContext.getApikey(), 
                                        logger, chatSafetyCheckService, property)
    );
    
    return sse;
}

```

The `SseEmitter` returned to Spring MVC handles the SSE protocol overhead, ensuring each chunk is formatted as `data: {...}\n\n`.

### The Callback Chain for Chat Streaming

The `StreamCallbackProvider.provide()` method in [`api/server/src/main/java/com/ke/bella/openapi/protocol/completion/callback/StreamCallbackProvider.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/server/src/main/java/com/ke/bella/openapi/protocol/completion/callback/StreamCallbackProvider.java) constructs a linked list of callbacks:

1. **Split reasoning callback** – separates reasoning content from main output
2. **Tool-call simulator** – handles function call streaming
3. **Merge reasoning callback** – recombines streams if needed
4. **StreamCompletionCallback** – final node that writes to the `SseEmitter`

Each callback implements `StreamCallback<StreamCompletionResponse>` and processes chunks asynchronously as they arrive from the provider.

## Text-to-Speech Streaming Architecture

For TTS streaming, Bella OpenAPI uses raw HTTP chunked transfer encoding rather than SSE. The `AudioController` starts an asynchronous servlet context and writes audio bytes directly to the response `OutputStream` as they arrive from the provider.

### Chunked HTTP Streaming in AudioController

In [`api/server/src/main/java/com/ke/bella/openapi/endpoints/AudioController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/server/src/main/java/com/ke/bella/openapi/endpoints/AudioController.java), the `speech` method checks `request.isStream()` and initializes an `AsyncContext` with a 20-minute timeout:

```java
if (request.isStream()) {
    AsyncContext async = httpRequest.startAsync();
    async.setTimeout(20 * 60 * 1000L);
    response.setContentType(getContentType(request.getResponseFormat()));
    OutputStream out = response.getOutputStream();
    
    ttsAdaptor.streamTts(
        request,
        ttsUrl,
        ttsProperty,
        ttsAdaptor.buildCallback(request, new StreamByteSender(async, out), 
                                processData, logger)
    );
    return;
}

```

The `StreamByteSender` wrapper manages the lifecycle of the async context, ensuring `async.complete()` is called when the stream terminates.

### TTS Callback and Byte Delivery

The `OpenAIAdaptor` for TTS (located in the TTS protocol package) implements `buildCallback()` to return an `OpenAIStreamTtsCallback`. This callback receives byte arrays from the provider and delegates to the `StreamByteSender`, which invokes `OutputStream.write(chunk)` for each audio frame. The client receives raw audio data (e.g., `audio/mpeg`) suitable for immediate playback or file storage.

## Core Streaming Infrastructure

Both chat and TTS streaming rely on a common HTTP utility that normalizes provider-specific streaming protocols into the Bella callback system.

### HttpUtils.streamRequest

The `streamRequest` method in [`api/spi/src/main/java/com/ke/bella/openapi/utils/HttpUtils.java`](https://github.com/lianjiatech/bella-openapi/blob/main/api/spi/src/main/java/com/ke/bella/openapi/utils/HttpUtils.java) opens a streaming HTTP connection to the AI provider and binds the response to a `BellaStreamCallback`. This utility handles:

- Connection pooling and timeouts
- Chunked transfer decoding
- Error propagation from the provider stream

Adaptor implementations such as `OpenAIAdaptor` (for chat) and the TTS variant call this utility after building their provider-specific `okhttp3.Request` objects.

### Protocol Adaptor Pattern

**`CompletionAdaptor`** and **`TtsAdaptor`** interfaces define the contract for streaming:

- `streamCompletion(request, url, property, callback)` – chat adaptors implement this to forward SSE-formatted provider responses
- `streamTts(request, url, property, callback)` – TTS adaptors implement this for binary audio streaming

The `OpenAIAdaptor` classes for both protocols demonstrate this pattern by translating Bella's internal request objects into OpenAI-compatible HTTP calls, then piping the response through `HttpUtils.streamRequest`.

## Practical Code Examples

### Client-Side Chat Streaming Request

```bash
curl -N -X POST https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Explain quantum computing"}],
    "stream": true
  }'

```

The `-N` flag disables buffering. Each SSE chunk arrives as:

```json
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","choices":[{"delta":{"content":"Quantum"}}]}

```

### Client-Side TTS Streaming Request

```bash
curl -X POST https://api.example.com/v1/audio/speech \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Hello from Bella OpenAPI",
    "response_format": "mp3",
    "stream": true
  }' \
  --output stream.mp3

```

The response body contains raw MP3 bytes streamed incrementally to `stream.mp3`.

## Summary

- **Chat streaming** uses **Server-Sent Events** via `SseEmitter` in [`ChatController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/ChatController.java), with a multi-stage callback chain provided by `StreamCallbackProvider` to handle reasoning, tool calls, and safety checks before client delivery.
- **TTS streaming** uses **HTTP chunked transfer** via `AsyncContext` in [`AudioController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/AudioController.java), writing raw audio bytes through `StreamByteSender` as they arrive from the provider.
- Both pathways delegate HTTP streaming mechanics to **`HttpUtils.streamRequest`**, ensuring consistent connection handling across different AI providers.
- Protocol adaptors like **`OpenAIAdaptor`** normalize provider-specific streaming formats into the Bella callback interface, enabling plug-and-play support for new AI backends.

## Frequently Asked Questions

### How does Bella OpenAPI choose between SSE and chunked HTTP for streaming?

The protocol is determined by the endpoint type. Chat completions in [`ChatController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/ChatController.java) always use **Server-Sent Events** because the data is text-based JSON chunks requiring event boundaries. TTS in [`AudioController.java`](https://github.com/lianjiatech/bella-openapi/blob/main/AudioController.java) uses **chunked HTTP** because it transmits binary audio bytes where SSE framing would be unnecessary overhead. Both methods rely on `HttpUtils.streamRequest` to manage the underlying HTTP connection.

### What is the purpose of the StreamCallbackProvider in chat streaming?

`StreamCallbackProvider.provide()` assembles a decorator chain of `StreamCallback` instances that process each chunk sequentially. This allows Bella OpenAPI to split reasoning content, simulate tool calls, merge streams, and run safety checks before the final `StreamCompletionCallback` writes to the `SseEmitter`. The pattern keeps the controller thin and delegates stream transformation to composable callback classes.

### Can custom adaptors implement streaming for other AI providers?

Yes. By implementing the `CompletionAdaptor` or `TtsAdaptor` interfaces and utilizing `HttpUtils.streamRequest`, developers can add streaming support for any provider that offers chunked HTTP responses. The adaptor only needs to map Bella's internal request format to the provider's API, then pass a provider-specific callback (usually extending `BellaStreamCallback`) to handle response chunking and error propagation.

### How does Bella OpenAPI handle streaming timeouts and errors?

For chat, `SseHelper.createSse()` configures a 30-minute timeout on the `SseEmitter`; for TTS, `AudioController` sets a 20-minute timeout on the `AsyncContext`. If the provider stream stalls or fails, `HttpUtils.streamRequest` propagates the exception through the callback chain, triggering the adaptor's error handling logic and closing the client connection gracefully.