How Bella OpenAPI Handles Streaming Responses for Chat Completions and TTS

Bella OpenAPI delivers real-time chat completions via Server-Sent Events (SSE) and streams Text-to-Speech audio through chunked HTTP responses, using a unified callback architecture anchored by HttpUtils.streamRequest and protocol-specific adaptors.

Bella OpenAPI supports real-time streaming for both conversational AI and voice synthesis workloads. According to the lianjiatech/bella-openapi source code, the platform routes streaming requests through protocol adaptors that bridge provider-specific APIs to standardized SSE or raw byte delivery mechanisms. This article explains the exact code paths in ChatController.java and AudioController.java that enable low-latency streaming for chat completions and TTS.

Chat Completions Streaming Architecture

When a client sends a chat request with stream: true, Bella OpenAPI initiates a Server-Sent Events (SSE) connection that remains open for up to 30 minutes. The controller creates an SseEmitter, delegates the provider request to a CompletionAdaptor, and chains multiple callback processors to transform and safety-check each chunk before forwarding it to the client.

SSE Implementation in ChatController

In api/server/src/main/java/com/ke/bella/openapi/endpoints/ChatController.java, the processCompletionRequest method detects streaming mode through request.isStream(). When true, it instantiates an emitter via SseHelper.createSse() and invokes the adaptor's streaming method:

if (request.isStream()) {
    SseEmitter sse = SseHelper.createSse(30 * 60 * 1000L, processData.getRequestId());
    
    adaptor.streamCompletion(
        request,
        ctx.url,
        property,
        StreamCallbackProvider.provide(sse, processData, EndpointContext.getApikey(), 
                                        logger, chatSafetyCheckService, property)
    );
    
    return sse;
}

The SseEmitter returned to Spring MVC handles the SSE protocol overhead, ensuring each chunk is formatted as data: {...}\n\n.

The Callback Chain for Chat Streaming

The StreamCallbackProvider.provide() method in api/server/src/main/java/com/ke/bella/openapi/protocol/completion/callback/StreamCallbackProvider.java constructs a linked list of callbacks:

  1. Split reasoning callback – separates reasoning content from main output
  2. Tool-call simulator – handles function call streaming
  3. Merge reasoning callback – recombines streams if needed
  4. StreamCompletionCallback – final node that writes to the SseEmitter

Each callback implements StreamCallback<StreamCompletionResponse> and processes chunks asynchronously as they arrive from the provider.

Text-to-Speech Streaming Architecture

For TTS streaming, Bella OpenAPI uses raw HTTP chunked transfer encoding rather than SSE. The AudioController starts an asynchronous servlet context and writes audio bytes directly to the response OutputStream as they arrive from the provider.

Chunked HTTP Streaming in AudioController

In api/server/src/main/java/com/ke/bella/openapi/endpoints/AudioController.java, the speech method checks request.isStream() and initializes an AsyncContext with a 20-minute timeout:

if (request.isStream()) {
    AsyncContext async = httpRequest.startAsync();
    async.setTimeout(20 * 60 * 1000L);
    response.setContentType(getContentType(request.getResponseFormat()));
    OutputStream out = response.getOutputStream();
    
    ttsAdaptor.streamTts(
        request,
        ttsUrl,
        ttsProperty,
        ttsAdaptor.buildCallback(request, new StreamByteSender(async, out), 
                                processData, logger)
    );
    return;
}

The StreamByteSender wrapper manages the lifecycle of the async context, ensuring async.complete() is called when the stream terminates.

TTS Callback and Byte Delivery

The OpenAIAdaptor for TTS (located in the TTS protocol package) implements buildCallback() to return an OpenAIStreamTtsCallback. This callback receives byte arrays from the provider and delegates to the StreamByteSender, which invokes OutputStream.write(chunk) for each audio frame. The client receives raw audio data (e.g., audio/mpeg) suitable for immediate playback or file storage.

Core Streaming Infrastructure

Both chat and TTS streaming rely on a common HTTP utility that normalizes provider-specific streaming protocols into the Bella callback system.

HttpUtils.streamRequest

The streamRequest method in api/spi/src/main/java/com/ke/bella/openapi/utils/HttpUtils.java opens a streaming HTTP connection to the AI provider and binds the response to a BellaStreamCallback. This utility handles:

  • Connection pooling and timeouts
  • Chunked transfer decoding
  • Error propagation from the provider stream

Adaptor implementations such as OpenAIAdaptor (for chat) and the TTS variant call this utility after building their provider-specific okhttp3.Request objects.

Protocol Adaptor Pattern

CompletionAdaptor and TtsAdaptor interfaces define the contract for streaming:

  • streamCompletion(request, url, property, callback) – chat adaptors implement this to forward SSE-formatted provider responses
  • streamTts(request, url, property, callback) – TTS adaptors implement this for binary audio streaming

The OpenAIAdaptor classes for both protocols demonstrate this pattern by translating Bella's internal request objects into OpenAI-compatible HTTP calls, then piping the response through HttpUtils.streamRequest.

Practical Code Examples

Client-Side Chat Streaming Request

curl -N -X POST https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Explain quantum computing"}],
    "stream": true
  }'

The -N flag disables buffering. Each SSE chunk arrives as:

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","choices":[{"delta":{"content":"Quantum"}}]}

Client-Side TTS Streaming Request

curl -X POST https://api.example.com/v1/audio/speech \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Hello from Bella OpenAPI",
    "response_format": "mp3",
    "stream": true
  }' \
  --output stream.mp3

The response body contains raw MP3 bytes streamed incrementally to stream.mp3.

Summary

  • Chat streaming uses Server-Sent Events via SseEmitter in ChatController.java, with a multi-stage callback chain provided by StreamCallbackProvider to handle reasoning, tool calls, and safety checks before client delivery.
  • TTS streaming uses HTTP chunked transfer via AsyncContext in AudioController.java, writing raw audio bytes through StreamByteSender as they arrive from the provider.
  • Both pathways delegate HTTP streaming mechanics to HttpUtils.streamRequest, ensuring consistent connection handling across different AI providers.
  • Protocol adaptors like OpenAIAdaptor normalize provider-specific streaming formats into the Bella callback interface, enabling plug-and-play support for new AI backends.

Frequently Asked Questions

How does Bella OpenAPI choose between SSE and chunked HTTP for streaming?

The protocol is determined by the endpoint type. Chat completions in ChatController.java always use Server-Sent Events because the data is text-based JSON chunks requiring event boundaries. TTS in AudioController.java uses chunked HTTP because it transmits binary audio bytes where SSE framing would be unnecessary overhead. Both methods rely on HttpUtils.streamRequest to manage the underlying HTTP connection.

What is the purpose of the StreamCallbackProvider in chat streaming?

StreamCallbackProvider.provide() assembles a decorator chain of StreamCallback instances that process each chunk sequentially. This allows Bella OpenAPI to split reasoning content, simulate tool calls, merge streams, and run safety checks before the final StreamCompletionCallback writes to the SseEmitter. The pattern keeps the controller thin and delegates stream transformation to composable callback classes.

Can custom adaptors implement streaming for other AI providers?

Yes. By implementing the CompletionAdaptor or TtsAdaptor interfaces and utilizing HttpUtils.streamRequest, developers can add streaming support for any provider that offers chunked HTTP responses. The adaptor only needs to map Bella's internal request format to the provider's API, then pass a provider-specific callback (usually extending BellaStreamCallback) to handle response chunking and error propagation.

How does Bella OpenAPI handle streaming timeouts and errors?

For chat, SseHelper.createSse() configures a 30-minute timeout on the SseEmitter; for TTS, AudioController sets a 20-minute timeout on the AsyncContext. If the provider stream stalls or fails, HttpUtils.streamRequest propagates the exception through the callback chain, triggering the adaptor's error handling logic and closing the client connection gracefully.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →