How Bella OpenAPI Handles Streaming Responses for Chat Completions and TTS
Bella OpenAPI delivers real-time chat completions via Server-Sent Events (SSE) and streams Text-to-Speech audio through chunked HTTP responses, using a unified callback architecture anchored by HttpUtils.streamRequest and protocol-specific adaptors.
Bella OpenAPI supports real-time streaming for both conversational AI and voice synthesis workloads. According to the lianjiatech/bella-openapi source code, the platform routes streaming requests through protocol adaptors that bridge provider-specific APIs to standardized SSE or raw byte delivery mechanisms. This article explains the exact code paths in ChatController.java and AudioController.java that enable low-latency streaming for chat completions and TTS.
Chat Completions Streaming Architecture
When a client sends a chat request with stream: true, Bella OpenAPI initiates a Server-Sent Events (SSE) connection that remains open for up to 30 minutes. The controller creates an SseEmitter, delegates the provider request to a CompletionAdaptor, and chains multiple callback processors to transform and safety-check each chunk before forwarding it to the client.
SSE Implementation in ChatController
In api/server/src/main/java/com/ke/bella/openapi/endpoints/ChatController.java, the processCompletionRequest method detects streaming mode through request.isStream(). When true, it instantiates an emitter via SseHelper.createSse() and invokes the adaptor's streaming method:
if (request.isStream()) {
SseEmitter sse = SseHelper.createSse(30 * 60 * 1000L, processData.getRequestId());
adaptor.streamCompletion(
request,
ctx.url,
property,
StreamCallbackProvider.provide(sse, processData, EndpointContext.getApikey(),
logger, chatSafetyCheckService, property)
);
return sse;
}
The SseEmitter returned to Spring MVC handles the SSE protocol overhead, ensuring each chunk is formatted as data: {...}\n\n.
The Callback Chain for Chat Streaming
The StreamCallbackProvider.provide() method in api/server/src/main/java/com/ke/bella/openapi/protocol/completion/callback/StreamCallbackProvider.java constructs a linked list of callbacks:
- Split reasoning callback – separates reasoning content from main output
- Tool-call simulator – handles function call streaming
- Merge reasoning callback – recombines streams if needed
- StreamCompletionCallback – final node that writes to the
SseEmitter
Each callback implements StreamCallback<StreamCompletionResponse> and processes chunks asynchronously as they arrive from the provider.
Text-to-Speech Streaming Architecture
For TTS streaming, Bella OpenAPI uses raw HTTP chunked transfer encoding rather than SSE. The AudioController starts an asynchronous servlet context and writes audio bytes directly to the response OutputStream as they arrive from the provider.
Chunked HTTP Streaming in AudioController
In api/server/src/main/java/com/ke/bella/openapi/endpoints/AudioController.java, the speech method checks request.isStream() and initializes an AsyncContext with a 20-minute timeout:
if (request.isStream()) {
AsyncContext async = httpRequest.startAsync();
async.setTimeout(20 * 60 * 1000L);
response.setContentType(getContentType(request.getResponseFormat()));
OutputStream out = response.getOutputStream();
ttsAdaptor.streamTts(
request,
ttsUrl,
ttsProperty,
ttsAdaptor.buildCallback(request, new StreamByteSender(async, out),
processData, logger)
);
return;
}
The StreamByteSender wrapper manages the lifecycle of the async context, ensuring async.complete() is called when the stream terminates.
TTS Callback and Byte Delivery
The OpenAIAdaptor for TTS (located in the TTS protocol package) implements buildCallback() to return an OpenAIStreamTtsCallback. This callback receives byte arrays from the provider and delegates to the StreamByteSender, which invokes OutputStream.write(chunk) for each audio frame. The client receives raw audio data (e.g., audio/mpeg) suitable for immediate playback or file storage.
Core Streaming Infrastructure
Both chat and TTS streaming rely on a common HTTP utility that normalizes provider-specific streaming protocols into the Bella callback system.
HttpUtils.streamRequest
The streamRequest method in api/spi/src/main/java/com/ke/bella/openapi/utils/HttpUtils.java opens a streaming HTTP connection to the AI provider and binds the response to a BellaStreamCallback. This utility handles:
- Connection pooling and timeouts
- Chunked transfer decoding
- Error propagation from the provider stream
Adaptor implementations such as OpenAIAdaptor (for chat) and the TTS variant call this utility after building their provider-specific okhttp3.Request objects.
Protocol Adaptor Pattern
CompletionAdaptor and TtsAdaptor interfaces define the contract for streaming:
streamCompletion(request, url, property, callback)– chat adaptors implement this to forward SSE-formatted provider responsesstreamTts(request, url, property, callback)– TTS adaptors implement this for binary audio streaming
The OpenAIAdaptor classes for both protocols demonstrate this pattern by translating Bella's internal request objects into OpenAI-compatible HTTP calls, then piping the response through HttpUtils.streamRequest.
Practical Code Examples
Client-Side Chat Streaming Request
curl -N -X POST https://api.example.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Explain quantum computing"}],
"stream": true
}'
The -N flag disables buffering. Each SSE chunk arrives as:
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","choices":[{"delta":{"content":"Quantum"}}]}
Client-Side TTS Streaming Request
curl -X POST https://api.example.com/v1/audio/speech \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "Hello from Bella OpenAPI",
"response_format": "mp3",
"stream": true
}' \
--output stream.mp3
The response body contains raw MP3 bytes streamed incrementally to stream.mp3.
Summary
- Chat streaming uses Server-Sent Events via
SseEmitterinChatController.java, with a multi-stage callback chain provided byStreamCallbackProviderto handle reasoning, tool calls, and safety checks before client delivery. - TTS streaming uses HTTP chunked transfer via
AsyncContextinAudioController.java, writing raw audio bytes throughStreamByteSenderas they arrive from the provider. - Both pathways delegate HTTP streaming mechanics to
HttpUtils.streamRequest, ensuring consistent connection handling across different AI providers. - Protocol adaptors like
OpenAIAdaptornormalize provider-specific streaming formats into the Bella callback interface, enabling plug-and-play support for new AI backends.
Frequently Asked Questions
How does Bella OpenAPI choose between SSE and chunked HTTP for streaming?
The protocol is determined by the endpoint type. Chat completions in ChatController.java always use Server-Sent Events because the data is text-based JSON chunks requiring event boundaries. TTS in AudioController.java uses chunked HTTP because it transmits binary audio bytes where SSE framing would be unnecessary overhead. Both methods rely on HttpUtils.streamRequest to manage the underlying HTTP connection.
What is the purpose of the StreamCallbackProvider in chat streaming?
StreamCallbackProvider.provide() assembles a decorator chain of StreamCallback instances that process each chunk sequentially. This allows Bella OpenAPI to split reasoning content, simulate tool calls, merge streams, and run safety checks before the final StreamCompletionCallback writes to the SseEmitter. The pattern keeps the controller thin and delegates stream transformation to composable callback classes.
Can custom adaptors implement streaming for other AI providers?
Yes. By implementing the CompletionAdaptor or TtsAdaptor interfaces and utilizing HttpUtils.streamRequest, developers can add streaming support for any provider that offers chunked HTTP responses. The adaptor only needs to map Bella's internal request format to the provider's API, then pass a provider-specific callback (usually extending BellaStreamCallback) to handle response chunking and error propagation.
How does Bella OpenAPI handle streaming timeouts and errors?
For chat, SseHelper.createSse() configures a 30-minute timeout on the SseEmitter; for TTS, AudioController sets a 20-minute timeout on the AsyncContext. If the provider stream stalls or fails, HttpUtils.streamRequest propagates the exception through the callback chain, triggering the adaptor's error handling logic and closing the client connection gracefully.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →