# How Does the VoiceStudio Backend Handle Real-Time Audio Streaming?

> Discover how VoiceStudio's backend manages real-time audio streaming using WebSockets for low-latency transcriptions with ASR engines like Sherpa-ONNX and WhisperX.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: internals
- Published: 2026-09-11

---

**VoiceStudio handles real-time audio streaming through a dedicated WebSocket endpoint at `/v1/audio/transcriptions/stream`, using binary audio frames and bidirectional JSON events to deliver low-latency partial and final transcriptions via ASR engines like Sherpa-ONNX or WhisperX.**

The VoiceStudio backend implements a high-performance streaming architecture designed for live dictation and speech-to-text applications. Unlike traditional batch transcription APIs, the system maintains persistent WebSocket connections to process audio chunks as they arrive, yielding incremental results before the stream concludes. This design leverages the `SpeechCapabilities` discovery contract to communicate transport requirements and event schemas to clients.

## WebSocket Endpoint and Transport Contract

### Endpoint Registration

The streaming entry point is registered in **[`backend/api/routers/speech_platform.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/speech_platform.py)** at lines 96-100, where the route `/v1/audio/transcriptions/stream` is exposed as a WebSocket handler. The router advertises this capability through the `STREAM_PATH` constant, making it discoverable via the platform's capability negotiation protocol. This path serves as the single entry point for all real-time audio ingestion.

### Audio Format Specifications

Incoming WebSocket frames must conform to specific binary formats defined in the `StreamInputCapability` class within the same router file. The backend accepts two primary encodings:
- **`audio/webm;codecs=opus`** for compressed Opus streams
- **`audio/pcm;encoding=s16le;channels=1`** for raw 16-bit little-endian PCM mono audio

Clients must send binary frames containing actual audio data rather than base64-encoded strings, minimizing serialization overhead.

### Control Messages and Event Types

The transport contract requires clients to signal stream completion explicitly. When audio ingestion finishes, the client must transmit a JSON control message with the key **`type: "input_audio.end"`**. This triggers the ASR engine to finalize processing and emit the concluding transcription.

Outgoing events from the backend follow the `StreamOutputCapability` schema and include:
- **`session.started`**: Confirms successful session initialization
- **`partial`**: Contains interim transcription results as speech is detected
- **`final`**: Delivers completed utterances or summary transcripts
- **`error`**: Communicates processing failures or format violations

## Session Management and State Handling

When a client initiates a WebSocket connection, the backend creates a dedicated streaming session object managed in **[`backend/worker/transport/server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/worker/transport/server.py)**. The server enforces strict session hygiene by tracking the boolean flag `session.stream_open` to ensure only one active stream exists per session. This prevents race conditions where multiple simultaneous streams might corrupt the ASR state or exhaust processing resources.

The session lifecycle follows this sequence:
1. WebSocket handshake completes and session object initializes
2. `session.stream_open` is set to `True` upon first audio frame receipt
3. Binary chunks flow to the ASR engine while the server maintains connection state
4. Upon receiving `input_audio.end`, the server sets `session.stream_open` to `False` and begins graceful shutdown

## ASR Processing Pipeline

### Engine Selection and Routing

Audio chunks arriving via WebSocket are forwarded to the active ASR implementation based on configuration. The backend supports multiple transcription engines:
- **Sherpa-ONNX**: Implemented in **[`backend/services/sherpa_dictation.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/sherpa_dictation.py)** for on-device, low-latency dictation
- **WhisperX**: Accessed through **[`backend/services/asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py)** for higher-accuracy GPU-accelerated transcription

The [`asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/asr_backend.py) module acts as a routing layer that selects the appropriate engine based on the model specified in the session configuration, allowing seamless switching between engines without client-side changes.

### Real-Time Transcription Flow

As binary audio frames accumulate, the selected ASR engine processes them in windows and yields incremental results. The server wraps these raw outputs into standardized JSON events:
- **Partial results** are streamed immediately to provide live feedback (low latency)
- **Final results** are emitted after the `input_audio.end` signal triggers the engine's completion logic, with `kind` fields set to `"utterance"` or `"summary"` depending on the segmentation strategy

This architecture ensures clients receive usable text within milliseconds of speech occurrence while maintaining accuracy for completed phrases.

## Client Integration Examples

### OpenAI-Compatible Client

For applications using OpenAI SDKs, VoiceStudio exposes a compatible streaming interface. The client automatically negotiates the WebSocket upgrade and handles binary frame transmission:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:3900/v1", api_key="local")

with client.audio.transcriptions.with_streaming_response.create(
    model="whisper-1",
    file=open("speech.wav", "rb"),
    language="en",
) as response:
    for event in response.iter_events():
        if event.event == "partial":
            print("Partial:", event.text)
        elif event.event == "final":
            print("Final:", event.text)

```

### Raw WebSocket Implementation

For custom implementations, connect directly to the streaming endpoint using any WebSocket library:

```python
import asyncio
import json
import websockets

async def stream():
    async with websockets.connect("ws://localhost:3900/v1/audio/transcriptions/stream") as ws:
        # Send binary audio chunks (e.g., Opus)

        ws.send(b"...binary chunk 1...")
        ws.send(b"...binary chunk 2...")
        
        # Signal end of audio

        await ws.send(json.dumps({"type": "input_audio.end"}).encode())
        
        async for msg in ws:
            evt = json.loads(msg)
            print(evt)  # {'event': 'partial', 'text': '...'} etc.

asyncio.run(stream())

```

## Summary

- **WebSocket Transport**: The backend exposes `/v1/audio/transcriptions/stream` in [`speech_platform.py`](https://github.com/debpalash/VoiceStudio/blob/main/speech_platform.py) as the primary entry point for bidirectional streaming.
- **Binary Protocol**: Accepts raw Opus or PCM16 audio frames, rejecting base64 or text-encoded payloads to optimize throughput.
- **Session Safety**: The `session.stream_open` flag in [`worker/transport/server.py`](https://github.com/debpalash/VoiceStudio/blob/main/worker/transport/server.py) enforces single-stream-per-session semantics.
- **Engine Flexibility**: Routes audio through [`sherpa_dictation.py`](https://github.com/debpalash/VoiceStudio/blob/main/sherpa_dictation.py) or [`asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/asr_backend.py) to balance latency versus accuracy requirements.
- **Event-Driven**: Uses structured JSON events (`partial`, `final`, `error`) to communicate transcription state without polling overhead.

## Frequently Asked Questions

### What audio codecs does VoiceStudio support for real-time streaming?

VoiceStudio accepts two primary formats: `audio/webm;codecs=opus` for compressed streams and `audio/pcm;encoding=s16le;channels=1` for uncompressed 16-bit PCM mono audio. The codec requirement is advertised in the `StreamInputCapability` discovery payload, allowing clients to negotiate the optimal format before transmission.

### How does the backend prevent multiple simultaneous streams from a single client?

The session manager in [`backend/worker/transport/server.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/worker/transport/server.py) tracks a `session.stream_open` boolean flag. When a client initiates a stream, the flag is set to `True`, and any subsequent stream attempts are rejected until the current session receives the `input_audio.end` control message and properly closes. This prevents resource exhaustion and ASR state corruption.

### Can I use standard OpenAI client libraries with VoiceStudio's streaming endpoint?

Yes, VoiceStudio implements an OpenAI-compatible API surface. You can use `client.audio.transcriptions.with_streaming_response.create()` from the official OpenAI Python or Node.js clients by pointing the `base_url` to your VoiceStudio instance (e.g., `http://localhost:3900/v1`). The client handles WebSocket negotiation automatically while exposing the same partial/final event interface.

### What happens if the client disconnects without sending the end-of-audio signal?

If the WebSocket connection closes abruptly without the `input_audio.end` control message, the session manager detects the disconnect through the transport layer and triggers cleanup routines. Any buffered audio in the ASR pipeline is processed as a final utterance, and the `session.stream_open` flag is reset to allow new connections from the same session identifier.