How Does the VoiceStudio Backend Handle Real-Time Audio Streaming?

VoiceStudio handles real-time audio streaming through a dedicated WebSocket endpoint at /v1/audio/transcriptions/stream, using binary audio frames and bidirectional JSON events to deliver low-latency partial and final transcriptions via ASR engines like Sherpa-ONNX or WhisperX.

The VoiceStudio backend implements a high-performance streaming architecture designed for live dictation and speech-to-text applications. Unlike traditional batch transcription APIs, the system maintains persistent WebSocket connections to process audio chunks as they arrive, yielding incremental results before the stream concludes. This design leverages the SpeechCapabilities discovery contract to communicate transport requirements and event schemas to clients.

WebSocket Endpoint and Transport Contract

Endpoint Registration

The streaming entry point is registered in backend/api/routers/speech_platform.py at lines 96-100, where the route /v1/audio/transcriptions/stream is exposed as a WebSocket handler. The router advertises this capability through the STREAM_PATH constant, making it discoverable via the platform's capability negotiation protocol. This path serves as the single entry point for all real-time audio ingestion.

Audio Format Specifications

Incoming WebSocket frames must conform to specific binary formats defined in the StreamInputCapability class within the same router file. The backend accepts two primary encodings:

  • audio/webm;codecs=opus for compressed Opus streams
  • audio/pcm;encoding=s16le;channels=1 for raw 16-bit little-endian PCM mono audio

Clients must send binary frames containing actual audio data rather than base64-encoded strings, minimizing serialization overhead.

Control Messages and Event Types

The transport contract requires clients to signal stream completion explicitly. When audio ingestion finishes, the client must transmit a JSON control message with the key type: "input_audio.end". This triggers the ASR engine to finalize processing and emit the concluding transcription.

Outgoing events from the backend follow the StreamOutputCapability schema and include:

  • session.started: Confirms successful session initialization
  • partial: Contains interim transcription results as speech is detected
  • final: Delivers completed utterances or summary transcripts
  • error: Communicates processing failures or format violations

Session Management and State Handling

When a client initiates a WebSocket connection, the backend creates a dedicated streaming session object managed in backend/worker/transport/server.py. The server enforces strict session hygiene by tracking the boolean flag session.stream_open to ensure only one active stream exists per session. This prevents race conditions where multiple simultaneous streams might corrupt the ASR state or exhaust processing resources.

The session lifecycle follows this sequence:

  1. WebSocket handshake completes and session object initializes
  2. session.stream_open is set to True upon first audio frame receipt
  3. Binary chunks flow to the ASR engine while the server maintains connection state
  4. Upon receiving input_audio.end, the server sets session.stream_open to False and begins graceful shutdown

ASR Processing Pipeline

Engine Selection and Routing

Audio chunks arriving via WebSocket are forwarded to the active ASR implementation based on configuration. The backend supports multiple transcription engines:

The asr_backend.py module acts as a routing layer that selects the appropriate engine based on the model specified in the session configuration, allowing seamless switching between engines without client-side changes.

Real-Time Transcription Flow

As binary audio frames accumulate, the selected ASR engine processes them in windows and yields incremental results. The server wraps these raw outputs into standardized JSON events:

  • Partial results are streamed immediately to provide live feedback (low latency)
  • Final results are emitted after the input_audio.end signal triggers the engine's completion logic, with kind fields set to "utterance" or "summary" depending on the segmentation strategy

This architecture ensures clients receive usable text within milliseconds of speech occurrence while maintaining accuracy for completed phrases.

Client Integration Examples

OpenAI-Compatible Client

For applications using OpenAI SDKs, VoiceStudio exposes a compatible streaming interface. The client automatically negotiates the WebSocket upgrade and handles binary frame transmission:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:3900/v1", api_key="local")

with client.audio.transcriptions.with_streaming_response.create(
    model="whisper-1",
    file=open("speech.wav", "rb"),
    language="en",
) as response:
    for event in response.iter_events():
        if event.event == "partial":
            print("Partial:", event.text)
        elif event.event == "final":
            print("Final:", event.text)

Raw WebSocket Implementation

For custom implementations, connect directly to the streaming endpoint using any WebSocket library:

import asyncio
import json
import websockets

async def stream():
    async with websockets.connect("ws://localhost:3900/v1/audio/transcriptions/stream") as ws:
        # Send binary audio chunks (e.g., Opus)

        ws.send(b"...binary chunk 1...")
        ws.send(b"...binary chunk 2...")
        
        # Signal end of audio

        await ws.send(json.dumps({"type": "input_audio.end"}).encode())
        
        async for msg in ws:
            evt = json.loads(msg)
            print(evt)  # {'event': 'partial', 'text': '...'} etc.

asyncio.run(stream())

Summary

  • WebSocket Transport: The backend exposes /v1/audio/transcriptions/stream in speech_platform.py as the primary entry point for bidirectional streaming.
  • Binary Protocol: Accepts raw Opus or PCM16 audio frames, rejecting base64 or text-encoded payloads to optimize throughput.
  • Session Safety: The session.stream_open flag in worker/transport/server.py enforces single-stream-per-session semantics.
  • Engine Flexibility: Routes audio through sherpa_dictation.py or asr_backend.py to balance latency versus accuracy requirements.
  • Event-Driven: Uses structured JSON events (partial, final, error) to communicate transcription state without polling overhead.

Frequently Asked Questions

What audio codecs does VoiceStudio support for real-time streaming?

VoiceStudio accepts two primary formats: audio/webm;codecs=opus for compressed streams and audio/pcm;encoding=s16le;channels=1 for uncompressed 16-bit PCM mono audio. The codec requirement is advertised in the StreamInputCapability discovery payload, allowing clients to negotiate the optimal format before transmission.

How does the backend prevent multiple simultaneous streams from a single client?

The session manager in backend/worker/transport/server.py tracks a session.stream_open boolean flag. When a client initiates a stream, the flag is set to True, and any subsequent stream attempts are rejected until the current session receives the input_audio.end control message and properly closes. This prevents resource exhaustion and ASR state corruption.

Can I use standard OpenAI client libraries with VoiceStudio's streaming endpoint?

Yes, VoiceStudio implements an OpenAI-compatible API surface. You can use client.audio.transcriptions.with_streaming_response.create() from the official OpenAI Python or Node.js clients by pointing the base_url to your VoiceStudio instance (e.g., http://localhost:3900/v1). The client handles WebSocket negotiation automatically while exposing the same partial/final event interface.

What happens if the client disconnects without sending the end-of-audio signal?

If the WebSocket connection closes abruptly without the input_audio.end control message, the session manager detects the disconnect through the transport layer and triggers cleanup routines. Any buffered audio in the ASR pipeline is processed as a final utterance, and the session.stream_open flag is reset to allow new connections from the same session identifier.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →