How Does the VoiceStudio Backend Handle Real-Time Audio Streaming?
VoiceStudio handles real-time audio streaming through a dedicated WebSocket endpoint at /v1/audio/transcriptions/stream, using binary audio frames and bidirectional JSON events to deliver low-latency partial and final transcriptions via ASR engines like Sherpa-ONNX or WhisperX.
The VoiceStudio backend implements a high-performance streaming architecture designed for live dictation and speech-to-text applications. Unlike traditional batch transcription APIs, the system maintains persistent WebSocket connections to process audio chunks as they arrive, yielding incremental results before the stream concludes. This design leverages the SpeechCapabilities discovery contract to communicate transport requirements and event schemas to clients.
WebSocket Endpoint and Transport Contract
Endpoint Registration
The streaming entry point is registered in backend/api/routers/speech_platform.py at lines 96-100, where the route /v1/audio/transcriptions/stream is exposed as a WebSocket handler. The router advertises this capability through the STREAM_PATH constant, making it discoverable via the platform's capability negotiation protocol. This path serves as the single entry point for all real-time audio ingestion.
Audio Format Specifications
Incoming WebSocket frames must conform to specific binary formats defined in the StreamInputCapability class within the same router file. The backend accepts two primary encodings:
audio/webm;codecs=opusfor compressed Opus streamsaudio/pcm;encoding=s16le;channels=1for raw 16-bit little-endian PCM mono audio
Clients must send binary frames containing actual audio data rather than base64-encoded strings, minimizing serialization overhead.
Control Messages and Event Types
The transport contract requires clients to signal stream completion explicitly. When audio ingestion finishes, the client must transmit a JSON control message with the key type: "input_audio.end". This triggers the ASR engine to finalize processing and emit the concluding transcription.
Outgoing events from the backend follow the StreamOutputCapability schema and include:
session.started: Confirms successful session initializationpartial: Contains interim transcription results as speech is detectedfinal: Delivers completed utterances or summary transcriptserror: Communicates processing failures or format violations
Session Management and State Handling
When a client initiates a WebSocket connection, the backend creates a dedicated streaming session object managed in backend/worker/transport/server.py. The server enforces strict session hygiene by tracking the boolean flag session.stream_open to ensure only one active stream exists per session. This prevents race conditions where multiple simultaneous streams might corrupt the ASR state or exhaust processing resources.
The session lifecycle follows this sequence:
- WebSocket handshake completes and session object initializes
session.stream_openis set toTrueupon first audio frame receipt- Binary chunks flow to the ASR engine while the server maintains connection state
- Upon receiving
input_audio.end, the server setssession.stream_opentoFalseand begins graceful shutdown
ASR Processing Pipeline
Engine Selection and Routing
Audio chunks arriving via WebSocket are forwarded to the active ASR implementation based on configuration. The backend supports multiple transcription engines:
- Sherpa-ONNX: Implemented in
backend/services/sherpa_dictation.pyfor on-device, low-latency dictation - WhisperX: Accessed through
backend/services/asr_backend.pyfor higher-accuracy GPU-accelerated transcription
The asr_backend.py module acts as a routing layer that selects the appropriate engine based on the model specified in the session configuration, allowing seamless switching between engines without client-side changes.
Real-Time Transcription Flow
As binary audio frames accumulate, the selected ASR engine processes them in windows and yields incremental results. The server wraps these raw outputs into standardized JSON events:
- Partial results are streamed immediately to provide live feedback (low latency)
- Final results are emitted after the
input_audio.endsignal triggers the engine's completion logic, withkindfields set to"utterance"or"summary"depending on the segmentation strategy
This architecture ensures clients receive usable text within milliseconds of speech occurrence while maintaining accuracy for completed phrases.
Client Integration Examples
OpenAI-Compatible Client
For applications using OpenAI SDKs, VoiceStudio exposes a compatible streaming interface. The client automatically negotiates the WebSocket upgrade and handles binary frame transmission:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:3900/v1", api_key="local")
with client.audio.transcriptions.with_streaming_response.create(
model="whisper-1",
file=open("speech.wav", "rb"),
language="en",
) as response:
for event in response.iter_events():
if event.event == "partial":
print("Partial:", event.text)
elif event.event == "final":
print("Final:", event.text)
Raw WebSocket Implementation
For custom implementations, connect directly to the streaming endpoint using any WebSocket library:
import asyncio
import json
import websockets
async def stream():
async with websockets.connect("ws://localhost:3900/v1/audio/transcriptions/stream") as ws:
# Send binary audio chunks (e.g., Opus)
ws.send(b"...binary chunk 1...")
ws.send(b"...binary chunk 2...")
# Signal end of audio
await ws.send(json.dumps({"type": "input_audio.end"}).encode())
async for msg in ws:
evt = json.loads(msg)
print(evt) # {'event': 'partial', 'text': '...'} etc.
asyncio.run(stream())
Summary
- WebSocket Transport: The backend exposes
/v1/audio/transcriptions/streaminspeech_platform.pyas the primary entry point for bidirectional streaming. - Binary Protocol: Accepts raw Opus or PCM16 audio frames, rejecting base64 or text-encoded payloads to optimize throughput.
- Session Safety: The
session.stream_openflag inworker/transport/server.pyenforces single-stream-per-session semantics. - Engine Flexibility: Routes audio through
sherpa_dictation.pyorasr_backend.pyto balance latency versus accuracy requirements. - Event-Driven: Uses structured JSON events (
partial,final,error) to communicate transcription state without polling overhead.
Frequently Asked Questions
What audio codecs does VoiceStudio support for real-time streaming?
VoiceStudio accepts two primary formats: audio/webm;codecs=opus for compressed streams and audio/pcm;encoding=s16le;channels=1 for uncompressed 16-bit PCM mono audio. The codec requirement is advertised in the StreamInputCapability discovery payload, allowing clients to negotiate the optimal format before transmission.
How does the backend prevent multiple simultaneous streams from a single client?
The session manager in backend/worker/transport/server.py tracks a session.stream_open boolean flag. When a client initiates a stream, the flag is set to True, and any subsequent stream attempts are rejected until the current session receives the input_audio.end control message and properly closes. This prevents resource exhaustion and ASR state corruption.
Can I use standard OpenAI client libraries with VoiceStudio's streaming endpoint?
Yes, VoiceStudio implements an OpenAI-compatible API surface. You can use client.audio.transcriptions.with_streaming_response.create() from the official OpenAI Python or Node.js clients by pointing the base_url to your VoiceStudio instance (e.g., http://localhost:3900/v1). The client handles WebSocket negotiation automatically while exposing the same partial/final event interface.
What happens if the client disconnects without sending the end-of-audio signal?
If the WebSocket connection closes abruptly without the input_audio.end control message, the session manager detects the disconnect through the transport layer and triggers cleanup routines. Any buffered audio in the ASR pipeline is processed as a final utterance, and the session.stream_open flag is reset to allow new connections from the same session identifier.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →