How Streaming Text Input Works in VibeVoice-Realtime: Architecture and Code
VibeVoice-Realtime processes streaming text input by tokenizing incoming text into fixed-size windows and interleaving language model inference with speech diffusion, enabling audio generation to begin after the first few tokens while subsequent text continues to arrive.
VibeVoice-Realtime is Microsoft's open-source text-to-speech system designed for real-time voice synthesis. The streaming text input architecture allows the model to generate audible speech without waiting for complete scripts, achieving approximately 200ms latency from text arrival to audio output. This article examines the windowed processing loop, interleaved diffusion mechanism, and WebSocket integration that make live "type-as-you-talk" experiences possible.
High-Level Architecture
The streaming pipeline consists of four coordinated components that handle text ingestion, model inference, audio generation, and network delivery:
-
VibeVoiceStreamingProcessor.process_input_with_cached_prompt– Prepares the initial cached voice prompt containing the speaker's KV-cache. Located invibevoice/processor/vibevoice_streaming_processor.py(lines 162-170), this runs once per speaker and does not handle the actual streaming text windows. -
VibeVoiceStreamingForConditionalGenerationInference.generate– The core generation engine invibevoice/modular/modeling_vibevoice_streaming_inference.py(lines 595-604). This method slices the full tokenized script into windows of lengthTTS_TEXT_WINDOW_SIZE, feeding each window sequentially into the language model while simultaneously diffusing speech tokens for previous windows. -
Audio Streamer – An optional callback interface that receives generated speech latents chunk-by-chunk and yields PCM audio frames. The demo service wraps this generator to stream binary audio over WebSocket connections.
-
Web UI / Demo – The FastAPI application in
demo/web/app.pyexposes a/streamWebSocket endpoint that forwards user input toStreamingTTSService.stream, continuously yielding audio chunks as the model processes new text windows (lines 59-71).
Windowed Text Processing
The streaming mechanism relies on fixed-size text windows defined by the constant TTS_TEXT_WINDOW_SIZE. Rather than requiring the entire tokenized script upfront, the generator processes text incrementally through a sliding window approach.
At initialization, the first window size is computed based on available tokens:
first_text_window_size = TTS_TEXT_WINDOW_SIZE if tts_text_ids.shape[1] >= TTS_TEXT_WINDOW_SIZE else tts_text_ids.shape[1]
This logic appears at lines 671-672 in modeling_vibevoice_streaming_inference.py. Inside the main generation loop, the current window is extracted via tensor slicing:
cur_input_tts_text_ids = tts_text_ids[:, tts_text_window_index*TTS_TEXT_WINDOW_SIZE:(tts_text_window_index+1)*TTS_TEXT_WINDOW_SIZE]
The implementation at line 726 handles the slicing, while line 727 pre-computes the next window size to prepare for state updates. After each forward pass, the LM and TTS-LM caches are updated (lines 730-734), allowing the model to maintain context across windows while releasing audio chunks incrementally.
Interleaved Speech Diffusion
For every text window processed, the model executes a four-stage pipeline that interleaves text encoding with speech synthesis:
- Prefill – The new token window is prefixed to the current language model input, updating the KV-cache states.
- Forward Pass –
self.forward_lmprocesses the concatenated input, producing hidden states for speech token prediction. - Diffusion Sampling –
sample_speech_tokensgenerates a block of speech tokens using the diffusion model conditioned on the current LM states. - Audio Yield – The decoded audio chunk is pushed to the
audio_streamerinterface, making it available to the client immediately.
This loop continues until an EOS token is generated, the maximum sequence length is reached, or an external stop signal interrupts execution. The architecture ensures that audio generation begins after processing the first window (typically achieving ~200ms latency) while subsequent windows are processed in parallel with playback.
External Stop Handling
The generate method accepts a stop_check_fn callback that enables live cancellation from client applications. If this function returns True, the generation loop aborts immediately, the audio streamer is closed, and resources are released gracefully (lines 4-13 in the inference module). This mechanism allows front-end applications to stop synthesis instantly when users interrupt the stream or navigate away from the session.
Implementation Example
The following Python pattern mirrors the production implementation used in the WebSocket demo, demonstrating how to wire the processor and inference components for streaming text input:
from vibevoice.modular.modeling_vibevoice_streaming_inference import (
VibeVoiceStreamingForConditionalGenerationInference,
)
from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor
# 1️⃣ Build the cached voice prompt (runs once per speaker)
processor = VibeVoiceStreamingProcessor(tokenizer, speech_encoder)
cached_prompt = processor.process_input_with_cached_prompt(
text=None, cached_prompt=voice_prompt_dict
)
# 2️⃣ Tokenise the script (can be incremental)
tts_text_ids = tokenizer(text, return_tensors="pt")["input_ids"]
# 3️⃣ Run streaming generation
model = VibeVoiceStreamingForConditionalGenerationInference.from_pretrained(
"microsoft/VibeVoice-Realtime-0.5B", device="cuda"
)
for audio_chunk in model.generate(
tts_text_ids=tts_text_ids,
cached_prompt=cached_prompt,
audio_streamer=my_streamer, # implements .stream() and .end()
stop_check_fn=my_stop_flag,
):
# audio_chunk is a NumPy array (PCM16) that can be sent to a client
send_to_client(audio_chunk)
The model.generate method returns an iterator that yields np.ndarray chunks containing PCM16 audio data, allowing your application to stream bytes to clients as they become available.
WebSocket Demo Integration
The repository includes a complete FastAPI demonstration in demo/web/app.py that exposes the streaming generator via WebSocket. The implementation wraps the model iterator in an async-compatible interface:
def streaming_tts(text: str, **kwargs) -> Iterator[np.ndarray]:
service: StreamingTTSService = app.state.tts_service
yield from service.stream(text, **kwargs)
@app.websocket("/stream")
async def websocket_stream(ws: WebSocket) -> None:
await ws.accept()
text = ws.query_params.get("text", "")
iterator = streaming_tts(text, cfg_scale=cfg_scale, ...)
while ws.client_state == WebSocketState.CONNECTED:
chunk = await asyncio.to_thread(next, iterator, sentinel)
if chunk is sentinel:
break
await ws.send_bytes(chunk.tobytes())
Clients connect to ws://<host>/stream?text=Hello+world and receive binary PCM audio chunks in real-time. The server handles text windowing and diffusion asynchronously, demonstrating how the streaming text input architecture translates to production WebSocket services.
Summary
- Windowed Processing – VibeVoice-Realtime splits tokenized input into fixed-size windows using
TTS_TEXT_WINDOW_SIZE, processing each slice independently while maintaining KV-cache continuity across iterations. - Interleaved Generation – The main loop in
VibeVoiceStreamingForConditionalGenerationInference.generatealternates between language model forward passes and diffusion-based speech token sampling, yielding audio after each window. - Low Latency – Audio generation begins after the first text window (approximately 200ms), enabling true real-time synthesis without requiring complete transcripts upfront.
- Cancellable Streams – The
stop_check_fncallback provides immediate cancellation capabilities, allowing clients to abort generation gracefully. - Production Ready – The WebSocket demo in
demo/web/app.pydemonstrates complete integration, streaming PCM audio chunks via/streamendpoints as text windows are processed.
Frequently Asked Questions
What is the latency for streaming text input in VibeVoice-Realtime?
According to the source code in vibevoice/modular/modeling_vibevoice_streaming_inference.py, the system achieves approximately 200ms latency from text arrival to first audio output. This is possible because the model begins processing the first TTS_TEXT_WINDOW_SIZE tokens immediately, without waiting for the complete input sequence to arrive.
How does VibeVoice-Realtime handle partial text inputs?
The architecture uses tensor slicing operations at lines 726-727 to extract windows from the full tts_text_ids tensor. As new text arrives, it is tokenized and appended to the input sequence, while the generation loop continues processing subsequent windows. The cached LM and TTS-LM states maintain context across these incremental updates.
Can streaming generation be cancelled mid-sentence?
Yes. The generate method accepts a stop_check_fn parameter that is checked within the main generation loop. When this function returns True, the loop exits immediately, the audio streamer is closed, and the iterator stops yielding chunks. This mechanism is implemented in the inference module between lines 4-13 and is utilized by the WebSocket demo for live cancellation.
What determines the size of text windows in the streaming processor?
The window size is controlled by the TTS_TEXT_WINDOW_SIZE constant defined in the inference module. The first window may be smaller if the initial input contains fewer tokens than this constant (as shown in lines 671-672), but subsequent windows adhere to the fixed size until the final chunk, which contains the remaining tokens. This fixed-size approach balances computational efficiency with streaming responsiveness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →