# How Streaming Text Input Works in VibeVoice-Realtime: Architecture and Code

> Discover how VibeVoice-Realtime handles streaming text input. Learn about its architecture, tokenization, and interleaved inference for immediate audio generation.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: architecture
- Published: 2026-03-28

---

**VibeVoice-Realtime processes streaming text input by tokenizing incoming text into fixed-size windows and interleaving language model inference with speech diffusion, enabling audio generation to begin after the first few tokens while subsequent text continues to arrive.**

VibeVoice-Realtime is Microsoft's open-source text-to-speech system designed for real-time voice synthesis. The streaming text input architecture allows the model to generate audible speech without waiting for complete scripts, achieving approximately 200ms latency from text arrival to audio output. This article examines the windowed processing loop, interleaved diffusion mechanism, and WebSocket integration that make live "type-as-you-talk" experiences possible.

## High-Level Architecture

The streaming pipeline consists of four coordinated components that handle text ingestion, model inference, audio generation, and network delivery:

- **`VibeVoiceStreamingProcessor.process_input_with_cached_prompt`** – Prepares the initial cached voice prompt containing the speaker's KV-cache. Located in [`vibevoice/processor/vibevoice_streaming_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_streaming_processor.py) (lines 162-170), this runs once per speaker and does not handle the actual streaming text windows.

- **`VibeVoiceStreamingForConditionalGenerationInference.generate`** – The core generation engine in [`vibevoice/modular/modeling_vibevoice_streaming_inference.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_streaming_inference.py) (lines 595-604). This method slices the full tokenized script into windows of length `TTS_TEXT_WINDOW_SIZE`, feeding each window sequentially into the language model while simultaneously diffusing speech tokens for previous windows.

- **Audio Streamer** – An optional callback interface that receives generated speech latents chunk-by-chunk and yields PCM audio frames. The demo service wraps this generator to stream binary audio over WebSocket connections.

- **Web UI / Demo** – The FastAPI application in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py) exposes a `/stream` WebSocket endpoint that forwards user input to `StreamingTTSService.stream`, continuously yielding audio chunks as the model processes new text windows (lines 59-71).

## Windowed Text Processing

The streaming mechanism relies on **fixed-size text windows** defined by the constant `TTS_TEXT_WINDOW_SIZE`. Rather than requiring the entire tokenized script upfront, the generator processes text incrementally through a sliding window approach.

At initialization, the first window size is computed based on available tokens:

```python
first_text_window_size = TTS_TEXT_WINDOW_SIZE if tts_text_ids.shape[1] >= TTS_TEXT_WINDOW_SIZE else tts_text_ids.shape[1]

```

This logic appears at lines 671-672 in [`modeling_vibevoice_streaming_inference.py`](https://github.com/microsoft/VibeVoice/blob/main/modeling_vibevoice_streaming_inference.py). Inside the main generation loop, the current window is extracted via tensor slicing:

```python
cur_input_tts_text_ids = tts_text_ids[:, tts_text_window_index*TTS_TEXT_WINDOW_SIZE:(tts_text_window_index+1)*TTS_TEXT_WINDOW_SIZE]

```

The implementation at line 726 handles the slicing, while line 727 pre-computes the next window size to prepare for state updates. After each forward pass, the LM and TTS-LM caches are updated (lines 730-734), allowing the model to maintain context across windows while releasing audio chunks incrementally.

## Interleaved Speech Diffusion

For every text window processed, the model executes a four-stage pipeline that interleaves text encoding with speech synthesis:

1. **Prefill** – The new token window is prefixed to the current language model input, updating the KV-cache states.
2. **Forward Pass** – `self.forward_lm` processes the concatenated input, producing hidden states for speech token prediction.
3. **Diffusion Sampling** – `sample_speech_tokens` generates a block of speech tokens using the diffusion model conditioned on the current LM states.
4. **Audio Yield** – The decoded audio chunk is pushed to the `audio_streamer` interface, making it available to the client immediately.

This loop continues until an EOS token is generated, the maximum sequence length is reached, or an external stop signal interrupts execution. The architecture ensures that audio generation begins after processing the first window (typically achieving ~200ms latency) while subsequent windows are processed in parallel with playback.

## External Stop Handling

The `generate` method accepts a `stop_check_fn` callback that enables **live cancellation** from client applications. If this function returns `True`, the generation loop aborts immediately, the audio streamer is closed, and resources are released gracefully (lines 4-13 in the inference module). This mechanism allows front-end applications to stop synthesis instantly when users interrupt the stream or navigate away from the session.

## Implementation Example

The following Python pattern mirrors the production implementation used in the WebSocket demo, demonstrating how to wire the processor and inference components for streaming text input:

```python
from vibevoice.modular.modeling_vibevoice_streaming_inference import (
    VibeVoiceStreamingForConditionalGenerationInference,
)
from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor

# 1️⃣ Build the cached voice prompt (runs once per speaker)

processor = VibeVoiceStreamingProcessor(tokenizer, speech_encoder)
cached_prompt = processor.process_input_with_cached_prompt(
    text=None, cached_prompt=voice_prompt_dict
)

# 2️⃣ Tokenise the script (can be incremental)

tts_text_ids = tokenizer(text, return_tensors="pt")["input_ids"]

# 3️⃣ Run streaming generation

model = VibeVoiceStreamingForConditionalGenerationInference.from_pretrained(
    "microsoft/VibeVoice-Realtime-0.5B", device="cuda"
)
for audio_chunk in model.generate(
    tts_text_ids=tts_text_ids,
    cached_prompt=cached_prompt,
    audio_streamer=my_streamer,      # implements .stream() and .end()

    stop_check_fn=my_stop_flag,
):
    # audio_chunk is a NumPy array (PCM16) that can be sent to a client

    send_to_client(audio_chunk)

```

The `model.generate` method returns an iterator that yields `np.ndarray` chunks containing PCM16 audio data, allowing your application to stream bytes to clients as they become available.

## WebSocket Demo Integration

The repository includes a complete FastAPI demonstration in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py) that exposes the streaming generator via WebSocket. The implementation wraps the model iterator in an async-compatible interface:

```python
def streaming_tts(text: str, **kwargs) -> Iterator[np.ndarray]:
    service: StreamingTTSService = app.state.tts_service
    yield from service.stream(text, **kwargs)

@app.websocket("/stream")
async def websocket_stream(ws: WebSocket) -> None:
    await ws.accept()
    text = ws.query_params.get("text", "")
    iterator = streaming_tts(text, cfg_scale=cfg_scale, ...)
    while ws.client_state == WebSocketState.CONNECTED:
        chunk = await asyncio.to_thread(next, iterator, sentinel)
        if chunk is sentinel:
            break
        await ws.send_bytes(chunk.tobytes())

```

Clients connect to `ws://<host>/stream?text=Hello+world` and receive binary PCM audio chunks in real-time. The server handles text windowing and diffusion asynchronously, demonstrating how the streaming text input architecture translates to production WebSocket services.

## Summary

- **Windowed Processing** – VibeVoice-Realtime splits tokenized input into fixed-size windows using `TTS_TEXT_WINDOW_SIZE`, processing each slice independently while maintaining KV-cache continuity across iterations.
- **Interleaved Generation** – The main loop in `VibeVoiceStreamingForConditionalGenerationInference.generate` alternates between language model forward passes and diffusion-based speech token sampling, yielding audio after each window.
- **Low Latency** – Audio generation begins after the first text window (approximately 200ms), enabling true real-time synthesis without requiring complete transcripts upfront.
- **Cancellable Streams** – The `stop_check_fn` callback provides immediate cancellation capabilities, allowing clients to abort generation gracefully.
- **Production Ready** – The WebSocket demo in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py) demonstrates complete integration, streaming PCM audio chunks via `/stream` endpoints as text windows are processed.

## Frequently Asked Questions

### What is the latency for streaming text input in VibeVoice-Realtime?

According to the source code in [`vibevoice/modular/modeling_vibevoice_streaming_inference.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_streaming_inference.py), the system achieves approximately 200ms latency from text arrival to first audio output. This is possible because the model begins processing the first `TTS_TEXT_WINDOW_SIZE` tokens immediately, without waiting for the complete input sequence to arrive.

### How does VibeVoice-Realtime handle partial text inputs?

The architecture uses tensor slicing operations at lines 726-727 to extract windows from the full `tts_text_ids` tensor. As new text arrives, it is tokenized and appended to the input sequence, while the generation loop continues processing subsequent windows. The cached LM and TTS-LM states maintain context across these incremental updates.

### Can streaming generation be cancelled mid-sentence?

Yes. The `generate` method accepts a `stop_check_fn` parameter that is checked within the main generation loop. When this function returns `True`, the loop exits immediately, the audio streamer is closed, and the iterator stops yielding chunks. This mechanism is implemented in the inference module between lines 4-13 and is utilized by the WebSocket demo for live cancellation.

### What determines the size of text windows in the streaming processor?

The window size is controlled by the `TTS_TEXT_WINDOW_SIZE` constant defined in the inference module. The first window may be smaller if the initial input contains fewer tokens than this constant (as shown in lines 671-672), but subsequent windows adhere to the fixed size until the final chunk, which contains the remaining tokens. This fixed-size approach balances computational efficiency with streaming responsiveness.