# How to Use VibeVoice for Real-Time Streaming TTS: Architecture and Implementation Guide

> Learn how to use VibeVoice for real-time streaming TTS. Explore its three-layer pipeline for efficient text-to-speech generation and audio streaming.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**VibeVoice implements a three-layer streaming pipeline—Processor, Streaming Model, and AudioStreamer—that ingests text in 5-token windows, generates speech via diffusion, and yields audio chunks through a thread-safe queue before the full utterance completes.**

VibeVoice is Microsoft's open-source neural text-to-speech system designed for low-latency streaming synthesis. This guide explains how to use VibeVoice for real-time streaming TTS by leveraging its modular inference classes, windowed generation strategy, and queue-based audio delivery, as implemented in the `microsoft/VibeVoice` repository.

## The Three-Layer Streaming Architecture

VibeVoice separates streaming concerns into distinct layers to enable incremental processing without blocking:

- **Processor Layer** ([`vibevoice/processor/vibevoice_streaming_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_streaming_processor.py)): The `VibeVoiceStreamingProcessor` class loads the tokenizer and audio normalizer from a pretrained checkpoint, preparing inputs for the streaming model.

- **Streaming Model Layer** ([`vibevoice/modular/modeling_vibevoice_streaming_inference.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_streaming_inference.py)): The `VibeVoiceStreamingForConditionalGenerationInference` class manages a dual-stage inference loop—alternating between a text language model (for windowed prefill) and a TTS language model (for diffusion-based speech token generation).

- **Streamer Layer** ([`vibevoice/modular/streamer.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/streamer.py)): The `AudioStreamer` provides a thread-safe queue via `put()` and an iterator interface via `get_stream()`, allowing audio chunks to be consumed as they are produced rather than waiting for the full waveform.

## Data Flow for Real-Time TTS

The streaming pipeline follows a six-stage pipeline that interleaves text ingestion with audio generation.

### 1. Initialize Processor and Model

The service loads the processor and model using `from_pretrained`, configures the device (CUDA, CPU, or MPS), and sets the diffusion inference step count (default 5). In [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py) (lines 68-84), the startup sequence creates these components:

```python
from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor
from vibevoice.modular.modeling_vibevoice_streaming_inference import VibeVoiceStreamingForConditionalGenerationInference

processor = VibeVoiceStreamingProcessor.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")
model = VibeVoiceStreamingForConditionalGenerationInference.from_pretrained(
    "microsoft/VibeVoice-Realtime-0.5B",
    torch_dtype="bfloat16",
    device_map="auto",
    attn_implementation="flash_attention_2",
)

```

### 2. Cache Voice Prompts

Before streaming begins, a pre-filled voice prompt (`.pt` file) is loaded and cached via `_ensure_voice_cached`. This cache contains the KV tensors for the speaker's latent style, enabling zero-latency voice switching during generation (see [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py) lines 150-166).

### 3. Windowed Text Ingestion

Incoming text is tokenized and sliced into **text windows** of size `TTS_TEXT_WINDOW_SIZE` (5 tokens). Each window is concatenated to the previously generated sequence and fed to the text LM via `forward_lm`. The model updates its cache using `_update_model_kwargs_for_generation` so subsequent windows reuse past KV states (see [`vibevoice/modular/modeling_vibevoice_streaming_inference.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_streaming_inference.py) lines 1048-1064).

### 4. Diffusion-Based Speech Generation

After each text window, the system enters a **speech-window loop** (`TTS_SPEECH_WINDOW_SIZE = 6`). For each step:

1. **Sample speech tokens** using `sample_speech_tokens` with classifier-free guidance.
2. **Decode latents** to raw audio via `self.model.acoustic_tokenizer.decode`, reusing an `acoustic_cache` for throughput.
3. **Push to streamer** via `audio_streamer.put()` so clients receive audio immediately (see lines 886-898 and 998-1002).

### 5. Detect End-of-Speech

A binary classifier (`tts_eos_classifier`) evaluates the TTS LM's last hidden state. When confidence exceeds 0.5, generation stops and `audio_streamer.end()` signals completion (lines 847-853).

### 6. Stream Audio to Client

The FastAPI WebSocket endpoint (`/stream` in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py) lines 63-71) converts float32 tensors to 16-bit PCM via `chunk_to_pcm16` and transmits via `ws.send_bytes()`. The browser client plays the PCM stream immediately without buffering the entire file.

## Tuning Parameters for Low Latency

| Parameter | Default | Impact | Tuning Guidance |
|-----------|---------|--------|-----------------|
| **Text window size** (`TTS_TEXT_WINDOW_SIZE`) | 5 tokens | Smaller values reduce prefill delay but increase overhead. | Keep at 5 for balanced latency; increase for very long prompts. |
| **Speech window size** (`TTS_SPEECH_WINDOW_SIZE`) | 6 steps | Controls diffusion iterations per text chunk. | 6 is optimized for the 0.5B realtime model; larger models may use 8. |
| **Inference steps** (`--inference_steps`) | 5 DDPM steps | More steps improve quality at the cost of latency. | Use 5 for real-time; raise to 8 for offline high-fidelity. |
| **Device & Attention** | CUDA | Flash Attention 2 on GPU maximizes throughput. | Set `attn_implementation="flash_attention_2"` with `device_map="auto"`. |

## Implementation Examples

### Pure Python Streaming Iterator

Use this pattern to integrate VibeVoice into a Python application without a web server:

```python
import torch
import numpy as np
from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor
from vibevoice.modular.modeling_vibevoice_streaming_inference import VibeVoiceStreamingForConditionalGenerationInference
from vibevoice.modular.streamer import AudioStreamer

# Load components

processor = VibeVoiceStreamingProcessor.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")
model = VibeVoiceStreamingForConditionalGenerationInference.from_pretrained(
    "microsoft/VibeVoice-Realtime-0.5B",
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",
)

# Load voice prompt cache

voice = torch.load("voices/streaming_model/en-Carter_man.pt", map_location="cpu")
prefilled = {
    "lm": voice["lm"],
    "tts_lm": voice["tts_lm"],
    "neg_lm": voice["neg_lm"],
    "neg_tts_lm": voice["neg_tts_lm"],
}

# Prepare text

text = "Hello, this is real-time streaming TTS with VibeVoice."
tts_ids = torch.tensor([processor.tokenizer.encode(text, add_special_tokens=False)], dtype=torch.long)

# Initialize streamer

streamer = AudioStreamer(batch_size=1, stop_signal=None)

# Generate (non-blocking)

model.generate(
    inputs=None,
    tts_text_ids=tts_ids,
    audio_streamer=streamer,
    all_prefilled_outputs=prefilled,
    tokenizer=processor.tokenizer,
)

# Consume chunks

for chunk in streamer.get_stream(0):
    wav = (chunk.cpu().numpy() * 32767).astype("int16")
    # Feed to audio device or file writer

```

### FastAPI WebSocket Server

Deploy a streaming endpoint that mirrors the official demo:

```python
from fastapi import FastAPI, WebSocket
from vibevoice.modular.streamer import AudioStreamer

app = FastAPI()
service = None  # Initialize on startup with StreamingTTSService

@app.websocket("/stream")
async def websocket_endpoint(websocket: WebSocket):
    await websocket.accept()
    text = websocket.query_params.get("text", "")
    
    streamer = AudioStreamer(batch_size=1, stop_signal=None)
    
    # Trigger generation (run in background thread in production)

    service.stream(text, audio_streamer=streamer)
    
    # Forward PCM-16 chunks as they arrive

    for audio_chunk in streamer.get_stream(0):
        pcm16 = (audio_chunk.cpu().numpy() * 32767).astype("int16").tobytes()
        await websocket.send_bytes(pcm16)
    
    await websocket.close()

```

### Command-Line Utility

Stream to a WAV file for testing:

```python
import argparse
import wave
import numpy as np
import torch
from vibevoice.processor.vibevoice_streaming_processor import VibeVoiceStreamingProcessor
from vibevoice.modular.modeling_vibevoice_streaming_inference import VibeVoiceStreamingForConditionalGenerationInference
from vibevoice.modular.streamer import AudioStreamer

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--text", required=True)
    parser.add_argument("--output", default="output.wav")
    args = parser.parse_args()
    
    processor = VibeVoiceStreamingProcessor.from_pretrained("microsoft/VibeVoice-Realtime-0.5B")
    model = VibeVoiceStreamingForConditionalGenerationInference.from_pretrained(
        "microsoft/VibeVoice-Realtime-0.5B",
        torch_dtype=torch.bfloat16,
        device_map="auto",
    )
    
    voice = torch.load("voices/streaming_model/en-Carter_man.pt", map_location="cpu")
    prefilled = {
        "lm": voice["lm"], "tts_lm": voice["tts_lm"],
        "neg_lm": voice["neg_lm"], "neg_tts_lm": voice["neg_tts_lm"]
    }
    
    tts_ids = torch.tensor([processor.tokenizer.encode(args.text, add_special_tokens=False)], dtype=torch.long)
    streamer = AudioStreamer(batch_size=1, stop_signal=None)
    
    model.generate(
        inputs=None,
        tts_text_ids=tts_ids,
        audio_streamer=streamer,
        all_prefilled_outputs=prefilled,
        tokenizer=processor.tokenizer,
    )
    
    # Collect and write

    audio = [chunk.cpu().numpy() for chunk in streamer.get_stream(0)]
    wav = np.concatenate(audio)
    
    with wave.open(args.output, "wb") as f:
        f.setnchannels(1)
        f.setsampwidth(2)
        f.setframerate(24000)
        f.writeframes((wav * 32767).astype("int16").tobytes())

if __name__ == "__main__":
    main()

```

## Summary

- **VibeVoice** implements real-time streaming TTS through a **Processor → Streaming Model → AudioStreamer** architecture.
- **Windowed generation** processes text in 5-token chunks (`TTS_TEXT_WINDOW_SIZE`) and speech in 6-step diffusion windows (`TTS_SPEECH_WINDOW_SIZE`).
- **Immediate delivery** occurs via `AudioStreamer.put()` and `get_stream()`, yielding audio before the full utterance completes.
- **Voice caching** via pre-filled `.pt` files eliminates speaker-switching latency.
- **EOS detection** uses a binary classifier on the TTS LM hidden states to terminate generation cleanly.
- The **FastAPI demo** in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py) provides a complete WebSocket reference implementation.

## Frequently Asked Questions

### How does VibeVoice achieve low latency in streaming mode?

VibeVoice achieves low latency by interleaving text processing and audio generation. The model processes incoming text in small 5-token windows while simultaneously running 6-step diffusion loops for speech tokens. Because the `AudioStreamer` queue yields chunks immediately via `put()` and `get_stream()`, audio begins transmitting before the full text is processed or the complete waveform is generated.

### What file formats does VibeVoice use for voice prompts?

VibeVoice uses PyTorch serialized tensors (`.pt` files) for voice prompts. These files contain pre-computed KV caches for both the text LM and TTS LM, stored under keys `"lm"`, `"tts_lm"`, `"neg_lm"`, and `"neg_tts_lm"` for classifier-free guidance. The demo repository provides reference voices like `en-Carter_man.pt` in the `voices/streaming_model/` directory.

### Can I adjust the trade-off between audio quality and generation speed?

Yes. The primary tuning lever is `--inference_steps` (default 5), which controls the number of DDPM diffusion steps. Reducing this value decreases latency but may reduce audio fidelity. Additionally, you can adjust `TTS_TEXT_WINDOW_SIZE` and `TTS_SPEECH_WINDOW_SIZE` in the model configuration, though the defaults (5 and 6 respectively) are optimized for the 0.5B realtime model.

### How do I deploy VibeVoice for production WebSocket streaming?

Deploy using the FastAPI pattern shown in [`demo/web/app.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/web/app.py). Load the `VibeVoiceStreamingProcessor` and `VibeVoiceStreamingForConditionalGenerationInference` during startup (not per-request), cache voice prompts via `_ensure_voice_cached`, and run the `model.generate()` call in a background thread. Stream PCM-16 bytes to the client as they arrive from `AudioStreamer.get_stream()`, ensuring your client can handle 24kHz mono PCM audio.