# How to Deploy VibeVoice with Tensor Parallelism: A Complete Guide

> Deploy VibeVoice with tensor parallelism easily using the --tp N flag. Automatically shard model weights across GPUs without manual setup for efficient performance.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**Deploy VibeVoice with tensor parallelism by passing the `--tp N` flag to the [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) launcher, which automatically configures vLLM to shard model weights across N GPUs without requiring manual torch.distributed setup.**

The microsoft/VibeVoice repository provides a production-ready automatic speech recognition (ASR) system that supports tensor parallelism to serve large models exceeding single-GPU memory limits. This guide explains how to deploy VibeVoice with tensor parallelism using the official one-click launcher, covering pure tensor parallel, data parallel, and hybrid deployment modes.

## Understanding Tensor Parallelism in VibeVoice

**Tensor Parallelism (TP)** splits individual model layers across multiple GPUs, allowing you to load checkpoints that would otherwise overflow a single device’s VRAM. When you deploy VibeVoice with tensor parallelism, the launcher script [`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py) handles all orchestration internally.

At lines 105-106 of the launcher, the script translates your `--tp N` argument into vLLM’s native `--tensor-parallel-size N` flag. The script then invokes `os.execvp()` to replace the Python process with a single vLLM server instance that internally manages NCCL communication between shards. No manual `torch.distributed.launch` or environment variable configuration is required.

## Deployment Modes: TP, DP, and Hybrid

The [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) launcher supports three distinct parallel strategies controlled via CLI flags:

### Tensor Parallel Only (`--tp`)

Use `--tp N` when serving one large model instance across N GPUs. This mode minimizes latency for individual requests and is ideal when the 3B or 7B VibeVoice checkpoint cannot fit on a single GPU.

### Data Parallel Only (`--dp`)

Use `--dp N` to spawn N independent model replicas, each residing on its own GPU. The launcher automatically configures an nginx reverse proxy to load-balance incoming requests across replicas, maximizing throughput for concurrent client connections.

### Hybrid Parallel (`--dp` and `--tp`)

Combine both flags (`--dp N --tp M`) to create N replicas, where each replica is tensor-parallelized across M GPUs. This requires N × M total GPUs and offers both high throughput via replication and large model support via sharding.

## Step-by-Step Deployment Examples

### Basic Tensor Parallel on 2 GPUs

Deploy the VibeVoice model across GPUs 0 and 1 using the official vLLM Docker image:

```bash
docker run -d --gpus '"device=0,1"' --name vibevoice-vllm \
  --ipc=host \
  -p 8000:8000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v "$(pwd)":/app \
  -w /app \
  --entrypoint bash \
  vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py --tp 2"

```

The `--tp 2` argument instructs the launcher to invoke vLLM with `--tensor-parallel-size 2`, sharding the model weights evenly across the two specified devices.

### Advanced Configuration with Custom Ports

For a 4-way tensor split with customized server settings and 90% GPU memory utilization:

```bash
docker run -d --gpus all --name vibevoice-vllm \
  --ipc=host \
  -p 9000:9000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=32 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v "$(pwd)":/app \
  -w /app \
  --entrypoint bash \
  vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py \
      --tp 4 \
      --port 9000 \
      --gpu-memory-utilization 0.9"

```

This configuration exposes the API on port 9000 and aggressively utilizes GPU memory for maximum batching capacity.

### Hybrid Deployment for Maximum Throughput

Run two independent replicas, each split across two GPUs, totaling four devices:

```bash
docker run -d --gpus '"device=0,1,2,3"' --name vibevoice-vllm \
  --ipc=host \
  -p 8000:8000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v "$(pwd)":/app \
  -w /app \
  --entrypoint bash \
  vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py \
      --dp 2 \
      --tp 2"

```

The launcher creates two vLLM workers (data parallel size 2), assigns each a tensor parallel size of 2, and configures nginx to distribute requests across both replicas.

## Verifying Your Deployment

Confirm the server is serving the VibeVoice model correctly:

```bash
curl http://localhost:8000/v1/models

```

You should receive a JSON response listing the model name as "vibevoice", indicating that the tensor-parallel weights loaded successfully and the API is ready for ASR inference.

## Key Components and Source Files

The deployment pipeline relies on these specific files from the microsoft/VibeVoice repository:

- **[`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py)** – Main entry point that parses `--tp` and `--dp` flags, builds the vLLM command via `_build_vllm_cmd()`, and executes the server.
- **[`vllm_plugin/tools/generate_tokenizer_files.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/tools/generate_tokenizer_files.py)** – Generates the extended tokenizer with audio tokens before server initialization.
- **[`vllm_plugin/model.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/model.py)** – Implements the `VibeVoiceModel` wrapper that integrates audio tensor processing into vLLM’s continuous batching engine.
- **[`docs/vibevoice-vllm-asr.md`](https://github.com/microsoft/VibeVoice/blob/main/docs/vibevoice-vllm-asr.md)** – Official documentation describing tensor parallel deployment options and performance tuning.

## Summary

- **Use `--tp N`** in [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) to deploy VibeVoice with tensor parallelism across N GPUs without manual distributed setup.
- **Tensor parallelism** shards model layers internally via vLLM’s NCCL backend, while **data parallelism** (`--dp`) creates independent replicas behind nginx.
- **Hybrid mode** (`--dp N --tp M`) scales to N × M GPUs for high-throughput serving of large checkpoints.
- The launcher handles dependency installation, tokenizer generation, and vLLM command construction automatically.

## Frequently Asked Questions

### What is the minimum GPU requirement for tensor parallelism in VibeVoice?

You need at least two GPUs to utilize tensor parallelism (`--tp 2`), though the specific count depends on your checkpoint size. The 3B parameter model typically requires `--tp 2` on consumer GPUs, while the 7B variant may need `--tp 4` or higher to fit within standard VRAM limits.

### Does VibeVoice require manual torch.distributed configuration for tensor parallelism?

No. The [`start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/start_server.py) launcher abstracts all distributed configuration. When you pass `--tp N`, the script automatically appends `--tensor-parallel-size N` to the vLLM command and launches a single process that internally handles inter-GPU communication via NCCL.

### Can I combine tensor parallelism with data parallelism?

Yes. Pass both `--dp N` and `--tp M` flags to create N replicas, each tensor-parallelized across M GPUs. The launcher configures nginx to load-balance across the N replicas, effectively multiplying throughput while maintaining the ability to serve large models.

### Where does the launcher configure the tensor parallel size?

The tensor parallel size is set in [`vllm_plugin/scripts/start_server.py`](https://github.com/microsoft/VibeVoice/blob/main/vllm_plugin/scripts/start_server.py) at lines 105-106, where the launcher constructs the vLLM command list and appends the `--tensor-parallel-size` argument based on your `--tp` input.