How to Deploy VibeVoice with Tensor Parallelism: A Complete Guide

Deploy VibeVoice with tensor parallelism by passing the --tp N flag to the start_server.py launcher, which automatically configures vLLM to shard model weights across N GPUs without requiring manual torch.distributed setup.

The microsoft/VibeVoice repository provides a production-ready automatic speech recognition (ASR) system that supports tensor parallelism to serve large models exceeding single-GPU memory limits. This guide explains how to deploy VibeVoice with tensor parallelism using the official one-click launcher, covering pure tensor parallel, data parallel, and hybrid deployment modes.

Understanding Tensor Parallelism in VibeVoice

Tensor Parallelism (TP) splits individual model layers across multiple GPUs, allowing you to load checkpoints that would otherwise overflow a single device’s VRAM. When you deploy VibeVoice with tensor parallelism, the launcher script vllm_plugin/scripts/start_server.py handles all orchestration internally.

At lines 105-106 of the launcher, the script translates your --tp N argument into vLLM’s native --tensor-parallel-size N flag. The script then invokes os.execvp() to replace the Python process with a single vLLM server instance that internally manages NCCL communication between shards. No manual torch.distributed.launch or environment variable configuration is required.

Deployment Modes: TP, DP, and Hybrid

The start_server.py launcher supports three distinct parallel strategies controlled via CLI flags:

Tensor Parallel Only (--tp)

Use --tp N when serving one large model instance across N GPUs. This mode minimizes latency for individual requests and is ideal when the 3B or 7B VibeVoice checkpoint cannot fit on a single GPU.

Data Parallel Only (--dp)

Use --dp N to spawn N independent model replicas, each residing on its own GPU. The launcher automatically configures an nginx reverse proxy to load-balance incoming requests across replicas, maximizing throughput for concurrent client connections.

Hybrid Parallel (--dp and --tp)

Combine both flags (--dp N --tp M) to create N replicas, where each replica is tensor-parallelized across M GPUs. This requires N × M total GPUs and offers both high throughput via replication and large model support via sharding.

Step-by-Step Deployment Examples

Basic Tensor Parallel on 2 GPUs

Deploy the VibeVoice model across GPUs 0 and 1 using the official vLLM Docker image:

docker run -d --gpus '"device=0,1"' --name vibevoice-vllm \
  --ipc=host \
  -p 8000:8000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v "$(pwd)":/app \
  -w /app \
  --entrypoint bash \
  vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py --tp 2"

The --tp 2 argument instructs the launcher to invoke vLLM with --tensor-parallel-size 2, sharding the model weights evenly across the two specified devices.

Advanced Configuration with Custom Ports

For a 4-way tensor split with customized server settings and 90% GPU memory utilization:

docker run -d --gpus all --name vibevoice-vllm \
  --ipc=host \
  -p 9000:9000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=32 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v "$(pwd)":/app \
  -w /app \
  --entrypoint bash \
  vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py \
      --tp 4 \
      --port 9000 \
      --gpu-memory-utilization 0.9"

This configuration exposes the API on port 9000 and aggressively utilizes GPU memory for maximum batching capacity.

Hybrid Deployment for Maximum Throughput

Run two independent replicas, each split across two GPUs, totaling four devices:

docker run -d --gpus '"device=0,1,2,3"' --name vibevoice-vllm \
  --ipc=host \
  -p 8000:8000 \
  -e VIBEVOICE_FFMPEG_MAX_CONCURRENCY=64 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  -v "$(pwd)":/app \
  -w /app \
  --entrypoint bash \
  vllm/vllm-openai:v0.14.1 \
  -c "python3 /app/vllm_plugin/scripts/start_server.py \
      --dp 2 \
      --tp 2"

The launcher creates two vLLM workers (data parallel size 2), assigns each a tensor parallel size of 2, and configures nginx to distribute requests across both replicas.

Verifying Your Deployment

Confirm the server is serving the VibeVoice model correctly:

curl http://localhost:8000/v1/models

You should receive a JSON response listing the model name as "vibevoice", indicating that the tensor-parallel weights loaded successfully and the API is ready for ASR inference.

Key Components and Source Files

The deployment pipeline relies on these specific files from the microsoft/VibeVoice repository:

Summary

  • Use --tp N in start_server.py to deploy VibeVoice with tensor parallelism across N GPUs without manual distributed setup.
  • Tensor parallelism shards model layers internally via vLLM’s NCCL backend, while data parallelism (--dp) creates independent replicas behind nginx.
  • Hybrid mode (--dp N --tp M) scales to N × M GPUs for high-throughput serving of large checkpoints.
  • The launcher handles dependency installation, tokenizer generation, and vLLM command construction automatically.

Frequently Asked Questions

What is the minimum GPU requirement for tensor parallelism in VibeVoice?

You need at least two GPUs to utilize tensor parallelism (--tp 2), though the specific count depends on your checkpoint size. The 3B parameter model typically requires --tp 2 on consumer GPUs, while the 7B variant may need --tp 4 or higher to fit within standard VRAM limits.

Does VibeVoice require manual torch.distributed configuration for tensor parallelism?

No. The start_server.py launcher abstracts all distributed configuration. When you pass --tp N, the script automatically appends --tensor-parallel-size N to the vLLM command and launches a single process that internally handles inter-GPU communication via NCCL.

Can I combine tensor parallelism with data parallelism?

Yes. Pass both --dp N and --tp M flags to create N replicas, each tensor-parallelized across M GPUs. The launcher configures nginx to load-balance across the N replicas, effectively multiplying throughput while maintaining the ability to serve large models.

Where does the launcher configure the tensor parallel size?

The tensor parallel size is set in vllm_plugin/scripts/start_server.py at lines 105-106, where the launcher constructs the vLLM command list and appends the --tensor-parallel-size argument based on your --tp input.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →