How to Process Long-Form Audio with VibeVoice-ASR: A Complete Technical Guide
VibeVoice-ASR transcribes audio up to 60 minutes (or longer) in a single forward pass by streaming acoustic and semantic tokenizers and inserting speech representations only once during the first generation step.
The microsoft/VibeVoice repository provides a production-ready automatic speech recognition system designed specifically for long-form content. Unlike traditional ASR models that struggle with memory constraints on extended recordings, VibeVoice-ASR implements segment-wise streaming tokenization that keeps memory usage bounded while maintaining full contextual awareness across hour-long inputs.
How VibeVoice-ASR Handles Long-Form Audio
VibeVoice-ASR achieves long-form capability through two architectural innovations: streaming tokenization and single-pass feature insertion. The system automatically detects audio duration and switches to a chunked processing mode when inputs exceed 60 seconds, splitting the waveform into manageable segments without losing semantic coherence.
The processor first loads audio using FFmpeg (when available) and resamples to the model's 24 kHz target rate. In vibevoice/processor/vibevoice_asr_processor.py, the _process_single_audio method calculates duration as len(audio_array) / self.target_sample_rate and sets use_streaming=True when the threshold is exceeded. This triggers the streaming pathway in the model's encoder.
The Streaming Architecture
The long-form pipeline relies on three core components that work together to process extended sequences without exhausting GPU memory.
Audio Segmentation and Tokenization
In vibevoice/modular/modeling_vibevoice_asr.py, the encode_speech method implements the core streaming logic. The method splits incoming audio into 60-second segments (segment_samples = 60 × 24000 samples) and processes each through the acoustic and semantic tokenizers separately.
For each segment, the code calls acoustic_tokenizer.encode(..., cache=..., use_cache=True, is_final_chunk=...) to produce mean representations without sampling. After all segments complete, the system concatenates the means and performs sampling exactly once using the learned standard deviation. This approach prevents the model from building excessively large convolution buffers that would otherwise exhaust memory on hour-long recordings.
Single-Pass Feature Insertion
The prepare_inputs_for_generation function (lines 91-106 in vibevoice/modular/modeling_vibevoice_asr.py) implements a critical optimization: it checks if cache_position[0] == 0 to identify the first generation step. Only on this initial pass does it insert the speech feature tensors and acoustic_input_mask at the <|speech_pad|> placeholder positions. Subsequent generation steps receive None for speech inputs, allowing the language model to continue autoregressively without重复处理 the acoustic features.
This design ensures that heavy acoustic and semantic encoding occurs only once per input, while the language model can generate up to 64,000 tokens of transcription using standard causal attention.
Processing Single Long Audio Files
To transcribe a single lengthy recording, use the VibeVoiceASRProcessor with automatic streaming detection:
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from vibevoice.modular.modeling_vibevoice_asr import VibeVoiceASRForConditionalGeneration
import torch
# Initialize processor with pretrained weights
processor = VibeVoiceASRProcessor.from_pretrained(
"microsoft/VibeVoice-ASR",
language_model_pretrained_name="Qwen/Qwen2.5-7B"
)
# Load model with automatic device mapping
model = VibeVoiceASRForConditionalGeneration.from_pretrained(
"microsoft/VibeVoice-ASR",
torch_dtype=torch.bfloat16,
device_map="auto",
attn_implementation="sdpa",
trust_remote_code=True,
)
model.eval()
# Process audio of any length - streaming activates automatically for >60s
inputs = processor(
audio="path/to/60_minute_podcast.wav",
return_tensors="pt",
padding=True,
add_generation_prompt=True,
)
# Move to device and generate
device = next(model.parameters()).device
inputs = {k: v.to(device) if isinstance(v, torch.Tensor) else v
for k, v in inputs.items()}
generated_ids = model.generate(**inputs, max_new_tokens=32768)
raw_text = processor.decode(generated_ids[0], skip_special_tokens=True)
# Extract structured segments with timestamps
segments = processor.post_process_transcription(raw_text)
for seg in segments:
print(f"[{seg.get('start_time'):.2f}s – {seg.get('end_time'):.2f}s] "
f"Speaker {seg.get('speaker_id')}: {seg.get('text')}")
The processor automatically constructs the token sequence containing speech placeholders, and the model handles the segment-wise encoding internally.
Batch Processing Multiple Files
For production workflows involving multiple long recordings, use the VibeVoiceASRBatchInference class from the demo scripts:
from demo.vibevoice_asr_inference_from_file import VibeVoiceASRBatchInference
# Initialize inference helper
asr = VibeVoiceASRBatchInference(
model_path="microsoft/VibeVoice-ASR",
device="cuda",
dtype=torch.bfloat16,
attn_implementation="sdpa",
)
# Process multiple lengthy files with automatic batching
audio_files = [
"meeting_recording_1.wav",
"meeting_recording_2.mp4", # Video files supported via FFmpeg
"interview_session.mp3",
]
results = asr.transcribe_with_batching(
audio_inputs=audio_files,
batch_size=2, # Adjust based on GPU memory
max_new_tokens=32768,
temperature=0.0, # Greedy decoding for accuracy
do_sample=False,
)
# Access structured results
for r in results:
print(f"File: {r['file']} ({r['generation_time']:.2f}s)")
for seg in r["segments"]:
print(f" [{seg.get('start_time'):.2f}s] {seg.get('text')}")
This implementation in demo/vibevoice_asr_inference_from_file.py handles the complete pipeline including loading, streaming tokenization, and structured output generation.
Generating Synthetic Long-Form Data for Testing
To benchmark the system or test memory constraints, concatenate existing datasets into extended sequences:
from demo.vibevoice_asr_inference_from_file import load_dataset_and_concatenate
# Create a 3-hour synthetic audio from Librispeech
long_audios = load_dataset_and_concatenate(
dataset_name="openslr/librispeech_asr",
split="test",
max_duration=10800, # 3 hours in seconds
num_audios=1,
target_sr=24000,
)
# Transcribe the synthetic long-form input
results = asr.transcribe_with_batching(
audio_inputs=long_audios,
batch_size=1,
max_new_tokens=64000, # Extended context for very long inputs
)
The load_dataset_and_concatenate function demonstrates how the streaming pipeline handles arbitrarily concatenated audio while maintaining accurate timestamps across segment boundaries.
Summary
- Automatic streaming: The processor detects audio >60 seconds and activates the streaming pathway in
encode_speechwithout user intervention. - Memory-bounded processing: By tokenizing 60-second segments separately and sampling once, the system processes hour-long audio without excessive memory growth.
- Single insertion point: Speech features enter the language model only when
cache_position[0] == 0, allowing standard autoregressive generation afterward. - Production-ready: The
VibeVoiceASRBatchInferenceclass indemo/vibevoice_asr_inference_from_file.pyprovides optimized batch processing for multiple long files.
Frequently Asked Questions
What is the maximum audio length VibeVoice-ASR can process?
VibeVoice-ASR can process audio lasting several hours in a single forward pass, limited primarily by the language model's 64,000 token context window rather than memory constraints. The acoustic encoder handles audio of any duration by processing 60-second segments sequentially, while the transcription length limit depends on the max_new_tokens parameter (default 32,768).
Does streaming mode reduce transcription accuracy compared to short-form processing?
No, streaming mode maintains full accuracy because the system concatenates mean representations from all segments before sampling once, ensuring the acoustic and semantic connectors receive complete audio statistics. The acoustic_connector and semantic_connector process the aggregated features identically to short-form inputs.
How does the system handle speaker diarization in long recordings?
The processor's post_process_transcription method extracts structured JSON containing start time, end time, speaker ID, and text for each segment. The model inserts speaker change tokens during generation, which the processor parses to identify different speakers across the timeline without requiring separate diarization passes.
Can I force streaming mode for short audio files?
While the system automatically disables streaming for audio under 60 seconds to reduce overhead, you can modify the use_streaming parameter in vibevoice/processor/vibevoice_asr_processor.py (lines 30-44) if you need consistent memory profiling or are processing batches with mixed durations. However, the single-pass path is optimized for short audio and uses less compute for inputs under the threshold.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →