# How to Process Long-Form Audio with VibeVoice-ASR: A Complete Technical Guide

> Learn to process long-form audio with VibeVoice-ASR. This guide details how to transcribe audio up to an hour in a single forward pass using advanced streaming techniques.

- Repository: [Microsoft/VibeVoice](https://github.com/microsoft/VibeVoice)
- Tags: how-to-guide
- Published: 2026-03-28

---

**VibeVoice-ASR transcribes audio up to 60 minutes (or longer) in a single forward pass by streaming acoustic and semantic tokenizers and inserting speech representations only once during the first generation step.**

The **microsoft/VibeVoice** repository provides a production-ready automatic speech recognition system designed specifically for long-form content. Unlike traditional ASR models that struggle with memory constraints on extended recordings, VibeVoice-ASR implements segment-wise streaming tokenization that keeps memory usage bounded while maintaining full contextual awareness across hour-long inputs.

## How VibeVoice-ASR Handles Long-Form Audio

VibeVoice-ASR achieves long-form capability through two architectural innovations: **streaming tokenization** and **single-pass feature insertion**. The system automatically detects audio duration and switches to a chunked processing mode when inputs exceed 60 seconds, splitting the waveform into manageable segments without losing semantic coherence.

The processor first loads audio using FFmpeg (when available) and resamples to the model's **24 kHz** target rate. In [`vibevoice/processor/vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_asr_processor.py), the `_process_single_audio` method calculates duration as `len(audio_array) / self.target_sample_rate` and sets `use_streaming=True` when the threshold is exceeded. This triggers the streaming pathway in the model's encoder.

## The Streaming Architecture

The long-form pipeline relies on three core components that work together to process extended sequences without exhausting GPU memory.

### Audio Segmentation and Tokenization

In [`vibevoice/modular/modeling_vibevoice_asr.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_asr.py), the `encode_speech` method implements the core streaming logic. The method splits incoming audio into **60-second segments** (`segment_samples = 60 × 24000` samples) and processes each through the acoustic and semantic tokenizers separately.

For each segment, the code calls `acoustic_tokenizer.encode(..., cache=..., use_cache=True, is_final_chunk=...)` to produce mean representations without sampling. After all segments complete, the system concatenates the means and performs **sampling exactly once** using the learned standard deviation. This approach prevents the model from building excessively large convolution buffers that would otherwise exhaust memory on hour-long recordings.

### Single-Pass Feature Insertion

The `prepare_inputs_for_generation` function (lines 91-106 in [`vibevoice/modular/modeling_vibevoice_asr.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/modular/modeling_vibevoice_asr.py)) implements a critical optimization: it checks `if cache_position[0] == 0` to identify the first generation step. Only on this initial pass does it insert the speech feature tensors and `acoustic_input_mask` at the `<|speech_pad|>` placeholder positions. Subsequent generation steps receive `None` for speech inputs, allowing the language model to continue autoregressively without重复处理 the acoustic features.

This design ensures that heavy acoustic and semantic encoding occurs only once per input, while the language model can generate up to **64,000 tokens** of transcription using standard causal attention.

## Processing Single Long Audio Files

To transcribe a single lengthy recording, use the `VibeVoiceASRProcessor` with automatic streaming detection:

```python
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor
from vibevoice.modular.modeling_vibevoice_asr import VibeVoiceASRForConditionalGeneration
import torch

# Initialize processor with pretrained weights

processor = VibeVoiceASRProcessor.from_pretrained(
    "microsoft/VibeVoice-ASR",
    language_model_pretrained_name="Qwen/Qwen2.5-7B"
)

# Load model with automatic device mapping

model = VibeVoiceASRForConditionalGeneration.from_pretrained(
    "microsoft/VibeVoice-ASR",
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="sdpa",
    trust_remote_code=True,
)
model.eval()

# Process audio of any length - streaming activates automatically for >60s

inputs = processor(
    audio="path/to/60_minute_podcast.wav",
    return_tensors="pt",
    padding=True,
    add_generation_prompt=True,
)

# Move to device and generate

device = next(model.parameters()).device
inputs = {k: v.to(device) if isinstance(v, torch.Tensor) else v 
          for k, v in inputs.items()}

generated_ids = model.generate(**inputs, max_new_tokens=32768)
raw_text = processor.decode(generated_ids[0], skip_special_tokens=True)

# Extract structured segments with timestamps

segments = processor.post_process_transcription(raw_text)
for seg in segments:
    print(f"[{seg.get('start_time'):.2f}s – {seg.get('end_time'):.2f}s] "
          f"Speaker {seg.get('speaker_id')}: {seg.get('text')}")

```

The processor automatically constructs the token sequence containing speech placeholders, and the model handles the segment-wise encoding internally.

## Batch Processing Multiple Files

For production workflows involving multiple long recordings, use the `VibeVoiceASRBatchInference` class from the demo scripts:

```python
from demo.vibevoice_asr_inference_from_file import VibeVoiceASRBatchInference

# Initialize inference helper

asr = VibeVoiceASRBatchInference(
    model_path="microsoft/VibeVoice-ASR",
    device="cuda",
    dtype=torch.bfloat16,
    attn_implementation="sdpa",
)

# Process multiple lengthy files with automatic batching

audio_files = [
    "meeting_recording_1.wav",
    "meeting_recording_2.mp4",  # Video files supported via FFmpeg

    "interview_session.mp3",
]

results = asr.transcribe_with_batching(
    audio_inputs=audio_files,
    batch_size=2,           # Adjust based on GPU memory

    max_new_tokens=32768,
    temperature=0.0,        # Greedy decoding for accuracy

    do_sample=False,
)

# Access structured results

for r in results:
    print(f"File: {r['file']} ({r['generation_time']:.2f}s)")
    for seg in r["segments"]:
        print(f"  [{seg.get('start_time'):.2f}s] {seg.get('text')}")

```

This implementation in [`demo/vibevoice_asr_inference_from_file.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_asr_inference_from_file.py) handles the complete pipeline including loading, streaming tokenization, and structured output generation.

## Generating Synthetic Long-Form Data for Testing

To benchmark the system or test memory constraints, concatenate existing datasets into extended sequences:

```python
from demo.vibevoice_asr_inference_from_file import load_dataset_and_concatenate

# Create a 3-hour synthetic audio from Librispeech

long_audios = load_dataset_and_concatenate(
    dataset_name="openslr/librispeech_asr",
    split="test",
    max_duration=10800,   # 3 hours in seconds

    num_audios=1,
    target_sr=24000,
)

# Transcribe the synthetic long-form input

results = asr.transcribe_with_batching(
    audio_inputs=long_audios,
    batch_size=1,
    max_new_tokens=64000,  # Extended context for very long inputs

)

```

The `load_dataset_and_concatenate` function demonstrates how the streaming pipeline handles arbitrarily concatenated audio while maintaining accurate timestamps across segment boundaries.

## Summary

- **Automatic streaming**: The processor detects audio >60 seconds and activates the streaming pathway in `encode_speech` without user intervention.
- **Memory-bounded processing**: By tokenizing 60-second segments separately and sampling once, the system processes hour-long audio without excessive memory growth.
- **Single insertion point**: Speech features enter the language model only when `cache_position[0] == 0`, allowing standard autoregressive generation afterward.
- **Production-ready**: The `VibeVoiceASRBatchInference` class in [`demo/vibevoice_asr_inference_from_file.py`](https://github.com/microsoft/VibeVoice/blob/main/demo/vibevoice_asr_inference_from_file.py) provides optimized batch processing for multiple long files.

## Frequently Asked Questions

### What is the maximum audio length VibeVoice-ASR can process?

VibeVoice-ASR can process audio lasting **several hours** in a single forward pass, limited primarily by the language model's **64,000 token context window** rather than memory constraints. The acoustic encoder handles audio of any duration by processing 60-second segments sequentially, while the transcription length limit depends on the `max_new_tokens` parameter (default 32,768).

### Does streaming mode reduce transcription accuracy compared to short-form processing?

No, streaming mode maintains full accuracy because the system concatenates mean representations from all segments before sampling once, ensuring the acoustic and semantic connectors receive complete audio statistics. The `acoustic_connector` and `semantic_connector` process the aggregated features identically to short-form inputs.

### How does the system handle speaker diarization in long recordings?

The processor's `post_process_transcription` method extracts structured JSON containing **start time**, **end time**, **speaker ID**, and **text** for each segment. The model inserts speaker change tokens during generation, which the processor parses to identify different speakers across the timeline without requiring separate diarization passes.

### Can I force streaming mode for short audio files?

While the system automatically disables streaming for audio under 60 seconds to reduce overhead, you can modify the `use_streaming` parameter in [`vibevoice/processor/vibevoice_asr_processor.py`](https://github.com/microsoft/VibeVoice/blob/main/vibevoice/processor/vibevoice_asr_processor.py) (lines 30-44) if you need consistent memory profiling or are processing batches with mixed durations. However, the single-pass path is optimized for short audio and uses less compute for inputs under the threshold.