# Building Streaming Audio and Video Generation Pipelines: A Technical Guide to Real-Time AI Media

> Learn to build real-time AI media generation pipelines. Explore streaming audio and video techniques using neural codecs and transformer models for low latency.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Streaming audio and video generation pipelines use neural codecs to compress media into discrete token streams (50–75 Hz for audio, 5–10 Hz for video), which transformer or diffusion models process in real-time to achieve end-to-end latencies of 150–500 ms.**

The `rohitg00/ai-engineering-from-scratch` repository provides a modular curriculum that implements these architectures through production-ready lessons on full-duplex speech synthesis and generative media. By combining neural audio codecs like Mimi with unified transformer models such as Moshi, developers can build streaming pipelines that consume and emit audio continuously without traditional pipeline bottlenecks.

## Core Architecture of Streaming Media Pipelines

Production streaming systems rely on a **stream-centric architecture** that minimizes latency between input and output. According to the repository's implementation guides, this stack consists of four primary components working in concert.

### Neural Codecs and Tokenization

The foundation of any streaming pipeline is the **neural codec**, which compresses raw waveforms into discrete tokens at rates between 50–75 Hz. In [`phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/docs/en.md), the curriculum specifies using **Mimi** or **EnCodec** to encode audio into Residual VQ (RVQ) indices across multiple codebooks—typically 8 codebooks at 12.5 Hz. This compression achieves **< 20 ms per frame** encoding latency, making it feasible for real-time applications.

### The Inner-Monologue Pattern (Moshi)

The Moshi architecture introduces an **inner-monologue text stream** that runs parallel to audio tokens. This stream emits text tokens every 80 ms, providing semantic alignment and free transcript generation while the model processes audio. As implemented in the repository's full-duplex examples, this pattern enables the model to maintain context across long conversations without waiting for sentence boundaries.

## Implementing Full-Duplex Speech-to-Speech Streaming

The repository's streaming speech-to-speech lesson demonstrates a **unified model** that achieves full-duplex dialogue with approximately **200 ms latency on a single L4 GPU**—half the latency of traditional pipelines (VAD → ASR → LLM → TTS).

### Architectural Flow

The system processes two simultaneous Mimi streams: user audio input and model-generated audio output. A **Temporal Transformer** consumes the latest tokens from both streams plus the previous inner-monologue text token, then predicts:

1. A new text token (inner monologue)
2. The next 8 Mimi tokens via a depth transformer (sequential per-codebook prediction)

Transport occurs over **WebSocket**, streaming 80 ms chunks bidirectionally to maintain real-time synchronization.

### Python Implementation

The [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py) file in the streaming S2S lesson provides this asyncio implementation that wires microphone input, WebSocket transport, and speaker output:

```python
import asyncio, websockets
from moshi.client_utils import encode_audio_mimi, decode_audio_mimi

async def moshi_chat():
    async with websockets.connect("ws://localhost:8998/api/chat") as ws:
        mic_task = asyncio.create_task(stream_mic_to(ws))
        spk_task = asyncio.create_task(stream_from_to_speaker(ws))
        await asyncio.gather(mic_task, spk_task)

async def stream_mic_to(ws):
    async for chunk_80ms in mic_stream_at_12_5_hz():
        mimi_tokens = encode_audio_mimi(chunk_80ms)
        await ws.send(serialize(mimi_tokens))

async def stream_from_to_speaker(ws):
    async for msg in ws:
        mimi_tokens, text_token = deserialize(msg)
        audio = decode_audio_mimi(mimi_tokens)
        await play(audio)

```

**Key constraint:** The generator must sustain **≥ 75 tokens per second** per active user (TPOT - Time Per Output Token) to prevent buffer underruns during streaming.

## Audio Generation: Token-AR vs Diffusion

For non-conversational audio generation (TTS, music, sound effects), the curriculum in [`phases/08-generative-ai/11-audio-generation/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/11-audio-generation/docs/en.md) describes two distinct approaches with different latency trade-offs.

### Token Autoregressive (AR) Generation

**Token-AR generators** use decoder-only transformers to predict the next codec token conditioned on style or text prompts. This method streams naturally at approximately **75 tokens per second**, making it ideal for live voice agents. The repository provides this toy example in the lesson code demonstrating deterministic token patterns:

```python
import numpy as np

def synth_tokens(style, length, vocab=1024):
    """Generate deterministic token patterns for demo purposes."""
    if style == 0:                      # speech-like alternating pattern

        return [i % vocab for i in range(length)]
    # music-like monotonic ramp

    return [(i * 7) % vocab for i in range(length)]

tokens = synth_tokens(style=1, length=200)
print("First 10 tokens:", tokens[:10])

```

### Flow-Matching Diffusion

**Flow-matching diffusion** (as used in Stable Audio 2.5 or AudioLDM 2) yields higher fidelity for music generation but incurs **fixed-clip latency** measured in seconds. To adapt diffusion for streaming, the pipeline must implement chunking strategies that break generation into overlapping windows, though this introduces complexity compared to native AR streaming.

## Extending to Video and Multimodal Streams

The curriculum addresses **video generation** in [`phases/08-generative-ai/10-video-generation/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/10-video-generation/docs/en.md) and multimodal grounding in Phase 12, applying identical streaming principles to visual media.

| Modality | Codec / Token Rate | Generator | Typical Streaming Latency |
|----------|-------------------|-----------|---------------------------|
| **Audio** | 50-75 Hz tokens (Mimi, EnCodec) | AR Transformer or Flow-Matching | 150-250 ms (AR) / seconds (diffusion) |
| **Video** | 3-D patch tokens (≈ 5-10 Hz) | DiT / Temporal Transformer | 300-500 ms per frame (real-time) |
| **Multimodal (Audio + Video)** | Synchronized token streams | Unified transformer (Moshi-style) | 200-400 ms end-to-end |

A production-ready unified pipeline ingests raw media through modality-specific codecs, fuses streams in a single transformer that attends across modalities (following the *Mio Any-to-Any Streaming* pattern), then decodes each stream separately through audio vocoders and video frame renderers.

## Key Files and Implementation Resources

The repository organizes content into documentation, executable code, and production artifacts (skills). These specific files contain the implementation details referenced above:

- **[`phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/docs/en.md)** – Architecture overview, latency benchmarks, and full-duplex dialogue concepts.
- **[`phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/code/main.py)** – Minimal symbolic simulation of the full-duplex asyncio loop.
- **[`phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/outputs/skill-duplex-pipeline.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/outputs/skill-duplex-pipeline.md)** – Decision framework for choosing between full-duplex and traditional pipeline architectures.
- **[`phases/08-generative-ai/11-audio-generation/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/11-audio-generation/docs/en.md)** – Codec-token pipeline specifications and generator selection criteria.
- **[`phases/08-generative-ai/11-audio-generation/outputs/skill-audio-brief.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/11-audio-generation/outputs/skill-audio-brief.md)** – Production planning template for audio generation projects.
- **[`phases/08-generative-ai/10-video-generation/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/10-video-generation/docs/en.md)** – 3-D patch tokenization and temporal DiT architecture for video streaming.
- **[`phases/08-generative-ai/10-video-generation/outputs/skill-video-brief.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/10-video-generation/outputs/skill-video-brief.md)** – Model-prompt-pipeline planning for video tasks.

## Summary

- **Neural codecs** (Mimi, EnCodec) compress audio to 50–75 Hz token streams with < 20 ms latency, enabling real-time processing.
- **Full-duplex streaming** via Moshi achieves 200 ms end-to-end latency on L4 GPUs by eliminating traditional pipeline stages and using an inner-monologue text stream for semantic coherence.
- **Token-AR generators** sustain 75+ tokens per second for live streaming, while diffusion models require chunking strategies for real-time video.
- **Unified multimodal pipelines** fuse audio and video token streams through transformers like Mio, decoding separately through vocoders and frame renderers served over WebSocket.
- The `ai-engineering-from-scratch` repository provides runnable Python implementations, architectural documentation, and production skills for each component.

## Frequently Asked Questions

### What is the minimum GPU requirement for running a full-duplex streaming pipeline?

According to the repository's benchmarks in the Moshi lesson, a single **L4 GPU** can sustain the 200 ms latency target for full-duplex speech-to-speech generation. However, the generator must maintain a throughput of at least **75 tokens per second** per active user to prevent audio dropouts, so GPU selection should account for concurrent user load.

### How does the inner-monologue pattern improve streaming quality?

The **inner-monologue** stream emits text tokens every 80 ms alongside audio tokens, providing a semantic "backbone" that keeps the model's audio generation coherent over long conversations. As implemented in [`phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/06-speech-and-audio/15-streaming-speech-to-speech-moshi-hibiki/docs/en.md), this pattern also generates free transcripts and creates a clean integration point for upgrading the language model component without disrupting the audio stream.

### Why use token-autoregressive models instead of diffusion for real-time audio?

**Token-AR models** naturally stream their output one token at a time (≈ 75 tokens/second), matching the consumption rate of neural audio codecs. **Diffusion models** (flow-matching) generate entire clips in parallel, resulting in seconds-scale latency that requires complex chunking and windowing strategies to adapt for real-time use, making AR preferable for voice agents.

### Can these streaming principles apply to video generation?

Yes. The repository's video generation lesson ([`phases/08-generative-ai/10-video-generation/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/08-generative-ai/10-video-generation/docs/en.md)) applies identical principles using **3-D patch tokenization** at 5–10 Hz and temporal diffusion transformers (DiT). Real-time video streaming requires **300–500 ms per frame** and uses the same WebSocket transport architecture, though bandwidth requirements are significantly higher than audio-only streams.