The Three Paradigms of Voice Agent Implementation: Cascaded, End-to-End, and Full-Duplex

Voice agents are built using three distinct architectural paradigms—cascaded pipelines, end-to-end full-modal models, and full-duplex streaming systems—each trading off between modularity, latency, and implementation complexity.

The bojieli/ai-agent-book repository documents these three paradigms of voice agent implementation in Chapter 6, providing concrete code experiments that demonstrate how speech recognition, reasoning, and synthesis can be orchestrated in fundamentally different ways. Whether you are optimizing for low-latency conversation or modular component upgrades, understanding these architectural patterns is essential for building production voice AI systems.

Cascaded (Pipeline) Paradigm

The cascaded paradigm treats Automatic Speech Recognition (ASR), Large Language Model (LLM) reasoning, and Text-to-Speech (TTS) as three separate, loosely-coupled components executed sequentially.

Architecture and Trade-offs

In this approach, each stage waits for the previous one to complete before starting, creating a linear data flow: audio input is transcribed to text, the text is processed by an LLM to generate a response, and that response is synthesized back into speech. According to the repository's chapter6/README.en.md, this paradigm leverages "mature, specialised models for each step" and offers the advantage that individual modules can be easily swapped or upgraded without affecting the others【/cache/repos/github.com/bojieli/ai-agent-book/main/chapter6/README.en.md†L3-L4】.

The primary weakness is end-to-end latency accumulation. Since each stage adds its own processing time, the total round-trip delay equals the sum of ASR, LLM inference, and TTS latencies. Additionally, maintaining conversational context across these discrete modules requires careful orchestration to prevent context loss between hand-offs.

Implementation Example

The chapter6/live-audio/ directory contains a working demonstration of this pipeline using Whisper for ASR and a configurable LLM for reasoning【/cache/repos/github.com/bojieli/ai-agent-book/main/chapter6/README.en.md†L21-L24】:


# 1️⃣ ASR stage

audio = record_microphone()
text = asr_model.transcribe(audio)          # Separate ASR model (e.g., Whisper)

# 2️⃣ LLM reasoning stage

reply = llm_chat.invoke(text)               # Separate LLM model

# 3️⃣ TTS synthesis stage

speech = tts_model.synthesize(reply)        # Separate TTS model (e.g., Tacotron-2)

play_audio(speech)

This orchestration is typically managed by a lightweight controller—often a FastAPI endpoint or simple script—that handles the serialization between specialized models like DeepSpeech for ASR or Vocos for TTS synthesis.

End-to-End Full-Modal Paradigm

The end-to-end full-modal paradigm collapses the entire voice agent pipeline into a single multimodal model that jointly learns speech recognition, reasoning, and speech synthesis.

Monolithic Architecture

Rather than chaining discrete components, this approach employs one monolithic model—such as MiniCPM-o or Qwen-VL—that receives raw audio input and directly outputs spoken responses in a single forward pass【/cache/repos/github.com/bojieli/ai-agent-book/main/chapter6/README.en.md†L25-L27】. The model learns a joint representation of speech and language, eliminating hand-off errors between modules and minimizing latency since the entire utterance can be processed at once without intermediate text representations.

The trade-off is debugging complexity and data requirements. Errors are internal to the model rather than isolated to specific pipeline stages, making troubleshooting more difficult. Additionally, training requires large volumes of diverse multimodal paired data covering both speech and language domains.

Implementation Example

The chapter6/end-to-end-speech/ experiment demonstrates this approach using MiniCPM-o:


# Single multimodal model handles perception, reasoning, and synthesis

audio = record_microphone()
speech = multimodal_model.run(audio)   # Input audio → output audio directly

play_audio(speech)

This paradigm eliminates explicit transcription and synthesis stages, with the model internally managing the conversion between modalities through learned cross-modal representations.

Full-Duplex (Streaming) Paradigm

The full-duplex paradigm enables simultaneous perception and synthesis, allowing the agent to speak while still listening through incremental, streaming processing.

Streaming Architecture

In this architecture, input audio is chunked (typically into 200ms segments) and fed to a streaming ASR module that emits partial transcripts as soon as possible. The LLM receives these growing transcripts and begins generating tokens before the user finishes speaking, streaming output to a low-latency TTS engine that starts synthesis immediately rather than waiting for the complete response【/cache/repos/github.com/bojieli/ai-agent-book/main/chapter6/README.en.md†L25-L26】.

This approach requires sophisticated handling of partial hypotheses, confidence thresholds, and correction mechanisms when the ASR updates its transcription (e.g., handling user self-corrections like "I meant..."). The chapter6/streaming-speech/ directory contains the concrete implementation of these streaming perception techniques.

Implementation Example

stream = microphone_stream()
partial_text = ""

for chunk in stream:
    # Incremental transcription with partial results

    partial_text += streaming_asr.transcribe(chunk, finalize=False)
    
    # Early generation trigger based on token confidence

    if len(partial_text) > MIN_TOKENS:
        partial_reply = llm_chat.stream_generate(partial_text)
        
        # Stream tokens to TTS as they arrive

        for token in partial_reply:
            tts_stream.add_token(token)

# Flush remaining audio buffer when user stops speaking

tts_stream.finalize()
play_audio(tts_stream.output())

The system continuously updates both perception and synthesis streams, creating the most conversational feel but requiring complex buffering and synchronization logic to manage competing audio streams.

Summary

  • Cascaded paradigm offers maximum modularity by separating ASR, LLM, and TTS into discrete components, making it ideal for systems requiring component specialization or frequent upgrades, though at the cost of accumulated latency.

  • End-to-end full-modal paradigm minimizes latency and hand-off errors by processing speech through a single multimodal model, but requires significant training data and makes debugging more challenging due to internalized error states.

  • Full-duplex streaming paradigm achieves the lowest perceived latency by processing audio chunks incrementally and synthesizing responses while the user is still speaking, though it demands sophisticated engineering for stream synchronization and hypothesis correction.

Frequently Asked Questions

What is the main latency difference between cascaded and end-to-end voice agents?

Cascaded systems accumulate latency sequentially across ASR, LLM, and TTS stages, resulting in a total round-trip time equal to the sum of all three components. End-to-end systems minimize latency by processing the entire utterance in a single forward pass through a multimodal model, eliminating inter-module hand-off delays.

When should I use the full-duplex streaming paradigm over the cascaded approach?

Choose the full-duplex paradigm when perceived latency is critical for user experience, such as in real-time conversational agents where users expect immediate vocal responses. However, this requires accepting increased engineering complexity for handling partial ASR hypotheses and maintaining synchronized audio streams, whereas the cascaded approach is simpler to implement and debug.

Can I combine elements from different voice agent paradigms?

While the paradigms represent distinct architectural philosophies, hybrid approaches exist—such as using a cascaded pipeline with streaming ASR to reduce the first-stage latency. However, as documented in chapter6/README.en.md, each paradigm fundamentally differs in how it balances modularity against latency, so mixing them requires careful consideration of where the boundaries between streaming and batch processing occur.

Which paradigm does the ai-agent-book repository recommend for beginners?

The repository's chapter6/live-audio/ implementation provides the most accessible starting point, demonstrating a cascaded pipeline using Whisper and standard LLM APIs. This approach is recommended for beginners because it allows independent debugging of ASR, reasoning, and synthesis components before attempting the more complex engineering required for full-duplex streaming or the specialized training needed for end-to-end models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →