How PersonaPlex Enables Full-Duplex Speech-to-Speech Conversation

PersonaPlex achieves real-time, full-duplex speech-to-speech conversation by streaming Opus-encoded audio over WebSockets, processing frames through a streaming Moshi language model on the server, and playing back responses via Web Audio worklets while simultaneously capturing microphone input.

NVIDIA's PersonaPlex repository implements a low-latency voice AI system that maintains natural, full-duplex speech-to-speech conversation without turn-taking barriers. The architecture combines a WebSocket binary transport layer, streaming neural audio codecs, and asynchronous client-side audio processing to enable simultaneous speaking and listening.

WebSocket Binary Protocol for Real-Time Audio Streaming

The transport layer uses a single WebSocket connection to exchange raw Opus packets bidirectionally between client and server.

Client-Side Frame Encoding

In client/src/protocol/encoder.ts, the encoder constructs binary frames by prefixing Opus pages with a type identifier. Audio packets are tagged with 0x01 to distinguish them from text messages:

// From client/src/protocol/encoder.ts (audio case)
case 'audio':
  // Buffer layout: [0x01, ...opusData]
  const buffer = new Uint8Array(1 + data.length);
  buffer[0] = 0x01;
  buffer.set(data, 1);
  return buffer;

Server-Side Packet Processing

The server receives these messages in moshi/moshi/server.py inside the recv_loop. The first byte (kind) determines the packet type, with kind == 1 indicating audio that should be handed to the Opus reader:


# From moshi/moshi/server.py recv_loop

kind = data[0]
if kind == 1:
    # Audio packet handling

    opus_reader.append_chunk(data[1:])
    pcm = opus_reader.decode()

Conversely, the send_loop prefixes outgoing Opus packets with 0x01 before writing them to the socket, creating a symmetrical full-duplex data channel.

Streaming Moshi Model Architecture

The server processes audio streams frame-by-frame using Moshi's Mimi encoder/decoder in continuous streaming mode, allowing inference to occur while speech is still being received.

Continuous Streaming Mode

Inside ServerState.__init__ in moshi/moshi/server.py, the model initializes with self.mimi.streaming_forever(1). This enables perpetual encode-decode cycles without resets between utterances.

The Opus Processing Loop

The opus_loop orchestrates the real-time inference pipeline:

  1. Pulls PCM audio from the Opus reader
  2. Encodes frames to VQ codes via the Mimi encoder
  3. Runs language model generation via LMGen.step
  4. Decodes generated tokens back to PCM audio using Mimi
  5. Re-encodes the PCM to Opus and queues it for transmission

This asynchronous loop processes incoming audio while simultaneously generating response audio, eliminating turn-taking latency. Text tokens are extracted and sent with prefix 0x02 for subtitle rendering, but the core conversation flow relies on this audio pipeline.

Low-Latency Client Audio Pipeline

The client implements parallel capture and playback using the Web Audio API and Web Workers to prevent main-thread blocking.

Recording and Transmission

The useUserAudio hook in client/src/pages/Conversation/hooks/useUserAudio.ts creates an opus-recorder instance that streams Opus pages as they are captured:

// From client/src/pages/Conversation/hooks/useUserAudio.ts
const recorder = new OpusRecorder({
  streamPages: true,
  ondataavailable: (typedArray: Uint8Array) => {
    // Forward raw bytes to server immediately
    sendMessage(typedArray);
  }
});

The ondataavailable callback forwards raw bytes to the server via sendMessage in useSocket, ensuring microphone audio flows continuously without waiting for recording to stop.

Decoding and Playback

Incoming audio packets are handled by useServerAudio.ts. The onSocketMessage callback extracts the payload and forwards it to decodeAudio, which runs the Opus-to-PCM pipeline in a Web Worker:

// From client/src/pages/Conversation/hooks/useServerAudio.ts
const decodeAudio = async (opusData: Uint8Array) => {
  // Post to decoder worker
  decoderWorker.postMessage({ type: 'decode', data: opusData });
};

// Worker posts PCM to AudioWorkletProcessor ('moshi-processor') for playback

Pre-Warming the Decoder

To avoid first-frame initialization delays, decoderWorker.ts implements prewarmDecoderWorker, which sends a minimal Ogg BOS (Beginning of Stream) page to the decoder on page load. This triggers the WASM decoder's internal initialization before actual conversation audio arrives.

Integration Example

Implementing the full-duplex conversation requires initializing the WebSocket, starting microphone capture, and handling server audio playback:

// 1️⃣ Open a WebSocket (full‑duplex channel)
const { socket, sendMessage, start } = useSocket({
  uri: "wss://your-host/api/chat",
});
start();

// 2️⃣ Start microphone capture (Opus packets go out)
const { startRecordingUser } = useUserAudio({
  constraints: { audio: true },
});
startRecordingUser();  // begins sending 0x01‑audio frames

// 3️⃣ Receive and play the assistant’s speech
const {} = useServerAudio({});
// decodeAudio handles incoming packets automatically via AudioWorklet

Summary

  • Binary WebSocket Protocol: Uses 0x01 prefixed Opus frames in client/src/protocol/encoder.ts and moshi/moshi/server.py for bidirectional audio streaming, distinguishing audio from text (0x02) packets.
  • Streaming Inference: The opus_loop in server.py processes audio through Mimi's streaming_forever mode and LMGen.step to generate responses incrementally without waiting for user speech to end.
  • Asynchronous Client Architecture: Separate threads for recording (useUserAudio.ts), decoding (decoderWorker.ts), and playback (moshi-processor AudioWorklet) enable simultaneous input and output.
  • Latency Optimization: Pre-warmed decoder workers and streaming packet transmission keep end-to-end latency under a few hundred milliseconds, creating natural conversation flow.

Frequently Asked Questions

How does PersonaPlex handle audio encoding without blocking the main thread?

The client offloads Opus decoding to a dedicated Web Worker defined in client/src/decoder/decoderWorker.ts. Incoming binary messages from the WebSocket are transferred to this worker, which runs the WASM decoder and posts PCM frames to an AudioWorkletProcessor for playback. This architecture keeps the main thread responsive while processing continuous audio streams.

What distinguishes audio packets from text messages in the WebSocket stream?

The protocol uses the first byte as a type identifier. Audio packets carry prefix 0x01 while text tokens for subtitles use 0x02. In moshi/moshi/server.py, the recv_loop inspects this kind byte to route audio to the Opus decoder and text to the subtitle handler.

Why is streaming mode critical for full-duplex conversation?

Standard turn-based systems buffer entire utterances before processing. PersonaPlex uses Mimi's streaming_forever(1) mode in server.py to process audio frames as they arrive, allowing the language model (LMGen.step) to generate response tokens incrementally. This overlapping processing eliminates the artificial pauses typical of half-duplex voice assistants.

How does the client minimize latency when starting a conversation?

The client calls prewarmDecoderWorker on page load, which sends a minimal Ogg BOS (Beginning of Stream) page to the decoder worker. This initializes the internal Opus decoder state before any real audio arrives, eliminating initialization overhead during the first user interaction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →