How PersonaPlex Enables Full-Duplex Speech-to-Speech Conversation
PersonaPlex achieves real-time, full-duplex speech-to-speech conversation by streaming Opus-encoded audio over WebSockets, processing frames through a streaming Moshi language model on the server, and playing back responses via Web Audio worklets while simultaneously capturing microphone input.
NVIDIA's PersonaPlex repository implements a low-latency voice AI system that maintains natural, full-duplex speech-to-speech conversation without turn-taking barriers. The architecture combines a WebSocket binary transport layer, streaming neural audio codecs, and asynchronous client-side audio processing to enable simultaneous speaking and listening.
WebSocket Binary Protocol for Real-Time Audio Streaming
The transport layer uses a single WebSocket connection to exchange raw Opus packets bidirectionally between client and server.
Client-Side Frame Encoding
In client/src/protocol/encoder.ts, the encoder constructs binary frames by prefixing Opus pages with a type identifier. Audio packets are tagged with 0x01 to distinguish them from text messages:
// From client/src/protocol/encoder.ts (audio case)
case 'audio':
// Buffer layout: [0x01, ...opusData]
const buffer = new Uint8Array(1 + data.length);
buffer[0] = 0x01;
buffer.set(data, 1);
return buffer;
Server-Side Packet Processing
The server receives these messages in moshi/moshi/server.py inside the recv_loop. The first byte (kind) determines the packet type, with kind == 1 indicating audio that should be handed to the Opus reader:
# From moshi/moshi/server.py recv_loop
kind = data[0]
if kind == 1:
# Audio packet handling
opus_reader.append_chunk(data[1:])
pcm = opus_reader.decode()
Conversely, the send_loop prefixes outgoing Opus packets with 0x01 before writing them to the socket, creating a symmetrical full-duplex data channel.
Streaming Moshi Model Architecture
The server processes audio streams frame-by-frame using Moshi's Mimi encoder/decoder in continuous streaming mode, allowing inference to occur while speech is still being received.
Continuous Streaming Mode
Inside ServerState.__init__ in moshi/moshi/server.py, the model initializes with self.mimi.streaming_forever(1). This enables perpetual encode-decode cycles without resets between utterances.
The Opus Processing Loop
The opus_loop orchestrates the real-time inference pipeline:
- Pulls PCM audio from the Opus reader
- Encodes frames to VQ codes via the Mimi encoder
- Runs language model generation via
LMGen.step - Decodes generated tokens back to PCM audio using Mimi
- Re-encodes the PCM to Opus and queues it for transmission
This asynchronous loop processes incoming audio while simultaneously generating response audio, eliminating turn-taking latency. Text tokens are extracted and sent with prefix 0x02 for subtitle rendering, but the core conversation flow relies on this audio pipeline.
Low-Latency Client Audio Pipeline
The client implements parallel capture and playback using the Web Audio API and Web Workers to prevent main-thread blocking.
Recording and Transmission
The useUserAudio hook in client/src/pages/Conversation/hooks/useUserAudio.ts creates an opus-recorder instance that streams Opus pages as they are captured:
// From client/src/pages/Conversation/hooks/useUserAudio.ts
const recorder = new OpusRecorder({
streamPages: true,
ondataavailable: (typedArray: Uint8Array) => {
// Forward raw bytes to server immediately
sendMessage(typedArray);
}
});
The ondataavailable callback forwards raw bytes to the server via sendMessage in useSocket, ensuring microphone audio flows continuously without waiting for recording to stop.
Decoding and Playback
Incoming audio packets are handled by useServerAudio.ts. The onSocketMessage callback extracts the payload and forwards it to decodeAudio, which runs the Opus-to-PCM pipeline in a Web Worker:
// From client/src/pages/Conversation/hooks/useServerAudio.ts
const decodeAudio = async (opusData: Uint8Array) => {
// Post to decoder worker
decoderWorker.postMessage({ type: 'decode', data: opusData });
};
// Worker posts PCM to AudioWorkletProcessor ('moshi-processor') for playback
Pre-Warming the Decoder
To avoid first-frame initialization delays, decoderWorker.ts implements prewarmDecoderWorker, which sends a minimal Ogg BOS (Beginning of Stream) page to the decoder on page load. This triggers the WASM decoder's internal initialization before actual conversation audio arrives.
Integration Example
Implementing the full-duplex conversation requires initializing the WebSocket, starting microphone capture, and handling server audio playback:
// 1️⃣ Open a WebSocket (full‑duplex channel)
const { socket, sendMessage, start } = useSocket({
uri: "wss://your-host/api/chat",
});
start();
// 2️⃣ Start microphone capture (Opus packets go out)
const { startRecordingUser } = useUserAudio({
constraints: { audio: true },
});
startRecordingUser(); // begins sending 0x01‑audio frames
// 3️⃣ Receive and play the assistant’s speech
const {} = useServerAudio({});
// decodeAudio handles incoming packets automatically via AudioWorklet
Summary
- Binary WebSocket Protocol: Uses
0x01prefixed Opus frames inclient/src/protocol/encoder.tsandmoshi/moshi/server.pyfor bidirectional audio streaming, distinguishing audio from text (0x02) packets. - Streaming Inference: The
opus_loopinserver.pyprocesses audio through Mimi'sstreaming_forevermode andLMGen.stepto generate responses incrementally without waiting for user speech to end. - Asynchronous Client Architecture: Separate threads for recording (
useUserAudio.ts), decoding (decoderWorker.ts), and playback (moshi-processorAudioWorklet) enable simultaneous input and output. - Latency Optimization: Pre-warmed decoder workers and streaming packet transmission keep end-to-end latency under a few hundred milliseconds, creating natural conversation flow.
Frequently Asked Questions
How does PersonaPlex handle audio encoding without blocking the main thread?
The client offloads Opus decoding to a dedicated Web Worker defined in client/src/decoder/decoderWorker.ts. Incoming binary messages from the WebSocket are transferred to this worker, which runs the WASM decoder and posts PCM frames to an AudioWorkletProcessor for playback. This architecture keeps the main thread responsive while processing continuous audio streams.
What distinguishes audio packets from text messages in the WebSocket stream?
The protocol uses the first byte as a type identifier. Audio packets carry prefix 0x01 while text tokens for subtitles use 0x02. In moshi/moshi/server.py, the recv_loop inspects this kind byte to route audio to the Opus decoder and text to the subtitle handler.
Why is streaming mode critical for full-duplex conversation?
Standard turn-based systems buffer entire utterances before processing. PersonaPlex uses Mimi's streaming_forever(1) mode in server.py to process audio frames as they arrive, allowing the language model (LMGen.step) to generate response tokens incrementally. This overlapping processing eliminates the artificial pauses typical of half-duplex voice assistants.
How does the client minimize latency when starting a conversation?
The client calls prewarmDecoderWorker on page load, which sends a minimal Ogg BOS (Beginning of Stream) page to the decoder worker. This initializes the internal Opus decoder state before any real audio arrives, eliminating initialization overhead during the first user interaction.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →