# How PersonaPlex Enables Full-Duplex Speech-to-Speech Conversation

> Discover how PersonaPlex enables full-duplex speech-to-speech conversation with real-time audio streaming, server-side AI processing, and simultaneous playback and capture.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: deep-dive
- Published: 2026-04-07

---

**PersonaPlex achieves real-time, full-duplex speech-to-speech conversation by streaming Opus-encoded audio over WebSockets, processing frames through a streaming Moshi language model on the server, and playing back responses via Web Audio worklets while simultaneously capturing microphone input.**

NVIDIA's PersonaPlex repository implements a low-latency voice AI system that maintains natural, full-duplex speech-to-speech conversation without turn-taking barriers. The architecture combines a WebSocket binary transport layer, streaming neural audio codecs, and asynchronous client-side audio processing to enable simultaneous speaking and listening.

## WebSocket Binary Protocol for Real-Time Audio Streaming

The transport layer uses a single WebSocket connection to exchange raw Opus packets bidirectionally between client and server.

### Client-Side Frame Encoding

In [`client/src/protocol/encoder.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/protocol/encoder.ts), the encoder constructs binary frames by prefixing Opus pages with a type identifier. Audio packets are tagged with `0x01` to distinguish them from text messages:

```typescript
// From client/src/protocol/encoder.ts (audio case)
case 'audio':
  // Buffer layout: [0x01, ...opusData]
  const buffer = new Uint8Array(1 + data.length);
  buffer[0] = 0x01;
  buffer.set(data, 1);
  return buffer;

```

### Server-Side Packet Processing

The server receives these messages in [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py) inside the `recv_loop`. The first byte (`kind`) determines the packet type, with `kind == 1` indicating audio that should be handed to the Opus reader:

```python

# From moshi/moshi/server.py recv_loop

kind = data[0]
if kind == 1:
    # Audio packet handling

    opus_reader.append_chunk(data[1:])
    pcm = opus_reader.decode()

```

Conversely, the `send_loop` prefixes outgoing Opus packets with `0x01` before writing them to the socket, creating a symmetrical full-duplex data channel.

## Streaming Moshi Model Architecture

The server processes audio streams frame-by-frame using Moshi's **Mimi** encoder/decoder in continuous streaming mode, allowing inference to occur while speech is still being received.

### Continuous Streaming Mode

Inside `ServerState.__init__` in [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py), the model initializes with `self.mimi.streaming_forever(1)`. This enables perpetual encode-decode cycles without resets between utterances.

### The Opus Processing Loop

The `opus_loop` orchestrates the real-time inference pipeline:

1. Pulls PCM audio from the Opus reader
2. Encodes frames to VQ codes via the Mimi encoder
3. Runs language model generation via `LMGen.step`
4. Decodes generated tokens back to PCM audio using Mimi
5. Re-encodes the PCM to Opus and queues it for transmission

This asynchronous loop processes incoming audio while simultaneously generating response audio, eliminating turn-taking latency. Text tokens are extracted and sent with prefix `0x02` for subtitle rendering, but the core conversation flow relies on this audio pipeline.

## Low-Latency Client Audio Pipeline

The client implements parallel capture and playback using the Web Audio API and Web Workers to prevent main-thread blocking.

### Recording and Transmission

The `useUserAudio` hook in [`client/src/pages/Conversation/hooks/useUserAudio.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/pages/Conversation/hooks/useUserAudio.ts) creates an opus-recorder instance that streams Opus pages as they are captured:

```typescript
// From client/src/pages/Conversation/hooks/useUserAudio.ts
const recorder = new OpusRecorder({
  streamPages: true,
  ondataavailable: (typedArray: Uint8Array) => {
    // Forward raw bytes to server immediately
    sendMessage(typedArray);
  }
});

```

The `ondataavailable` callback forwards raw bytes to the server via `sendMessage` in `useSocket`, ensuring microphone audio flows continuously without waiting for recording to stop.

### Decoding and Playback

Incoming audio packets are handled by [`useServerAudio.ts`](https://github.com/NVIDIA/personaplex/blob/main/useServerAudio.ts). The `onSocketMessage` callback extracts the payload and forwards it to `decodeAudio`, which runs the Opus-to-PCM pipeline in a Web Worker:

```typescript
// From client/src/pages/Conversation/hooks/useServerAudio.ts
const decodeAudio = async (opusData: Uint8Array) => {
  // Post to decoder worker
  decoderWorker.postMessage({ type: 'decode', data: opusData });
};

// Worker posts PCM to AudioWorkletProcessor ('moshi-processor') for playback

```

### Pre-Warming the Decoder

To avoid first-frame initialization delays, [`decoderWorker.ts`](https://github.com/NVIDIA/personaplex/blob/main/decoderWorker.ts) implements `prewarmDecoderWorker`, which sends a minimal Ogg BOS (Beginning of Stream) page to the decoder on page load. This triggers the WASM decoder's internal initialization before actual conversation audio arrives.

## Integration Example

Implementing the full-duplex conversation requires initializing the WebSocket, starting microphone capture, and handling server audio playback:

```typescript
// 1️⃣ Open a WebSocket (full‑duplex channel)
const { socket, sendMessage, start } = useSocket({
  uri: "wss://your-host/api/chat",
});
start();

// 2️⃣ Start microphone capture (Opus packets go out)
const { startRecordingUser } = useUserAudio({
  constraints: { audio: true },
});
startRecordingUser();  // begins sending 0x01‑audio frames

// 3️⃣ Receive and play the assistant’s speech
const {} = useServerAudio({});
// decodeAudio handles incoming packets automatically via AudioWorklet

```

## Summary

- **Binary WebSocket Protocol**: Uses `0x01` prefixed Opus frames in [`client/src/protocol/encoder.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/protocol/encoder.ts) and [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py) for bidirectional audio streaming, distinguishing audio from text (`0x02`) packets.
- **Streaming Inference**: The `opus_loop` in [`server.py`](https://github.com/NVIDIA/personaplex/blob/main/server.py) processes audio through Mimi's `streaming_forever` mode and `LMGen.step` to generate responses incrementally without waiting for user speech to end.
- **Asynchronous Client Architecture**: Separate threads for recording ([`useUserAudio.ts`](https://github.com/NVIDIA/personaplex/blob/main/useUserAudio.ts)), decoding ([`decoderWorker.ts`](https://github.com/NVIDIA/personaplex/blob/main/decoderWorker.ts)), and playback (`moshi-processor` AudioWorklet) enable simultaneous input and output.
- **Latency Optimization**: Pre-warmed decoder workers and streaming packet transmission keep end-to-end latency under a few hundred milliseconds, creating natural conversation flow.

## Frequently Asked Questions

### How does PersonaPlex handle audio encoding without blocking the main thread?

The client offloads Opus decoding to a dedicated Web Worker defined in [`client/src/decoder/decoderWorker.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/decoder/decoderWorker.ts). Incoming binary messages from the WebSocket are transferred to this worker, which runs the WASM decoder and posts PCM frames to an `AudioWorkletProcessor` for playback. This architecture keeps the main thread responsive while processing continuous audio streams.

### What distinguishes audio packets from text messages in the WebSocket stream?

The protocol uses the first byte as a type identifier. Audio packets carry prefix `0x01` while text tokens for subtitles use `0x02`. In [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py), the `recv_loop` inspects this `kind` byte to route audio to the Opus decoder and text to the subtitle handler.

### Why is streaming mode critical for full-duplex conversation?

Standard turn-based systems buffer entire utterances before processing. PersonaPlex uses Mimi's `streaming_forever(1)` mode in [`server.py`](https://github.com/NVIDIA/personaplex/blob/main/server.py) to process audio frames as they arrive, allowing the language model (`LMGen.step`) to generate response tokens incrementally. This overlapping processing eliminates the artificial pauses typical of half-duplex voice assistants.

### How does the client minimize latency when starting a conversation?

The client calls `prewarmDecoderWorker` on page load, which sends a minimal Ogg BOS (Beginning of Stream) page to the decoder worker. This initializes the internal Opus decoder state before any real audio arrives, eliminating initialization overhead during the first user interaction.