# PersonaPlex Offline Evaluation vs Live WebSocket Server: Architecture and Usage

> PersonaPlex offers offline evaluation via CLI and a live WebSocket server for real-time streaming. Explore their architecture and usage to choose the best fit for your needs.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: architecture
- Published: 2026-04-07

---

**PersonaPlex supports both offline evaluation through the `moshi.offline` CLI module for batch processing and a live WebSocket server via `moshi.server` for real-time audio streaming, sharing the same Mimi encoder and Moshi LM pipeline under the hood.**

The NVIDIA/personaplex repository provides a multimodal conversational AI framework that can operate in two distinct execution modes. Whether you need reproducible batch processing for automated testing or low-latency duplex audio for interactive demos, understanding the differences between offline evaluation and the live WebSocket server is essential for optimal deployment.

## Core Architectural Differences

The fundamental distinction lies in how data enters and exits the system. **Offline evaluation** runs as a self-contained Python process that reads local WAV files and writes results to disk, while the **live WebSocket server** exposes an HTTP endpoint that accepts real-time binary audio streams from remote clients.

### Entry Points and Execution Models

You initiate offline evaluation by running `python -m moshi.offline`, which parses command-line arguments such as `--input-wav`, `--voice-prompt`, and `--output-wav` before executing the inference pipeline synchronously. This mode requires no network configuration and processes the entire audio file from start to finish within a single process.

In contrast, the live server starts with `python -m moshi.server`, spinning up an `aiohttp`-based HTTP service that listens for WebSocket connections on `/api/chat`. The server maintains persistent connections with clients (typically the React frontend) and processes audio frame-by-frame as binary Opus-encoded packets arrive over the network.

### Data Flow Comparison

The offline pipeline in [`moshi/moshi/offline.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/offline.py) loads user audio from disk, iterates through frames using `lm_iterate_audio`, and concatenates generated PCM frames before writing the final WAV file. The server implementation in [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py) follows a continuous loop: receiving binary messages (type `0x01` for audio), decoding them through an `opus_reader`, encoding with Mimi, generating responses via `LMGen.step`, and streaming back Opus packets via `ws.send_bytes`.

## Offline Evaluation Mode Deep Dive

Offline evaluation is ideal for automated testing, batch processing, and environments where network services are undesirable. The entire lifecycle completes without external dependencies.

### Pipeline Architecture in [`offline.py`](https://github.com/NVIDIA/personaplex/blob/main/offline.py)

The evaluation script follows a strict linear progression through seven phases:

1.  **CLI Parsing** – Validates arguments including `--input-wav`, `--voice-prompt`, and `--text-prompt`
2.  **Model Loading** – Downloads or loads from cache the Mimi encoder/decoder, Moshi LM, and SentencePiece tokenizer
3.  **Warmup** – Calls `warmup(mimi, other_mimi, lm_gen, device, frame_size)` to compile CUDA graphs with dummy zero frames
4.  **Prompt Injection** – Uses `LMGen.step_system_prompts` to inject system text tokens and voice-prompt audio
5.  **Streaming Simulation** – Processes the input WAV frame-by-frame via `lm_iterate_audio`, encoding each frame and feeding it to the LM
6.  **Output Generation** – Decodes agent audio tokens, trims/pads to match original duration, and writes to disk using `sphn.write_wav`
7.  **Transcript Extraction** – Collects text tokens and saves them as JSON alongside the audio output

### Voice Prompt Handling

Both modes utilize the `_get_voice_prompt_dir` helper to locate the `voices.tgz` archive. In offline mode, this occurs immediately before processing begins, loading the requested `.pt` file (such as `NATM1.pt`) directly from disk into memory for injection into the prompt phase.

## Live WebSocket Server Mode Deep Dive

The server mode transforms PersonaPlex into a real-time conversational service accessible via browser or custom clients.

### Server Architecture in [`server.py`](https://github.com/NVIDIA/personaplex/blob/main/server.py)

The [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py) implementation manages concurrent connections through an async event loop:

- **Startup Phase** – Downloads voice prompts, loads model components identical to the offline script, and executes `state.warmup()` to prime CUDA graphs and streaming state
- **HTTP Routing** – Registers `/api/chat` to route WebSocket upgrade requests to `ServerState.handle_chat`
- **Connection Handshake** – Upon client connection, the server transmits a `0x00` handshake packet before accepting audio data
- **Audio Processing Loop** – Uses `opus_reader` to reconstruct PCM frames from binary type-`1` messages, encodes via Mimi, processes through `LMGen.step`, and re-encodes responses via `opus_writer`
- **Text Streaming** – Sends generated text tokens as UTF-8 payloads in messages with type identifier `0x02`
- **Transmission** – Prepends `b"\x01"` to Opus packets before sending via `ws.send_bytes()`

### Binary Protocol and Client Implementation

The protocol defined in [`client/src/protocol/types.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/protocol/types.ts) and implemented in [`client/src/protocol/encoder.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/protocol/encoder.ts) specifies message framing:

- **Handshake (`0x00`)**: Server-to-client initialization signal
- **Audio (`0x01`)**: Binary Opus-encoded PCM data flowing bidirectionally
- **Text (`0x02`)**: Control messages containing agent text responses
- **Control/Metadata/Ping**: Additional message types for connection management

The React frontend uses the `useSocket` hook in [`client/src/pages/Conversation/hooks/useSocket.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/pages/Conversation/hooks/useSocket.ts) to manage the WebSocket lifecycle, implement inactivity timeouts, and encode/decode messages using the shared protocol.

## Practical Usage Examples

### Running Batch Offline Evaluation

Execute the offline module with Hugging Face authentication to process a local audio file:

```bash
HF_TOKEN=$HF_TOKEN \
python -m moshi.offline \
  --voice-prompt "NATM1.pt" \
  --text-prompt "You are a helpful assistant." \
  --input-wav "assets/test/input_service.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

```

This produces `output.wav` containing the agent-generated audio and [`output.json`](https://github.com/NVIDIA/personaplex/blob/main/output.json) with the tokenized transcript. No network connection is required after initial model downloads.

### Launching the Real-Time Server

Start the WebSocket server with SSL and optional CPU offloading for memory-constrained environments:

```bash
SSL_DIR=$(mktemp -d)
python -m moshi.server --ssl "$SSL_DIR" --cpu-offload

```

The server outputs a URL (typically `https://<host>:8998`) where the React UI connects. The service remains active, handling concurrent conversational sessions until manually terminated.

### Implementing a Custom Client

Connect to the server using the binary protocol:

```javascript
import { encodeMessage, decodeMessage } from "./protocol/encoder";
import { WSMessage } from "./protocol/types";

const ws = new WebSocket("wss://localhost:8998/api/chat");
ws.binaryType = "arraybuffer";

ws.onmessage = (ev) => {
  const msg = decodeMessage(new Uint8Array(ev.data));
  if (msg.type === "handshake") {
    console.log("Connected – start streaming audio");
  } else if (msg.type === "audio") {
    // Decode Opus packet and play back
  } else if (msg.type === "text") {
    console.log("Agent says:", msg.data);
  }
};

function sendAudioChunk(chunk) {
  const message = { type: "audio", data: chunk };
  ws.send(encodeMessage(message));
}

```

This low-level implementation mirrors the logic found in [`useSocket.ts`](https://github.com/NVIDIA/personaplex/blob/main/useSocket.ts), allowing custom clients to interact with the PersonaPlex server without the React frontend.

## Summary

- **Offline evaluation** ([`moshi/moshi/offline.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/offline.py)) provides deterministic, file-based batch processing with no network requirements, ideal for automated testing and reproducible experiments.
- **Live WebSocket server** ([`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py)) enables real-time duplex audio communication via `aiohttp` and binary Opus streaming, suited for interactive demos and production conversational interfaces.
- Both modes share identical core components: the Mimi encoder, Moshi LM, voice prompt handling via `_get_voice_prompt_dir`, and warmup procedures.
- The WebSocket protocol uses typed binary messages (`0x00` handshake, `0x01` audio, `0x02` text) encoded through [`client/src/protocol/encoder.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/protocol/encoder.ts).

## Frequently Asked Questions

### Which mode should I use for automated testing in PersonaPlex?

Use **offline evaluation** for automated testing. The `python -m moshi.offline` command processes local WAV files deterministically, writes reproducible outputs to disk, and requires no network configuration. This eliminates variability from network latency or client-side timing issues, making it ideal for CI/CD pipelines and regression testing.

### How does the WebSocket protocol handle audio encoding in PersonaPlex?

The protocol transmits audio as **Opus-encoded binary packets** with message type `0x01`. The server receives these via `opus_reader`, decodes them to PCM, processes through the Mimi encoder and Moshi LM, then re-encodes the response as Opus via `opus_writer` before transmitting back to the client with the `0x01` prefix.

### Can I use the same voice prompts in both offline and server modes?

Yes. Both implementations call `_get_voice_prompt_dir` to locate the `voices.tgz` archive. In [`moshi/moshi/offline.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/offline.py), this occurs immediately before prompt injection, while in [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py), voice prompts are loaded at startup to serve any client request. The same `.pt` prompt files (such as `NATM1.pt`) work identically in both contexts.

### What is the purpose of the warmup phase in both modes?

The **warmup** phase compiles CUDA graphs and initializes streaming state to prevent latency spikes during actual inference. Offline mode calls `warmup(mimi, other_mimi, lm_gen, device, frame_size)` with dummy zero frames before processing input, while the server calls `state.warmup()` during startup. This ensures consistent performance when processing begins.