PersonaPlex Offline Evaluation vs Live WebSocket Server: Architecture and Usage

PersonaPlex supports both offline evaluation through the moshi.offline CLI module for batch processing and a live WebSocket server via moshi.server for real-time audio streaming, sharing the same Mimi encoder and Moshi LM pipeline under the hood.

The NVIDIA/personaplex repository provides a multimodal conversational AI framework that can operate in two distinct execution modes. Whether you need reproducible batch processing for automated testing or low-latency duplex audio for interactive demos, understanding the differences between offline evaluation and the live WebSocket server is essential for optimal deployment.

Core Architectural Differences

The fundamental distinction lies in how data enters and exits the system. Offline evaluation runs as a self-contained Python process that reads local WAV files and writes results to disk, while the live WebSocket server exposes an HTTP endpoint that accepts real-time binary audio streams from remote clients.

Entry Points and Execution Models

You initiate offline evaluation by running python -m moshi.offline, which parses command-line arguments such as --input-wav, --voice-prompt, and --output-wav before executing the inference pipeline synchronously. This mode requires no network configuration and processes the entire audio file from start to finish within a single process.

In contrast, the live server starts with python -m moshi.server, spinning up an aiohttp-based HTTP service that listens for WebSocket connections on /api/chat. The server maintains persistent connections with clients (typically the React frontend) and processes audio frame-by-frame as binary Opus-encoded packets arrive over the network.

Data Flow Comparison

The offline pipeline in moshi/moshi/offline.py loads user audio from disk, iterates through frames using lm_iterate_audio, and concatenates generated PCM frames before writing the final WAV file. The server implementation in moshi/moshi/server.py follows a continuous loop: receiving binary messages (type 0x01 for audio), decoding them through an opus_reader, encoding with Mimi, generating responses via LMGen.step, and streaming back Opus packets via ws.send_bytes.

Offline Evaluation Mode Deep Dive

Offline evaluation is ideal for automated testing, batch processing, and environments where network services are undesirable. The entire lifecycle completes without external dependencies.

Pipeline Architecture in offline.py

The evaluation script follows a strict linear progression through seven phases:

  1. CLI Parsing – Validates arguments including --input-wav, --voice-prompt, and --text-prompt
  2. Model Loading – Downloads or loads from cache the Mimi encoder/decoder, Moshi LM, and SentencePiece tokenizer
  3. Warmup – Calls warmup(mimi, other_mimi, lm_gen, device, frame_size) to compile CUDA graphs with dummy zero frames
  4. Prompt Injection – Uses LMGen.step_system_prompts to inject system text tokens and voice-prompt audio
  5. Streaming Simulation – Processes the input WAV frame-by-frame via lm_iterate_audio, encoding each frame and feeding it to the LM
  6. Output Generation – Decodes agent audio tokens, trims/pads to match original duration, and writes to disk using sphn.write_wav
  7. Transcript Extraction – Collects text tokens and saves them as JSON alongside the audio output

Voice Prompt Handling

Both modes utilize the _get_voice_prompt_dir helper to locate the voices.tgz archive. In offline mode, this occurs immediately before processing begins, loading the requested .pt file (such as NATM1.pt) directly from disk into memory for injection into the prompt phase.

Live WebSocket Server Mode Deep Dive

The server mode transforms PersonaPlex into a real-time conversational service accessible via browser or custom clients.

Server Architecture in server.py

The moshi/moshi/server.py implementation manages concurrent connections through an async event loop:

  • Startup Phase – Downloads voice prompts, loads model components identical to the offline script, and executes state.warmup() to prime CUDA graphs and streaming state
  • HTTP Routing – Registers /api/chat to route WebSocket upgrade requests to ServerState.handle_chat
  • Connection Handshake – Upon client connection, the server transmits a 0x00 handshake packet before accepting audio data
  • Audio Processing Loop – Uses opus_reader to reconstruct PCM frames from binary type-1 messages, encodes via Mimi, processes through LMGen.step, and re-encodes responses via opus_writer
  • Text Streaming – Sends generated text tokens as UTF-8 payloads in messages with type identifier 0x02
  • Transmission – Prepends b"\x01" to Opus packets before sending via ws.send_bytes()

Binary Protocol and Client Implementation

The protocol defined in client/src/protocol/types.ts and implemented in client/src/protocol/encoder.ts specifies message framing:

  • Handshake (0x00): Server-to-client initialization signal
  • Audio (0x01): Binary Opus-encoded PCM data flowing bidirectionally
  • Text (0x02): Control messages containing agent text responses
  • Control/Metadata/Ping: Additional message types for connection management

The React frontend uses the useSocket hook in client/src/pages/Conversation/hooks/useSocket.ts to manage the WebSocket lifecycle, implement inactivity timeouts, and encode/decode messages using the shared protocol.

Practical Usage Examples

Running Batch Offline Evaluation

Execute the offline module with Hugging Face authentication to process a local audio file:

HF_TOKEN=$HF_TOKEN \
python -m moshi.offline \
  --voice-prompt "NATM1.pt" \
  --text-prompt "You are a helpful assistant." \
  --input-wav "assets/test/input_service.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

This produces output.wav containing the agent-generated audio and output.json with the tokenized transcript. No network connection is required after initial model downloads.

Launching the Real-Time Server

Start the WebSocket server with SSL and optional CPU offloading for memory-constrained environments:

SSL_DIR=$(mktemp -d)
python -m moshi.server --ssl "$SSL_DIR" --cpu-offload

The server outputs a URL (typically https://<host>:8998) where the React UI connects. The service remains active, handling concurrent conversational sessions until manually terminated.

Implementing a Custom Client

Connect to the server using the binary protocol:

import { encodeMessage, decodeMessage } from "./protocol/encoder";
import { WSMessage } from "./protocol/types";

const ws = new WebSocket("wss://localhost:8998/api/chat");
ws.binaryType = "arraybuffer";

ws.onmessage = (ev) => {
  const msg = decodeMessage(new Uint8Array(ev.data));
  if (msg.type === "handshake") {
    console.log("Connected – start streaming audio");
  } else if (msg.type === "audio") {
    // Decode Opus packet and play back
  } else if (msg.type === "text") {
    console.log("Agent says:", msg.data);
  }
};

function sendAudioChunk(chunk) {
  const message = { type: "audio", data: chunk };
  ws.send(encodeMessage(message));
}

This low-level implementation mirrors the logic found in useSocket.ts, allowing custom clients to interact with the PersonaPlex server without the React frontend.

Summary

  • Offline evaluation (moshi/moshi/offline.py) provides deterministic, file-based batch processing with no network requirements, ideal for automated testing and reproducible experiments.
  • Live WebSocket server (moshi/moshi/server.py) enables real-time duplex audio communication via aiohttp and binary Opus streaming, suited for interactive demos and production conversational interfaces.
  • Both modes share identical core components: the Mimi encoder, Moshi LM, voice prompt handling via _get_voice_prompt_dir, and warmup procedures.
  • The WebSocket protocol uses typed binary messages (0x00 handshake, 0x01 audio, 0x02 text) encoded through client/src/protocol/encoder.ts.

Frequently Asked Questions

Which mode should I use for automated testing in PersonaPlex?

Use offline evaluation for automated testing. The python -m moshi.offline command processes local WAV files deterministically, writes reproducible outputs to disk, and requires no network configuration. This eliminates variability from network latency or client-side timing issues, making it ideal for CI/CD pipelines and regression testing.

How does the WebSocket protocol handle audio encoding in PersonaPlex?

The protocol transmits audio as Opus-encoded binary packets with message type 0x01. The server receives these via opus_reader, decodes them to PCM, processes through the Mimi encoder and Moshi LM, then re-encodes the response as Opus via opus_writer before transmitting back to the client with the 0x01 prefix.

Can I use the same voice prompts in both offline and server modes?

Yes. Both implementations call _get_voice_prompt_dir to locate the voices.tgz archive. In moshi/moshi/offline.py, this occurs immediately before prompt injection, while in moshi/moshi/server.py, voice prompts are loaded at startup to serve any client request. The same .pt prompt files (such as NATM1.pt) work identically in both contexts.

What is the purpose of the warmup phase in both modes?

The warmup phase compiles CUDA graphs and initializes streaming state to prevent latency spikes during actual inference. Offline mode calls warmup(mimi, other_mimi, lm_gen, device, frame_size) with dummy zero frames before processing input, while the server calls state.warmup() during startup. This ensures consistent performance when processing begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →