PersonaPlex Client-Side Web UI Architecture: React 18, Web Audio API, and Custom WebSocket Protocols

The PersonaPlex client-side Web UI is a React 18 application that uses React Router for navigation, React Contexts for state management, custom hooks for WebSocket and audio handling, and a Web Audio API Worklet for real-time duplex audio streaming.

The PersonaPlex client-side Web UI architecture powers NVIDIA's open-source conversational AI interface, delivering a browser-based full-duplex voice chat experience. Built with React 18 and TypeScript, the application runs entirely client-side and communicates with backend services via WebSocket using a custom binary protocol. This architecture enables real-time audio streaming and recording without server-side rendering dependencies.

Technology Stack and Build System

The UI rests on a modern React stack compiled by Vite and styled with Tailwind CSS. The build configuration in vite.config.ts bundles TypeScript source files while client/tailwind.config.js drives responsive layouts using utility classes like max-w-96 and md:max-w-screen-lg.

Environment configuration is centralized in client/src/env.ts, which parses import.meta.env.VITE_QUEUE_API_PATH to set the API endpoint and determines VITE_ENV mode ("development" or "production").

Routing and Page Structure

The entry point in client/src/app.tsx establishes a createBrowserRouter with a single root route that renders the <Queue /> component.

// client/src/app.tsx
import { createBrowserRouter, RouterProvider } from "react-router-dom";
import { Queue } from "./pages/Queue/Queue";

const router = createBrowserRouter([
  { path: "/", element: <Queue /> },
]);

ReactDOM.createRoot(document.getElementById("root")!).render(
  <RouterProvider router={router} />
);

Queue Page (Configuration UI)

client/src/pages/Queue/Queue.tsx serves as the landing interface for model configuration. It handles text prompt selection, voice model selection, and microphone permission requests. Once the user authorizes audio access, the component initializes the AudioContext and the MoshiProcessor worklet before conditionally rendering the <Conversation /> component.

Conversation Page (Chat Interface)

client/src/pages/Conversation/Conversation.tsx hosts the active chat session. It wraps child components with SocketContext and MediaContext providers, manages the WebSocket lifecycle, and renders download links for recorded conversations.

State Management with React Contexts

The architecture eliminates prop drilling through two primary context providers defined in the Conversation directory.

SocketContext

client/src/pages/Conversation/SocketContext.ts exposes the WebSocket instance, connection status, and message dispatch function:

export const SocketContext = createContext<{
  socketStatus: SocketStatus;
  sendMessage: (msg: WSMessage) => void;
  socket: WebSocket | null;
}>({ socketStatus: "disconnected", sendMessage: () => {}, socket: null });

MediaContext

client/src/pages/Conversation/MediaContext.ts holds references to the AudioContext, AudioWorkletNode (the "moshi-processor"), the recording destination node, and control functions (startRecording, stopRecording). Both ServerAudio and UserAudio components consume this context to handle playback and capture.

Custom Hooks for Reusable Logic

Complex interactions are encapsulated in custom hooks located under client/src/pages/Conversation/hooks/.

useSocket Hook

client/src/pages/Conversation/hooks/useSocket.ts manages the full WebSocket lifecycle. It handles binary message encoding/decoding, automatic reconnection logic, inactivity timeouts, and exposes socketStatus, sendMessage, start, and stop methods to consuming components.

useModelParams Hook

client/src/pages/Conversation/hooks/useModelParams.ts stores tunable inference parameters—including temperature, top-k values, prompt text, voice file selection, and random seed. The hook persists the random seed across sessions using useLocalStorage.

Supporting Hooks

The codebase includes useUserAudio for microphone stream management and useSystemTheme for detecting color preferences (currently forced to light mode in the Queue interface).

WebSocket Communication Protocol

The UI communicates with the backend endpoint /api/chat using a custom binary protocol defined in client/src/protocol/encoder.ts and client/src/protocol/types.ts.

Messages are encoded from plain objects into Uint8Array buffers for transmission, and incoming binary payloads are decoded back into typed WSMessage structures. The useSocket hook constructs the connection URL by appending model parameters as query strings:

const url = new URL(`${wsProtocol}://${workerAddr}/api/chat`);
url.searchParams.append("text_temperature", params.textTemperature.toString());
// Additional parameters appended...

Upon connection, the client sends a handshake message; once the server responds, socketStatus transitions to "connected" and the chat interface activates.

Real-Time Audio Processing Pipeline

The audio architecture leverages the Web Audio API with a dedicated AudioWorklet for low-latency performance.

MoshiProcessor Worklet

client/src/audio-processor.ts implements the MoshiProcessor, which runs in an isolated thread to avoid blocking the main UI. A separate client/src/decoder/decoderWorker.ts pre-warms the WASM decoder as soon as the AudioContext initializes.

Duplex Streaming and Recording

The system creates a MediaStreamAudioDestinationNode for recording output. A ChannelMergerNode composites the server audio stream onto the left channel and the user microphone input onto the right channel, enabling full-duplex conversation capture via the MediaRecorder API.

Initialization Flow

The Queue component triggers startProcessor() to instantiate the AudioContext, load the worklet module, and register the "moshi-processor" node. This worklet handles adaptive buffering, jitter compensation, and reports latency metrics back to the main thread via postMessage.

Summary

  • React 18 with TypeScript: The UI uses strict typing and modern hooks, bundled via Vite with Tailwind CSS for styling.
  • Context-Based Architecture: SocketContext and MediaContext in client/src/pages/Conversation/ provide global access to WebSocket connections and audio objects.
  • Custom Binary Protocol: Efficient WebSocket communication uses typed encode/decode utilities in client/src/protocol/encoder.ts rather than JSON.
  • Web Audio Worklets: The MoshiProcessor in client/src/audio-processor.ts handles real-time audio buffering and playback in a dedicated thread.
  • Full-Duplex Recording: The MediaRecorder captures mixed server and user audio through a stereo merger node for complete conversation archival.

Frequently Asked Questions

What frontend technologies power the PersonaPlex UI?

The PersonaPlex client-side Web UI architecture relies on React 18, TypeScript, and React Router v6 for the component framework and navigation. Vite handles the build process and hot module replacement, while Tailwind CSS provides the styling system through utility classes.

How does the application manage WebSocket connections?

Connection logic resides in the useSocket hook within client/src/pages/Conversation/hooks/useSocket.ts. This hook manages connection lifecycle, automatic reconnection, inactivity timeouts, and binary message encoding/decoding using the custom protocol defined in client/src/protocol/encoder.ts.

Where is the audio processing logic implemented?

Real-time audio processing occurs in the MoshiProcessor worklet defined in client/src/audio-processor.ts, which runs separate from the main thread to minimize latency. The decoder worker in client/src/decoder/decoderWorker.ts handles WASM codec initialization, while MediaContext coordinates recording and playback nodes.

How does the application route between configuration and chat views?

The createBrowserRouter in client/src/app.tsx defines a single root route rendering the Queue component. The Queue page manages conditional rendering: it displays the configuration UI until microphone permissions are granted and the audio worklet initializes, then renders the Conversation component for the active chat session.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →