# How PersonaPlex Manages Turn-Taking in Conversations: Architecture and Implementation

> Discover how PersonaPlex manages turn-taking in conversations using interleaved system prompts, configurable audio-silence periods, and user audio within a single streaming loop. Learn the architecture and implementation.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: architecture
- Published: 2026-04-07

---

**PersonaPlex manages turn-taking in conversations by interleaving system prompts, configurable audio-silence periods, and user audio within a single streaming loop orchestrated by the `LMGen` class and `ServerState` handler.**

PersonaPlex, an open-source conversational AI framework from NVIDIA, implements real-time voice interaction through a sophisticated **turn-taking mechanism** that synchronizes model speech and user input. Unlike simple request-response APIs, this system uses silence detection and prompt interleaving to create natural, full-duplex conversations. The implementation spans the `LMGen` class in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py) and the `ServerState` handler in [`moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/server.py).

## System Prompt Sequencing with `LMGen.step_system_prompts_async`

The turn-taking process initiates through the `step_system_prompts_async` method in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py) (lines 1117-1121), which establishes the conversation "script" before user interaction begins. This method executes a precise four-step sequence: voice prompt emission, initial silence insertion, text prompt processing, and final silence generation.

According to the PersonaPlex source code, the method chains four internal calls to demarcate speaking turns:

```python
await self._step_voice_prompt_async(mimi, is_alive)   # voice prompt

await self._step_audio_silence_async(is_alive)      # pause → user turn

await self._step_text_prompt_async(is_alive)        # text prompt

await self._step_audio_silence_async(is_alive)      # pause → model turn

```

This sequencing creates explicit **handoff points** where the system pauses, allowing the audio pipeline to switch between the model's voice synthesis and user input capture.

## Silence Gap Generation via `_step_audio_silence_core`

Between speaker transitions, PersonaPlex inserts configurable silent frames using `_step_audio_silence_core` (defined in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py), lines 1075-1077). The method generates `audio_silence_frame_cnt` silent frames—defaulting to 0.5 seconds—to demarcate turn boundaries.

The implementation yields `None` values for each silent frame:

```python
for _ in range(self.audio_silence_frame_cnt):
    # generate a silent audio frame (no tokens, just pause)

    yield None

```

These silent periods serve as **acoustic buffers** that prevent the model from interrupting user speech and signal to the streaming pipeline that a turn transition is complete.

## Full-Duplex Stream Coordination in `ServerState.handle_chat`

While the `LMGen` class handles prompt sequencing, the `ServerState.handle_chat` method in [`moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/server.py) (lines 289-298) manages the **concurrent audio streams** that execute the turn-taking in real time. The method spawns three asyncio tasks that operate simultaneously:

```python
tasks = [
    asyncio.create_task(recv_loop()),   # receives user audio

    asyncio.create_task(opus_loop()),   # processes audio → model output

    asyncio.create_task(send_loop()),   # streams model audio back

]
done, pending = await asyncio.wait(tasks, return_when=asyncio.FIRST_COMPLETED)

```

The `recv_loop` captures incoming user audio, `opus_loop` processes tokens through the Mimi codec and language model, and `send_loop` transmits generated speech back to the client. These loops respect the silence markers inserted by `LMGen`, ensuring that the model only generates output after detecting the user's silence period.

To prevent **cross-client interference**, the server maintains an `asyncio.Lock` (`self.lock`) that guarantees exclusive access to the conversation state, ensuring that only one client controls the turn-taking sequence at any moment.

## The Four-Phase Turn-Taking Sequence

PersonaPlex structures conversation flow through a rigid four-phase cycle that alternates between model and user control:

- **Phase 1: Voice Prompt** — The model encodes a short voice embedding (e.g., "You are Jane") via `_step_voice_prompt_async`, establishing the persona's acoustic signature.
- **Phase 2: Silence 1 (User Turn Signal)** — `audio_silence_frame_cnt` silent frames create a gap that gives the user an opportunity to speak, preventing immediate model interruption.
- **Phase 3: Text Prompt** — Optional contextual text is tokenized and fed to the model through `_step_text_prompt_async`, conditioning the response generation.
- **Phase 4: Silence 2 (Model Turn Signal)** — A final silence gap signals that the model should begin generating audio output, completing the turn initialization.

This sequence repeats throughout the conversation, with the `opus_loop` continuously checking for user audio during silence periods before triggering model generation.

## Client-Server Implementation of Turn-Taking

On the client side, the WebSocket connection established in [`client/src/pages/Conversation/hooks/useSocket.ts`](https://github.com/NVIDIA/personaplex/blob/main/client/src/pages/Conversation/hooks/useSocket.ts) transmits binary audio packets marked with kind `0x01`:

```typescript
const ws = new WebSocket(`ws://${host}/api/chat?text_prompt=${encodeURIComponent(prompt)}&voice_prompt=${voiceFile}`);
ws.binaryType = "arraybuffer";

// Send recorded audio packets (kind 0x01)
ws.send(new Uint8Array([0x01, ...opusChunk]));

```

Meanwhile, the server-side `handle_chat` method initializes the conversation by calling `step_system_prompts_async` before entering the streaming loop:

```python
await self.lm_gen.step_system_prompts_async(self.mimi, is_alive=is_alive)
self.mimi.reset_streaming()

# After system prompts the streaming loops start...

```

When the model generates text responses, the `opus_loop` transmits them with kind `0x02`:

```python
if text_token not in (0, 3):
    _text = self.text_tokenizer.id_to_piece(text_token).replace("▁", " ")
    await ws.send_bytes(b"\x02" + _text.encode())

```

This binary message protocol ensures that text and audio streams remain synchronized during turn transitions.

## Summary

- **PersonaPlex implements turn-taking in conversations** through a combination of system prompt sequencing, configurable silence frames, and full-duplex asyncio streaming.
- The `LMGen.step_system_prompts_async` method in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py) orchestrates the initial conversation script by alternating voice prompts, text prompts, and silence periods.
- Silent frames generated by `_step_audio_silence_core` create default 0.5-second gaps that signal turn transitions between the model and user.
- Three concurrent loops (`recv_loop`, `opus_loop`, `send_loop`) in [`moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/server.py) handle real-time audio streaming while respecting silence boundaries.
- An `asyncio.Lock` prevents overlapping conversations, ensuring that only one client controls the turn-taking stream at any given time.

## Frequently Asked Questions

### How does PersonaPlex prevent overlapping speech during turn-taking?

The system utilizes configurable silence frames generated by `_step_audio_silence_core` to create acoustic gaps between speaker turns. Additionally, the `asyncio.Lock` in `ServerState` ensures that only one client can control the conversation stream at a time, preventing cross-user interference or simultaneous speaking.

### What is the default silence duration between turns in PersonaPlex?

The default silence duration is **0.5 seconds**, controlled by the `audio_silence_frame_cnt` parameter in `LMGen`. This value determines how many silent frames are yielded between the voice prompt, text prompt, and active conversation turns to allow for natural speech cadence.

### How does the `LMGen` class coordinate voice and text prompts?

The `LMGen` class sequences prompts through `step_system_prompts_async`, which calls four internal methods in order: `_step_voice_prompt_async`, `_step_audio_silence_async`, `_step_text_prompt_async`, and a final `_step_audio_silence_async`. This chaining ensures that the model establishes its persona acoustically, pauses for user input, processes contextual text, and then signals readiness to generate output.

### How are client and server streams synchronized during turn transitions?

The client transmits audio packets with byte prefix `0x01` while the server uses `0x02` for text responses, creating a binary protocol that distinguishes message types. The server's `handle_chat` method runs three asyncio tasks (`recv_loop`, `opus_loop`, `send_loop`) that continuously process streams until `asyncio.FIRST_COMPLETED` triggers, ensuring that silence detection and audio generation remain time-synchronized.