How PersonaPlex Processes System Tags for Persona Control: Complete Implementation Guide

PersonaPlex automatically wraps persona instructions in <system> tags using the wrap_with_system_tags utility, tokenizes the wrapped text into LMGen.text_prompt_tokens, and injects these tokens during the system-prompt phase via step_system_prompts_async to control the model's persona.

NVIDIA's PersonaPlex repository implements a deterministic pipeline for handling system-level persona instructions. Understanding how PersonaPlex processes system tags for persona control is essential for developers integrating custom personalities into the Moshi dialogue model, as the model expects specific formatting to recognize system-level directives.

The Role of System Tags in PersonaPlex

PersonaPlex uses system tags to delimit the portion of a text prompt containing persona-control instructions. The model expects these tags to be present; otherwise, the instructions may be ignored or treated as user dialogue. The repository implements a small utility that automatically adds the tags when they are missing, then feeds the tagged text into the language model during the system-prompt phase.

Unlike standard XML-style tags that use opening and closing variants, PersonaPlex uses <system> at both the beginning and end of the persona instruction block. This formatting signals to the underlying language model that the enclosed text represents system-level configuration rather than conversational content.

Automatic Tag Injection with wrap_with_system_tags

Both the server (live chat) and offline inference pipeline call the same helper function to ensure consistent tag formatting.

The Tag Wrapping Implementation

The wrap_with_system_tags function in moshi/moshi/server.py (lines 79-87) and moshi/moshi/offline.py (lines 83-90) implements a sanitization check:

def wrap_with_system_tags(text: str) -> str:
    """Add system tags as the model expects if they are missing.
    Example: "<system> You enjoy having a good conversation. <system>"
    """
    cleaned = text.strip()
    if cleaned.startswith("<system>") and cleaned.endswith("<system>"):
        return cleaned
    return f"<system> {cleaned} <system>"

This function performs three critical operations:

  1. Strips whitespace from the input text to ensure clean tag placement.
  2. Checks for existing tags to prevent double-wrapping if the user already provided them.
  3. Wraps the content with <system> tags at both start and end if they are absent.

Server-Side Integration

When processing live chat requests, the server automatically applies this transformation to the text_prompt query parameter:


# Located in moshi/moshi/server.py around line 70-71

self.lm_gen.text_prompt_tokens = (
    self.text_tokenizer.encode(
        wrap_with_system_tags(request.query["text_prompt"])
    )
    if len(request.query["text_prompt"]) > 0 else None
)

The raw user-provided string is sanitized, wrapped, tokenized, and stored in LMGen.text_prompt_tokens for later injection into the generation pipeline.

Tokenization and Storage

After wrapping, the text undergoes tokenization using the model's text tokenizer. The resulting token IDs are stored in the text_prompt_tokens attribute of the LMGen class. This storage mechanism separates the persona configuration from the ongoing dialogue context, ensuring that system instructions maintain their authoritative position in the model's context window.

The tokenization happens immediately after tag wrapping in both the server and offline code paths, creating a deterministic pipeline from raw text to model-ready tokens.

System Prompt Execution Phase

The actual injection of persona instructions occurs during the system-prompt phase of the language model generator. This phase runs before the main dialogue loop begins, establishing the persona context for the entire conversation.

step_system_prompts_async Workflow

Located in moshi/moshi/models/lm.py (lines 1117-1122), the step_system_prompts_async method orchestrates the initialization sequence:

async def step_system_prompts_async(self, mimi, is_alive: Optional[Callable]=None):
    await self._step_voice_prompt_async(mimi, is_alive)
    await self._step_audio_silence_async(is_alive)
    await self._step_text_prompt_async(is_alive)
    await self._step_audio_silence_async(is_alive)

The _step_text_prompt_async routine iterates over self.text_prompt_tokens (the tokenized, system-tag-wrapped text) and inserts each token into the model's decoding stream. This is the precise moment when the persona instructions influence the generation, setting the behavioral baseline for the conversation.

End-to-End Usage Examples

Live Server Implementation

When a client sends a request containing a plain text prompt, the server handles the entire wrapping and tokenization pipeline automatically:


# Client sends request with persona description

request.query["text_prompt"] = "You are an enthusiastic tech tutor."

# Server automatically processes (internal implementation)

wrapped = wrap_with_system_tags(request.query["text_prompt"])

# Result: "<system> You are an enthusiastic tech tutor. <system>"

tokens = tokenizer.encode(wrapped)

# Tokens stored in LMGen.text_prompt_tokens for system prompt phase

Offline Inference Script

For batch processing or development workflows, the offline module provides the same functionality:

from moshi.moshi.offline import wrap_with_system_tags, run_inference

prompt = "You are a calm, helpful assistant."

# Manual wrapping (optional - run_inference handles this internally)

wrapped_prompt = wrap_with_system_tags(prompt)

# Returns: "<system> You are a calm, helpful assistant. <system>"

# Or pass raw prompt to run_inference, which calls wrap_with_system_tags internally

run_inference(
    input_wav="user.wav",
    output_wav="assistant.wav",
    output_text="assistant.txt",
    text_prompt=prompt,  # Automatically wrapped inside run_inference

    voice_prompt_path="voice.pt",
    tokenizer_path="tokenizer.model",
    moshi_weight="moshi.ckpt",
    mimi_weight="mimi.ckpt",
    hf_repo="nvidia/personaplex-7b-v1",
    device="cuda",
    seed=42,
    temp_audio=0.7,
    temp_text=0.8,
    topk_audio=10,
    topk_text=10,
    greedy=False,
)

Summary

  • Automatic Wrapping: The wrap_with_system_tags function in moshi/moshi/server.py and moshi/moshi/offline.py automatically adds <system> tags to persona instructions if missing, using <system> as both opening and closing delimiters.
  • Token Storage: Wrapped text is tokenized and stored in LMGen.text_prompt_tokens, separating system instructions from dialogue context.
  • Phase Injection: The step_system_prompts_async method in moshi/moshi/models/lm.py executes the system-prompt phase, calling _step_text_prompt_async to stream tokens into the model before conversation begins.
  • Dual Path: Both live server and offline inference pipelines share identical tag-wrapping logic, ensuring consistent persona control across deployment scenarios.

Frequently Asked Questions

What happens if I don't use system tags in my PersonaPlex prompt?

If you provide a text prompt without <system> tags, the wrap_with_system_tags function automatically adds them during preprocessing. However, if you manually bypass this utility and send untagged text directly to the tokenizer, the model may ignore the persona instructions or interpret them as user dialogue rather than system-level configuration, resulting in inconsistent persona adherence.

Why does PersonaPlex use <system> as both opening and closing tags instead of standard XML?

According to the implementation in moshi/moshi/server.py, PersonaPlex uses <system> at both the start and end of the persona instruction (e.g., <system> You are a tutor. <system>). This specific formatting convention matches the expectations of the underlying language model architecture, distinguishing system prompts from other structured data formats that might use </system> as a closing tag.

Can I provide pre-tagged text to skip the wrapping step?

Yes. The wrap_with_system_tags function checks if the input already starts and ends with <system> tags using cleaned.startswith("<system>") and cleaned.endswith("<system>"). If both conditions are met, the function returns the text unchanged, allowing advanced users to provide manually formatted prompts while preventing double-wrapping.

Where exactly are the system prompt tokens injected during generation?

The tokens are injected during the step_system_prompts_async method defined in moshi/moshi/models/lm.py at lines 1117-1122. Specifically, the _step_text_prompt_async sub-routine processes the text_prompt_tokens stored in the LMGen instance, feeding them into the model's stream before any user audio or text enters the generation pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →