How to Implement Custom Chat Prompts and Conversation Memory in Speech-to-Speech

You can customize system prompts by using build_voice_system_prompt or build_text_system_prompt from the prompt builder modules, and manage conversation history by instantiating the Chat class with a specified size limit, calling add_item() for each turn and trim_if_needed() to enforce bounds or trigger background compaction.

The huggingface/speech-to-speech framework provides a modular architecture for building real-time conversational agents. Implementing custom chat prompts and conversation memory requires understanding two core components: the prompt construction system in speech_to_speech/LLM/ and the bounded buffer implementation in the Chat class. This guide shows you how to configure persona-driven system prompts and implement scalable memory management that automatically evicts or summarizes old conversation turns.

Building Custom System Prompts for Voice and Text

The library separates prompt construction by channel (voice vs. text) to optimize for spoken versus written interaction patterns. Both builders share the same skeleton structure but inject channel-specific rules.

Voice Prompt Construction

For real-time voice conversations, use build_voice_system_prompt from speech_to_speech/LLM/voice_prompt.py. This function assembles a prompt using VOICE_SYSTEM_PROMPT_LEAD to establish spoken-context rules, your custom session prompt, and VOICE_SYSTEM_PROMPT_TAIL to enforce constraints like "keep replies brief" and "speak before a tool call".

from speech_to_speech.LLM.voice_prompt import build_voice_system_prompt

session_prompt = "You are a travel assistant. Keep answers concise and friendly."
tool_section = ""  # Optional: add tool instructions here

full_system_prompt = build_voice_system_prompt(session_prompt, tool_section=tool_section)

Source: voice_prompt.py – build_voice_system_prompt

Text Prompt Construction

For text-based interactions, use build_text_system_prompt from speech_to_speech/LLM/text_prompt.py. This variant uses TEXT_SYSTEM_PROMPT_LEAD ("You are a helpful assistant") and TEXT_SYSTEM_PROMPT_TAIL with markdown-friendly formatting rules instead of spoken-style prose.

from speech_to_speech.LLM.text_prompt import build_text_system_prompt

system_prompt = build_text_system_prompt(session_prompt, tool_section=tool_section)

Source: text_prompt.py – build_text_system_prompt

Tool-Call Integration

When your agent uses function calling, generate the tool instruction block using build_tool_system_prompt from speech_to_speech/LLM/tool_call/tool_prompt.py. This function accepts a text_only parameter to toggle between voice-friendly templates (with spoken instructions) and text-only variants.

from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool

weather_tool = FunctionTool(
    name="get_weather",
    description="Fetch current weather for a city.",
    parameters={
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
)

# Voice variant includes "speak first" prose; set text_only=True for text mode

tool_section = build_tool_system_prompt([weather_tool], text_only=False)

Source: tool_prompt.py – build_tool_system_prompt

In speech_to_speech/LLM/language_model.py, the LanguageModel class automatically selects the appropriate builder based on the wants_audio flag, assembling the full instructions by combining the session prompt with the tool section.

Managing Conversation Memory with the Chat Class

The Chat class in speech_to_speech/LLM/chat.py implements a bounded buffer that stores conversation history while ensuring the LLM never receives a payload exceeding your configured memory limit.

Initializing the Chat Buffer

Instantiate Chat with a size parameter representing the maximum number of user turns to retain. Initialize it with a system message using make_system_message:

from speech_to_speech.LLM.chat import Chat, make_system_message

chat = Chat(size=8)
system_msg = make_system_message(full_system_prompt)
chat.init_chat(system_msg)

The class tracks:

  • buffer: List of conversation items (the live history)
  • _user_turn_count: Current number of user messages in buffer
  • _pending_tool_calls: Tool calls awaiting output (preserved even if turns are evicted)

Source: Chat.init and fields

Adding Items and Enforcing Limits

Add conversation turns using add_item(), which validates and routes Realtime API items into the buffer:

from speech_to_speech.LLM.chat import make_user_message, make_assistant_message

chat.add_item(make_user_message("I want a beach trip near Seattle."))
chat.add_item(make_assistant_message("Sure, let me check the weather first."))

After each generation, enforce the size limit:

chat.trim_if_needed(compactor=None)  # Simple eviction of oldest turns

Source: add_item implementation

Memory Compaction and Summarization

Instead of losing old context entirely, you can enable background compaction. When trim_if_needed() receives a compactor function, it spawns a daemon thread to summarize eligible turns and replaces them with synthetic user_summary and assistant_summary messages.

First, build a compactor using build_compactor from speech_to_speech/LLM/compaction_prompt.py:

from speech_to_speech.LLM.compaction_prompt import build_compactor

def llm_generate(system: str, user: str) -> str:
    # Call a lightweight LLM (e.g., gpt-4o-mini) for summarization

    return '{"user_summary":"User wants weekend beach trip near Seattle.",' \
           '"assistant_summary":"Assistant offered to check weather."}'

my_compactor = build_compactor(llm_generate)

Then trigger compaction during trimming:

chat.trim_if_needed(compactor=my_compactor)

The compaction prompt uses templates defined in compaction_prompt.py: COMPACTION_SYSTEM_PROMPT provides instructions, while COMPACTION_USER_TEMPLATE wraps the transcript of turns being summarized. The Chat class handles splicing the results back into the buffer via _apply_compaction.

Source: build_compactor and _maybe_trigger_compaction

Complete Implementation Example

This example demonstrates assembling a custom voice prompt with tools, initializing memory, and enabling compaction:

from speech_to_speech.LLM.voice_prompt import build_voice_system_prompt
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
from speech_to_speech.LLM.chat import Chat, make_system_message, make_user_message, make_assistant_message
from speech_to_speech.LLM.compaction_prompt import build_compactor

# 1. Define persona and build system prompt

session_prompt = """
You are a travel assistant that helps users plan weekend getaways.
Keep answers concise and include a friendly tone.
"""

weather_tool = FunctionTool(
    name="get_weather",
    description="Fetch current weather for a city.",
    parameters={
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
)

tool_section = build_tool_system_prompt([weather_tool])
system_prompt = build_voice_system_prompt(session_prompt, tool_section=tool_section)

# 2. Initialize chat memory with compaction

chat = Chat(size=8)
chat.init_chat(make_system_message(system_prompt))

def summarizer(system: str, user: str) -> str:
    # Integration with your LLM client here

    return '{"user_summary":"Summarized user request", "assistant_summary":"Summarized assistant response"}'

compactor = build_compactor(summarizer)

# 3. Simulate conversation turns

chat.add_item(make_user_message("I want a beach trip near Seattle."))
chat.add_item(make_assistant_message("Sure, let me check the weather first."))

# 4. Enforce memory limits with background compaction

chat.trim_if_needed(compactor=compactor)

# 5. Serialize for LLM request

payload = chat.to_responses_api_chat()  # Ready for ResponsesApiModelHandler

Summary

  • Custom prompts: Use build_voice_system_prompt or build_text_system_prompt from their respective modules, injecting your session description and optional tool sections created via build_tool_system_prompt.
  • Tool integration: Pass FunctionTool objects to build_tool_system_prompt, choosing text_only=True for text channels and text_only=False for voice channels.
  • Memory initialization: Create a Chat instance with a specific size (max user turns) and seed it with make_system_message.
  • Turn management: Call add_item() for each user message, assistant reply, or function interaction, then invoke trim_if_needed() after every generation cycle.
  • Long-term memory: Supply a compactor function (built via build_compactor) to trim_if_needed() to trigger background summarization of old turns rather than simple eviction.

Frequently Asked Questions

What is the difference between voice and text system prompts in speech-to-speech?

Voice prompts (build_voice_system_prompt) insert channel-specific rules like "keep replies brief" and "speak before calling a tool" using templates stored in VOICE_SYSTEM_PROMPT_LEAD and VOICE_SYSTEM_PROMPT_TAIL. Text prompts (build_text_system_prompt) use different templates optimized for markdown readability without spoken-conversation constraints, as implemented in speech_to_speech/LLM/text_prompt.py.

How does the Chat class prevent memory from growing indefinitely?

The Chat class enforces a hard limit via the size parameter representing maximum user turns. When you call trim_if_needed(), it checks _user_turn_count against this limit and either evicts the oldest complete turn (if compactor=None) or triggers background compaction to replace old turns with summary messages.

What happens to pending tool calls when conversation turns are evicted?

The Chat class maintains a _pending_tool_calls registry that tracks function calls awaiting their output. If a turn containing a tool call is evicted before the function returns, the metadata is preserved. When the function output arrives via add_item(), the system handles re-injection of the call context so the LLM can still correlate the result with the original request.

How do I enable conversation memory compression instead of deletion?

Import build_compactor from speech_to_speech/LLM/compaction_prompt.py and provide a generate_fn callable that accepts (system_prompt, user_prompt) and returns a JSON string with user_summary and assistant_summary keys. Pass the resulting compactor to chat.trim_if_needed(compactor=my_compactor) to activate background thread summarization via the COMPACTION_SYSTEM_PROMPT template.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →