# How to Implement Custom Chat Prompts and Conversation Memory in Speech-to-Speech

> Learn to implement custom chat prompts and conversation memory in speech-to-speech with Hugging Face. Customize system prompts and manage chat history efficiently for better AI interactions.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**You can customize system prompts by using `build_voice_system_prompt` or `build_text_system_prompt` from the prompt builder modules, and manage conversation history by instantiating the `Chat` class with a specified `size` limit, calling `add_item()` for each turn and `trim_if_needed()` to enforce bounds or trigger background compaction.**

The huggingface/speech-to-speech framework provides a modular architecture for building real-time conversational agents. Implementing custom chat prompts and conversation memory requires understanding two core components: the prompt construction system in `speech_to_speech/LLM/` and the bounded buffer implementation in the `Chat` class. This guide shows you how to configure persona-driven system prompts and implement scalable memory management that automatically evicts or summarizes old conversation turns.

## Building Custom System Prompts for Voice and Text

The library separates prompt construction by channel (voice vs. text) to optimize for spoken versus written interaction patterns. Both builders share the same skeleton structure but inject channel-specific rules.

### Voice Prompt Construction

For real-time voice conversations, use `build_voice_system_prompt` from [`speech_to_speech/LLM/voice_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/voice_prompt.py). This function assembles a prompt using `VOICE_SYSTEM_PROMPT_LEAD` to establish spoken-context rules, your custom session prompt, and `VOICE_SYSTEM_PROMPT_TAIL` to enforce constraints like "keep replies brief" and "speak before a tool call".

```python
from speech_to_speech.LLM.voice_prompt import build_voice_system_prompt

session_prompt = "You are a travel assistant. Keep answers concise and friendly."
tool_section = ""  # Optional: add tool instructions here

full_system_prompt = build_voice_system_prompt(session_prompt, tool_section=tool_section)

```

Source: [voice_prompt.py – build_voice_system_prompt](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/voice_prompt.py#L32-L41)

### Text Prompt Construction

For text-based interactions, use `build_text_system_prompt` from [`speech_to_speech/LLM/text_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/text_prompt.py). This variant uses `TEXT_SYSTEM_PROMPT_LEAD` ("You are a helpful assistant") and `TEXT_SYSTEM_PROMPT_TAIL` with markdown-friendly formatting rules instead of spoken-style prose.

```python
from speech_to_speech.LLM.text_prompt import build_text_system_prompt

system_prompt = build_text_system_prompt(session_prompt, tool_section=tool_section)

```

Source: [text_prompt.py – build_text_system_prompt](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/text_prompt.py#L28-L36)

### Tool-Call Integration

When your agent uses function calling, generate the tool instruction block using `build_tool_system_prompt` from [`speech_to_speech/LLM/tool_call/tool_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/tool_call/tool_prompt.py). This function accepts a `text_only` parameter to toggle between voice-friendly templates (with spoken instructions) and text-only variants.

```python
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool

weather_tool = FunctionTool(
    name="get_weather",
    description="Fetch current weather for a city.",
    parameters={
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
)

# Voice variant includes "speak first" prose; set text_only=True for text mode

tool_section = build_tool_system_prompt([weather_tool], text_only=False)

```

Source: [tool_prompt.py – build_tool_system_prompt](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/tool_call/tool_prompt.py#L78-L95)

In [`speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/language_model.py), the `LanguageModel` class automatically selects the appropriate builder based on the `wants_audio` flag, assembling the full instructions by combining the session prompt with the tool section.

## Managing Conversation Memory with the Chat Class

The `Chat` class in [`speech_to_speech/LLM/chat.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/chat.py) implements a bounded buffer that stores conversation history while ensuring the LLM never receives a payload exceeding your configured memory limit.

### Initializing the Chat Buffer

Instantiate `Chat` with a `size` parameter representing the maximum number of user turns to retain. Initialize it with a system message using `make_system_message`:

```python
from speech_to_speech.LLM.chat import Chat, make_system_message

chat = Chat(size=8)
system_msg = make_system_message(full_system_prompt)
chat.init_chat(system_msg)

```

The class tracks:
- `buffer`: List of conversation items (the live history)
- `_user_turn_count`: Current number of user messages in buffer
- `_pending_tool_calls`: Tool calls awaiting output (preserved even if turns are evicted)

Source: [Chat.__init__ and fields](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py#L95-L104)

### Adding Items and Enforcing Limits

Add conversation turns using `add_item()`, which validates and routes Realtime API items into the buffer:

```python
from speech_to_speech.LLM.chat import make_user_message, make_assistant_message

chat.add_item(make_user_message("I want a beach trip near Seattle."))
chat.add_item(make_assistant_message("Sure, let me check the weather first."))

```

After each generation, enforce the size limit:

```python
chat.trim_if_needed(compactor=None)  # Simple eviction of oldest turns

```

Source: [add_item implementation](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py#L74-L118)

### Memory Compaction and Summarization

Instead of losing old context entirely, you can enable background compaction. When `trim_if_needed()` receives a compactor function, it spawns a daemon thread to summarize eligible turns and replaces them with synthetic `user_summary` and `assistant_summary` messages.

First, build a compactor using `build_compactor` from [`speech_to_speech/LLM/compaction_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/compaction_prompt.py):

```python
from speech_to_speech.LLM.compaction_prompt import build_compactor

def llm_generate(system: str, user: str) -> str:
    # Call a lightweight LLM (e.g., gpt-4o-mini) for summarization

    return '{"user_summary":"User wants weekend beach trip near Seattle.",' \
           '"assistant_summary":"Assistant offered to check weather."}'

my_compactor = build_compactor(llm_generate)

```

Then trigger compaction during trimming:

```python
chat.trim_if_needed(compactor=my_compactor)

```

The compaction prompt uses templates defined in [`compaction_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/compaction_prompt.py): `COMPACTION_SYSTEM_PROMPT` provides instructions, while `COMPACTION_USER_TEMPLATE` wraps the transcript of turns being summarized. The `Chat` class handles splicing the results back into the buffer via `_apply_compaction`.

Source: [build_compactor](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/compaction_prompt.py#L39-L57) and [_maybe_trigger_compaction](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat.py#L29-L36)

## Complete Implementation Example

This example demonstrates assembling a custom voice prompt with tools, initializing memory, and enabling compaction:

```python
from speech_to_speech.LLM.voice_prompt import build_voice_system_prompt
from speech_to_speech.LLM.tool_call.tool_prompt import build_tool_system_prompt
from speech_to_speech.LLM.tool_call.function_tool import FunctionTool
from speech_to_speech.LLM.chat import Chat, make_system_message, make_user_message, make_assistant_message
from speech_to_speech.LLM.compaction_prompt import build_compactor

# 1. Define persona and build system prompt

session_prompt = """
You are a travel assistant that helps users plan weekend getaways.
Keep answers concise and include a friendly tone.
"""

weather_tool = FunctionTool(
    name="get_weather",
    description="Fetch current weather for a city.",
    parameters={
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
)

tool_section = build_tool_system_prompt([weather_tool])
system_prompt = build_voice_system_prompt(session_prompt, tool_section=tool_section)

# 2. Initialize chat memory with compaction

chat = Chat(size=8)
chat.init_chat(make_system_message(system_prompt))

def summarizer(system: str, user: str) -> str:
    # Integration with your LLM client here

    return '{"user_summary":"Summarized user request", "assistant_summary":"Summarized assistant response"}'

compactor = build_compactor(summarizer)

# 3. Simulate conversation turns

chat.add_item(make_user_message("I want a beach trip near Seattle."))
chat.add_item(make_assistant_message("Sure, let me check the weather first."))

# 4. Enforce memory limits with background compaction

chat.trim_if_needed(compactor=compactor)

# 5. Serialize for LLM request

payload = chat.to_responses_api_chat()  # Ready for ResponsesApiModelHandler

```

## Summary

- **Custom prompts**: Use `build_voice_system_prompt` or `build_text_system_prompt` from their respective modules, injecting your session description and optional tool sections created via `build_tool_system_prompt`.
- **Tool integration**: Pass `FunctionTool` objects to `build_tool_system_prompt`, choosing `text_only=True` for text channels and `text_only=False` for voice channels.
- **Memory initialization**: Create a `Chat` instance with a specific `size` (max user turns) and seed it with `make_system_message`.
- **Turn management**: Call `add_item()` for each user message, assistant reply, or function interaction, then invoke `trim_if_needed()` after every generation cycle.
- **Long-term memory**: Supply a compactor function (built via `build_compactor`) to `trim_if_needed()` to trigger background summarization of old turns rather than simple eviction.

## Frequently Asked Questions

### What is the difference between voice and text system prompts in speech-to-speech?

Voice prompts (`build_voice_system_prompt`) insert channel-specific rules like "keep replies brief" and "speak before calling a tool" using templates stored in `VOICE_SYSTEM_PROMPT_LEAD` and `VOICE_SYSTEM_PROMPT_TAIL`. Text prompts (`build_text_system_prompt`) use different templates optimized for markdown readability without spoken-conversation constraints, as implemented in [`speech_to_speech/LLM/text_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/text_prompt.py).

### How does the Chat class prevent memory from growing indefinitely?

The `Chat` class enforces a hard limit via the `size` parameter representing maximum user turns. When you call `trim_if_needed()`, it checks `_user_turn_count` against this limit and either evicts the oldest complete turn (if `compactor=None`) or triggers background compaction to replace old turns with summary messages.

### What happens to pending tool calls when conversation turns are evicted?

The `Chat` class maintains a `_pending_tool_calls` registry that tracks function calls awaiting their output. If a turn containing a tool call is evicted before the function returns, the metadata is preserved. When the function output arrives via `add_item()`, the system handles re-injection of the call context so the LLM can still correlate the result with the original request.

### How do I enable conversation memory compression instead of deletion?

Import `build_compactor` from [`speech_to_speech/LLM/compaction_prompt.py`](https://github.com/huggingface/speech-to-speech/blob/main/speech_to_speech/LLM/compaction_prompt.py) and provide a `generate_fn` callable that accepts `(system_prompt, user_prompt)` and returns a JSON string with `user_summary` and `assistant_summary` keys. Pass the resulting compactor to `chat.trim_if_needed(compactor=my_compactor)` to activate background thread summarization via the `COMPACTION_SYSTEM_PROMPT` template.