How to Use Structured Output from an LLM to Guide TTS in LiveKit Agents
To use structured output from an LLM to guide TTS in LiveKit Agents, define a TypedDict schema containing voice instructions, request it via response_format in LLM.chat, parse the streaming JSON fragments with pydantic_core.from_json, and dynamically update the TTS configuration using tts.update_options() before synthesizing speech.
LiveKit Agents enables you to build voice-first assistants where the language model not only generates responses but also supplies metadata to control vocal expression. By leveraging structured output, you can instruct the LLM to return JSON objects containing both the spoken text and TTS directives (tone, emotion, speed), then apply those directives in real-time during audio synthesis.
Define the Structured Response Schema
First, create a TypedDict that specifies the JSON structure the LLM must return. This schema separates the vocal performance instructions from the actual content.
from typing import TypedDict, Annotated
from pydantic import Field
class ResponseEmotion(TypedDict):
voice_instructions: Annotated[
str,
Field(..., description="Concise TTS directive for tone, emotion, intonation, and speed"),
]
response: str
The voice_instructions field provides the directive for the TTS provider, while response contains the text to be spoken. This definition is found in examples/voice_agents/structured_output.py at lines 32-38.
Override the LLM Node to Request Structured Output
Override the llm_node method in your Agent subclass to request the schema via the response_format parameter. Cast the LLM to openai.LLM to access provider-specific features.
from typing import cast
from livekit.plugins import openai
from ._utils import NOT_GIVEN
async def llm_node(self, chat_ctx, tools, model_settings):
llm = cast(openai.LLM, self.llm)
tool_choice = model_settings.tool_choice if model_settings else NOT_GIVEN
async with llm.chat(
chat_ctx=chat_ctx,
tools=tools,
tool_choice=tool_choice,
response_format=ResponseEmotion, # Request structured JSON output
) as stream:
async for chunk in stream:
yield chunk
This implementation is located at lines 84-94 in structured_output.py.
Parse Streaming JSON with process_structured_output
Because the LLM streams partial JSON fragments, you need a helper to accumulate chunks, attempt deserialization, and separate the voice instructions from the spoken text. The process_structured_output helper handles this using pydantic_core.from_json with allow_partial="trailing-strings".
from typing import AsyncIterable, Callable
from pydantic_core import from_json
async def process_structured_output(
text: AsyncIterable[str],
callback: Callable[[ResponseEmotion], None] | None = None,
) -> AsyncIterable[str]:
last_response = ""
acc_text = ""
async for chunk in text:
acc_text += chunk
try:
resp: ResponseEmotion = from_json(acc_text, allow_partial="trailing-strings")
except ValueError:
continue
if callback:
callback(resp)
if not resp.get("response"):
continue
new_delta = resp["response"][len(last_response):]
if new_delta:
yield new_delta
last_response = resp["response"]
This function:
- Accumulates streaming text until valid JSON is parseable
- Invokes a callback with the full
ResponseEmotionobject once available - Yields only the
responsetext field for downstream processing
The source code resides at lines 40-63 in examples/voice_agents/structured_output.py.
Apply Voice Instructions in the TTS Node
Override tts_node to intercept the stream, parse the structured output, and update the TTS options before synthesis begins. The callback executes once—the first time both voice_instructions and a non-empty response are present—then calls tts.update_options() to inject the directives into the OpenAI TTS provider.
async def tts_node(self, text, model_settings):
instruction_updated = False
def output_processed(resp: ResponseEmotion):
nonlocal instruction_updated
if resp.get("voice_instructions") and resp.get("response") and not instruction_updated:
instruction_updated = True
logger.info(f'Applying TTS instructions: "{resp["voice_instructions"]}"')
tts = cast(openai.TTS, self.tts)
tts.update_options(instructions=resp["voice_instructions"])
# Strip instructions and pass only verbal content to the default TTS node
return Agent.default.tts_node(
self,
process_structured_output(text, callback=output_processed),
model_settings,
)
This pattern ensures the voice_instructions are applied to the TTS instance via update_options(), while the Agent.default.tts_node receives only the clean text for synthesis. See lines 96-117 in structured_output.py.
Filter Transcriptions for Clean Logs
To prevent the raw JSON schema from appearing in conversation transcripts, optionally override transcription_node to strip the metadata before logging.
async def transcription_node(self, text, model_settings):
return Agent.default.transcription_node(
self,
process_structured_output(text), # Remove TTS directives from logs
model_settings,
)
This implementation is available at lines 119-124 in the example file.
Summary
- Structured output allows the LLM to return JSON containing both
voice_instructions(TTS metadata) andresponse(spoken text). - Override
llm_nodeto passresponse_format=ResponseEmotionwhen callingLLM.chat(). - Use
process_structured_outputto handle streaming JSON parsing and separate instructions from content. - Override
tts_nodeto calltts.update_options(instructions=...)with the parsed voice directives before synthesizing audio. - The implementation relies on
openai.LLMandopenai.TTSfrom the LiveKit OpenAI plugin, specifically theupdate_optionsmethod available in that provider.
Frequently Asked Questions
Which TTS providers support dynamic voice instructions via update_options?
The update_options() method is implemented in the OpenAI TTS plugin (livekit-plugins/openai), which passes instructions directly to OpenAI's TTS API to control tone, speed, and style. Other TTS providers in the LiveKit ecosystem may implement similar methods, but you should verify their specific plugin implementations in livekit-plugins/.
Can I change voice instructions mid-conversation?
Yes. Because the pipeline processes streaming chunks asynchronously, the callback in tts_node can trigger update_options() at any point during generation. This enables dynamic emotional shifts—such as switching from calm to excited tone—without restarting the agent or interrupting the conversation flow.
What happens if the LLM returns malformed JSON?
The process_structured_output helper uses pydantic_core.from_json with allow_partial="trailing-strings" to handle incomplete JSON fragments gracefully. If the accumulated text is not yet valid JSON, it catches the ValueError and continues accumulating chunks until deserialization succeeds.
Is the ResponseEmotion schema flexible enough for custom TTS parameters?
Absolutely. The TypedDict schema is fully customizable—you can add fields for specific parameters like speed, pitch, or emotion_label as long as your TTS provider supports them in update_options(). Simply extend the schema definition and update the callback logic to map new fields to the appropriate TTS configuration options.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →