# How to Use Structured Output from an LLM to Guide TTS in LiveKit Agents

> Learn to guide TTS in LiveKit Agents using structured LLM output. Define schemas, parse JSON, and dynamically update TTS for better voice control. Integrate LLM responses seamlessly into your audio applications.

- Repository: [LiveKit/agents](https://github.com/livekit/agents)
- Tags: how-to-guide
- Published: 2026-03-06

---

**To use structured output from an LLM to guide TTS in LiveKit Agents, define a `TypedDict` schema containing voice instructions, request it via `response_format` in `LLM.chat`, parse the streaming JSON fragments with `pydantic_core.from_json`, and dynamically update the TTS configuration using `tts.update_options()` before synthesizing speech.**

LiveKit Agents enables you to build voice-first assistants where the language model not only generates responses but also supplies metadata to control vocal expression. By leveraging **structured output**, you can instruct the LLM to return JSON objects containing both the spoken text and TTS directives (tone, emotion, speed), then apply those directives in real-time during audio synthesis.

## Define the Structured Response Schema

First, create a `TypedDict` that specifies the JSON structure the LLM must return. This schema separates the vocal performance instructions from the actual content.

```python
from typing import TypedDict, Annotated
from pydantic import Field

class ResponseEmotion(TypedDict):
    voice_instructions: Annotated[
        str,
        Field(..., description="Concise TTS directive for tone, emotion, intonation, and speed"),
    ]
    response: str

```

The `voice_instructions` field provides the directive for the TTS provider, while `response` contains the text to be spoken. This definition is found in [`examples/voice_agents/structured_output.py`](https://github.com/livekit/agents/blob/main/examples/voice_agents/structured_output.py) at lines 32-38.

## Override the LLM Node to Request Structured Output

Override the `llm_node` method in your `Agent` subclass to request the schema via the `response_format` parameter. Cast the LLM to `openai.LLM` to access provider-specific features.

```python
from typing import cast
from livekit.plugins import openai
from ._utils import NOT_GIVEN

async def llm_node(self, chat_ctx, tools, model_settings):
    llm = cast(openai.LLM, self.llm)
    tool_choice = model_settings.tool_choice if model_settings else NOT_GIVEN
    
    async with llm.chat(
        chat_ctx=chat_ctx,
        tools=tools,
        tool_choice=tool_choice,
        response_format=ResponseEmotion,  # Request structured JSON output

    ) as stream:
        async for chunk in stream:
            yield chunk

```

This implementation is located at lines 84-94 in [`structured_output.py`](https://github.com/livekit/agents/blob/main/structured_output.py).

## Parse Streaming JSON with process_structured_output

Because the LLM streams partial JSON fragments, you need a helper to accumulate chunks, attempt deserialization, and separate the voice instructions from the spoken text. The `process_structured_output` helper handles this using `pydantic_core.from_json` with `allow_partial="trailing-strings"`.

```python
from typing import AsyncIterable, Callable
from pydantic_core import from_json

async def process_structured_output(
    text: AsyncIterable[str],
    callback: Callable[[ResponseEmotion], None] | None = None,
) -> AsyncIterable[str]:
    last_response = ""
    acc_text = ""
    
    async for chunk in text:
        acc_text += chunk
        try:
            resp: ResponseEmotion = from_json(acc_text, allow_partial="trailing-strings")
        except ValueError:
            continue

        if callback:
            callback(resp)

        if not resp.get("response"):
            continue

        new_delta = resp["response"][len(last_response):]
        if new_delta:
            yield new_delta
        last_response = resp["response"]

```

This function:
- Accumulates streaming text until valid JSON is parseable
- Invokes a callback with the full `ResponseEmotion` object once available
- Yields only the `response` text field for downstream processing

The source code resides at lines 40-63 in [`examples/voice_agents/structured_output.py`](https://github.com/livekit/agents/blob/main/examples/voice_agents/structured_output.py).

## Apply Voice Instructions in the TTS Node

Override `tts_node` to intercept the stream, parse the structured output, and update the TTS options before synthesis begins. The callback executes once—the first time both `voice_instructions` and a non-empty `response` are present—then calls `tts.update_options()` to inject the directives into the OpenAI TTS provider.

```python
async def tts_node(self, text, model_settings):
    instruction_updated = False

    def output_processed(resp: ResponseEmotion):
        nonlocal instruction_updated
        if resp.get("voice_instructions") and resp.get("response") and not instruction_updated:
            instruction_updated = True
            logger.info(f'Applying TTS instructions: "{resp["voice_instructions"]}"')
            tts = cast(openai.TTS, self.tts)
            tts.update_options(instructions=resp["voice_instructions"])

    # Strip instructions and pass only verbal content to the default TTS node

    return Agent.default.tts_node(
        self,
        process_structured_output(text, callback=output_processed),
        model_settings,
    )

```

This pattern ensures the `voice_instructions` are applied to the TTS instance via `update_options()`, while the `Agent.default.tts_node` receives only the clean text for synthesis. See lines 96-117 in [`structured_output.py`](https://github.com/livekit/agents/blob/main/structured_output.py).

## Filter Transcriptions for Clean Logs

To prevent the raw JSON schema from appearing in conversation transcripts, optionally override `transcription_node` to strip the metadata before logging.

```python
async def transcription_node(self, text, model_settings):
    return Agent.default.transcription_node(
        self,
        process_structured_output(text),  # Remove TTS directives from logs

        model_settings,
    )

```

This implementation is available at lines 119-124 in the example file.

## Summary

- **Structured output** allows the LLM to return JSON containing both `voice_instructions` (TTS metadata) and `response` (spoken text).
- Override **`llm_node`** to pass `response_format=ResponseEmotion` when calling `LLM.chat()`.
- Use **`process_structured_output`** to handle streaming JSON parsing and separate instructions from content.
- Override **`tts_node`** to call `tts.update_options(instructions=...)` with the parsed voice directives before synthesizing audio.
- The implementation relies on **`openai.LLM`** and **`openai.TTS`** from the LiveKit OpenAI plugin, specifically the `update_options` method available in that provider.

## Frequently Asked Questions

### Which TTS providers support dynamic voice instructions via update_options?

The `update_options()` method is implemented in the OpenAI TTS plugin (`livekit-plugins/openai`), which passes instructions directly to OpenAI's TTS API to control tone, speed, and style. Other TTS providers in the LiveKit ecosystem may implement similar methods, but you should verify their specific plugin implementations in `livekit-plugins/`.

### Can I change voice instructions mid-conversation?

Yes. Because the pipeline processes streaming chunks asynchronously, the callback in `tts_node` can trigger `update_options()` at any point during generation. This enables dynamic emotional shifts—such as switching from calm to excited tone—without restarting the agent or interrupting the conversation flow.

### What happens if the LLM returns malformed JSON?

The `process_structured_output` helper uses `pydantic_core.from_json` with `allow_partial="trailing-strings"` to handle incomplete JSON fragments gracefully. If the accumulated text is not yet valid JSON, it catches the `ValueError` and continues accumulating chunks until deserialization succeeds.

### Is the ResponseEmotion schema flexible enough for custom TTS parameters?

Absolutely. The `TypedDict` schema is fully customizable—you can add fields for specific parameters like `speed`, `pitch`, or `emotion_label` as long as your TTS provider supports them in `update_options()`. Simply extend the schema definition and update the callback logic to map new fields to the appropriate TTS configuration options.