# How to Implement Push-to-Talk Functionality in LiveKit Voice Agents

> Implement push-to-talk in LiveKit Agents. Learn to control microphone input with RPCs for seamless voice agent interactions. Get started today!

- Repository: [LiveKit/agents](https://github.com/livekit/agents)
- Tags: how-to-guide
- Published: 2026-03-06

---

**To implement push-to-talk in LiveKit Agents, create an `AgentSession` with manual turn detection, disable audio input by default, and expose RPC methods that toggle the microphone on button press and release.**

The `livekit/agents` repository provides a voice agent framework that supports explicit control over when the agent listens to participants. By disabling automatic voice activity detection (VAD) and managing audio input state programmatically, you can build a reliable push-to-talk (PTT) system for single or multi-participant rooms.

## Configure Manual Turn Detection

The foundation of push-to-talk functionality is disabling the automatic turn detection that normally triggers when a user speaks. In `AgentSession.__init__`, pass `turn_detection="manual"` to prevent the session from automatically starting or stopping user turns.

```python
from livekit.agents.voice import AgentSession

session = AgentSession(turn_detection="manual")

```

This configuration is defined in [`livekit-agents/livekit/agents/voice/agent_session.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/voice/agent_session.py) at lines 181–190, where the `turn_detection` parameter accepts `"manual"` or `"automatic"` modes.

## Control Audio Input State

With manual turn detection enabled, you must explicitly control when the microphone is active. Start the session with the audio input disabled to ensure no audio streams until the user explicitly requests it.

```python

# Disable microphone at startup

session.input.set_audio_enabled(False)

```

The `set_audio_enabled` method is implemented in [`livekit-agents/livekit/agents/voice/io.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/voice/io.py) at lines 56–72. This method toggles the underlying audio stream, allowing you to mute and unmute the input dynamically during the session.

## Expose RPC Methods for Client Control

To enable the front-end to control the push-to-talk state, register RPC methods that the client can invoke when the user presses or releases the talk button.

### Start Turn

When the user presses the talk button, the `start_turn` RPC method should interrupt any ongoing agent response, clear the previous user turn, bind the turn to the specific participant, and enable the microphone.

```python
from livekit import rtc

@ctx.room.local_participant.register_rpc_method("start_turn")
async def start_turn(data: rtc.RpcInvocationData):
    session.interrupt()
    session.clear_user_turn()
    session.room_io.set_participant(data.caller_identity)
    session.input.set_audio_enabled(True)

```

### End Turn

When the user releases the button, the `end_turn` RPC method disables the microphone and commits the captured audio to a transcript. The `commit_user_turn` method accepts `transcript_timeout` and `stt_flush_duration` parameters to handle STT finalization.

```python
@ctx.room.local_participant.register_rpc_method("end_turn")
async def end_turn(data: rtc.RpcInvocationData):
    session.input.set_audio_enabled(False)
    try:
        user_transcript = await session.commit_user_turn(
            transcript_timeout=5.0,       # wait for STT to finish

            stt_flush_duration=2.0,       # add silence so STT finalises

        )
        logger.info(f"user transcript: {user_transcript}")
    except Exception as e:
        logger.error("error committing user turn", exc_info=e)

```

### Cancel Turn

Optionally, expose a `cancel_turn` method that allows the UI to abort a turn without sending a transcript to the LLM.

```python
@ctx.room.local_participant.register_rpc_method("cancel_turn")
async def cancel_turn(data: rtc.RpcInvocationData):
    session.input.set_audio_enabled(False)
    session.clear_user_turn()
    logger.info("cancel turn")

```

## Complete Working Example

The `livekit/agents` repository provides a complete implementation in [`examples/voice_agents/push_to_talk.py`](https://github.com/livekit/agents/blob/main/examples/voice_agents/push_to_talk.py). The example demonstrates how to wire the session creation, RPC registration, and agent logic together.

Key implementation details from the example include:

- Setting `attributes={"push-to-talk": "1"}` in the job request handler to signal the front-end that PTT is available
- Using `session.room_io.set_participant()` to bind turns to specific caller identities in multi-user rooms
- Handling empty transcripts in `on_user_turn_completed` to prevent the agent from responding to silence

## Summary

- **Manual turn detection** is required to disable automatic VAD-based turn management; initialize `AgentSession` with `turn_detection="manual"`.
- **Audio input control** is handled via `session.input.set_audio_enabled()`, which must be called to mute the microphone at startup and unmute it during PTT activation.
- **RPC methods** bridge the front-end UI to the agent, with `start_turn` enabling the microphone and `end_turn` committing the audio to a transcript via `session.commit_user_turn()`.
- **Multi-participant support** is achieved by calling `session.room_io.set_participant()` with the caller identity from the RPC invocation data.

## Frequently Asked Questions

### What is the difference between manual and automatic turn detection in LiveKit Agents?

Automatic turn detection uses voice activity detection (VAD) to automatically start and stop user turns when speech is detected. Manual turn detection, configured with `turn_detection="manual"`, disables this automation and requires explicit RPC calls or method invocations to control when the agent listens and when the turn ends.

### How do I handle multiple participants with push-to-talk?

When a participant invokes the `start_turn` RPC method, extract their identity from `data.caller_identity` and pass it to `session.room_io.set_participant(data.caller_identity)`. This binds the current turn to that specific participant, ensuring that audio and transcripts are attributed correctly in multi-user rooms.

### What are the recommended timeout values for commit_user_turn?

The example implementation uses `transcript_timeout=5.0` seconds to wait for the speech-to-text (STT) provider to finalize the transcript, and `stt_flush_duration=2.0` seconds to append silence to the audio buffer, ensuring the STT engine recognizes the end of speech. Adjust these values based on your STT provider's latency and your application's responsiveness requirements.

### Can I implement push-to-talk without using RPC methods?

While RPC methods provide the cleanest integration with front-end clients, you could implement similar functionality using data messages or by managing state through a separate signaling mechanism. However, the `livekit/agents` framework is designed around RPC methods for this use case, as they provide type-safe, request-response semantics that are ideal for push-to-talk button states.