How to Handle Long-Running Function Calls in LiveKit Agents: A Complete Guide
LiveKit Agents handles long-running function calls by running LLM tools asynchronously alongside speech generation, allowing you to cancel operations on user interruption or protect critical tasks from cancellation using RunContext and SpeechHandle APIs.
When building voice AI agents with the LiveKit Agents framework, you will inevitably encounter operations that take several seconds to complete—web searches, database queries, or external API calls. The framework provides specific patterns to handle these long-running function calls without blocking the conversational flow, while giving you fine-grained control over interruption behavior.
Understanding the Architecture for Long-Running Operations
LiveKit Agents couples tool execution with the speech generation pipeline through two primary abstractions: RunContext and SpeechHandle. Understanding how these interact is essential for implementing robust long-running workflows.
The Role of RunContext
The RunContext object is passed automatically to every tool decorated with @function_tool. Located in livekit-agents/livekit/agents/voice/events.py, this class provides access to the current SpeechHandle and exposes methods to control interruption behavior. When a long-running function is invoked, the RunContext serves as the bridge between your async operation and the agent's speech state.
SpeechHandle and Interruption Management
The SpeechHandle class, defined in livekit-agents/livekit/agents/voice/speech_handle.py, manages the lifecycle of individual speech turns. It maintains internal futures for tracking interruptions (_interrupt_fut), completion (_done_fut), and scheduling (_scheduled_fut). The critical method wait_if_not_interrupted() allows you to await one or more futures only as long as the current speech hasn't been interrupted, providing the core mechanism for cancellable long-running operations.
Implementing Interruptible Long-Running Function Calls
For most use cases—such as web searches or data retrieval—you want the operation to continue while the assistant speaks, but cancel immediately if the user interrupts. This pattern prevents stale data from being injected into a conversation that has already moved on.
The implementation follows a four-step pattern as demonstrated in examples/voice_agents/long_running_function.py:
import asyncio
import logging
from livekit.agents import Agent, RunContext
from livekit.agents.llm import function_tool
logger = logging.getLogger("long-run-example")
class MyAgent(Agent):
def __init__(self) -> None:
super().__init__(instructions="You are a voice assistant.")
@function_tool
async def search_web(self, query: str, run_ctx: RunContext) -> str | None:
"""A tool that performs a potentially slow web search."""
logger.info(f"Searching the web for: {query}")
# 1. Start the long-running coroutine without awaiting it yet
long_task = asyncio.ensure_future(self._slow_search(query))
# 2. Wait for the task unless the speech gets interrupted
await run_ctx.speech_handle.wait_if_not_interrupted([long_task])
# 3. If interrupted, cancel the search and return None
if run_ctx.speech_handle.interrupted:
logger.info(f"Search for '{query}' was interrupted by the user.")
long_task.cancel()
return None
# 4. Otherwise, fetch the result and return it to the LLM
result = long_task.result()
logger.info(f"Search completed: {result}")
return result
async def _slow_search(self, query: str) -> str:
"""Simulated long-running operation (e.g., external API call)."""
await asyncio.sleep(5) # 5 seconds of work
return f"Fake results for "{query}""
Key implementation details:
asyncio.ensure_futurestarts the coroutine immediately without blocking, allowing the function to set up interruption handling before the work begins.wait_if_not_interruptedusesasyncio.shieldinternally to protect the underlying task from cancellation while still allowing the framework to detect interruptions. If the user speaks over the assistant, the method returns immediately, leaving the background task running until you explicitly cancel it.run_ctx.speech_handle.interruptedprovides a boolean check to determine whether the return fromwait_if_not_interruptedwas due to completion or interruption.
Protecting Critical Operations from Interruption
Some operations must complete regardless of user behavior—financial transactions, database writes, or state-changing commands. For these scenarios, LiveKit Agents provides disallow_interruptions().
Located in the same RunContext class in livekit-agents/livekit/agents/voice/events.py, this method disables interruption for the duration of the tool execution:
@function_tool
async def process_payment(self, amount: float, run_ctx: RunContext) -> str:
"""Process a payment that must complete even if user interrupts."""
# Prevent any interruption during this critical operation
run_ctx.disallow_interruptions()
# Now safe to await the long-running task
result = await self._charge_card(amount)
return f"Payment processed: {result}"
Important considerations:
- Once
disallow_interruptions()is called, the user cannot stop the current speech turn until the tool returns. - Use this sparingly—only for operations where partial completion or cancellation would leave the system in an inconsistent state.
- The method affects only the current
RunContext; subsequent tool calls or speech turns revert to the default interruptible behavior.
When to Use Each Pattern
Choosing the right interruption strategy depends on the nature of your long-running operation and the user experience you want to provide.
| Situation | Recommended Approach | Implementation |
|---|---|---|
| Typical background work (search, query, fetch) | Allow interruption, cancel task | await run_ctx.speech_handle.wait_if_not_interrupted([task]) + check interrupted flag |
| Critical state-changing operations (payments, writes) | Disallow interruption entirely | run_ctx.disallow_interruptions() before awaiting |
| Fire-and-forget tasks (logging, analytics) | Don't wait, return immediately | asyncio.create_task() without adding to wait_if_not_interrupted |
| Very long tasks (minutes-long processing) | Background worker + follow-up | Spawn separate worker, return job ID, use follow-up message |
Summary
- LiveKit Agents executes LLM tools asynchronously alongside speech generation, enabling non-blocking long-running operations.
- Use
run_ctx.speech_handle.wait_if_not_interrupted()to await tasks that should cancel if the user interrupts the assistant, checking theinterruptedproperty to handle cleanup. - Call
run_ctx.disallow_interruptions()before awaiting critical operations that must complete regardless of user behavior. - The framework uses
asyncio.shieldinternally to protect underlying tasks while still allowing interruption detection, with a 5-second safety timeout for cleanup.
Frequently Asked Questions
What happens if a user interrupts a long-running function call?
If the user speaks over the assistant while a tool is executing, the SpeechHandle sets its internal _interrupt_fut and marks the handle as interrupted. When using wait_if_not_interrupted(), the method returns immediately, allowing your code to check run_ctx.speech_handle.interrupted and cancel the background task. This prevents stale results from being injected into a conversation that has already moved on.
How do I prevent a function from being cancelled mid-execution?
For critical operations like payments or database writes, call run_ctx.disallow_interruptions() before awaiting your long-running task. This method, defined in livekit-agents/livekit/agents/voice/events.py, disables the interruption mechanism for the current tool execution, ensuring the operation completes even if the user attempts to speak over the assistant.
Can I run multiple long-running functions simultaneously?
Yes. The wait_if_not_interrupted() method accepts a list of awaitables, allowing you to pass multiple futures or tasks. The framework uses asyncio.gather() with asyncio.shield() internally to wait for all tasks while still monitoring for interruptions. If an interruption occurs, you receive control back immediately and can cancel individual tasks as needed.
Where can I find the reference implementation for long-running tools?
The official reference implementation is located at examples/voice_agents/long_running_function.py in the livekit/agents repository. This file demonstrates the complete pattern of starting a background task, using wait_if_not_interrupted(), and handling the interruption flag. For the underlying APIs, examine livekit-agents/livekit/agents/voice/speech_handle.py for SpeechHandle and livekit-agents/livekit/agents/voice/events.py for RunContext.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →