How to Configure Job Scheduling and Dispatch APIs in LiveKit Agents

LiveKit Agents uses a worker-based architecture where you configure scheduling behavior through ServerOptions and control job placement via the explicit AgentDispatch API.

The livekit/agents repository provides a Python framework for building voice and multimodal AI agents. Understanding how to configure job scheduling and dispatch APIs allows you to optimize resource utilization, implement custom load balancing, and selectively assign agents to specific rooms rather than accepting every incoming request.

Configuring Job Scheduling Behavior

The worker process in livekit-agents/livekit/agents/worker.py manages the lifecycle of agent jobs. You tune its scheduling decisions by passing a custom ServerOptions object to the AgentServer.

Load-Based Availability

The worker determines its availability based on a load value between 0 and 1. By default, it uses _DefaultLoadCalc.get_load() in worker.py (lines 4–13), which calculates a CPU-based moving average.

To implement custom load logic, provide a load_fnc callable:

def my_load_calc(server) -> float:
    # Example: measure queued jobs as percentage of capacity

    return len(server.active_jobs) / 20.0

options = ServerOptions(
    load_fnc=my_load_calc,
    load_threshold=0.5,  # Reject new jobs above 50% load

)

The load_threshold parameter (default 0.7 in production, disabled/∞ in development) determines when the worker marks itself unavailable. This is defined in ServerOptions at lines 65–70 of worker.py.

Idle Process Management

To minimize cold-start latency, the worker maintains a pool of idle processes. The num_idle_processes setting (lines 78–82 in worker.py) defaults to a value adapted to your CPU core count. Increase this for high-traffic scenarios or decrease it to conserve memory.

Request Filtering and Prewarming

The request_fnc parameter (lines 54–58) accepts an async function that inspects each JobRequest and decides whether to accept or reject it. By default, all jobs are accepted. For initialization tasks that should run once per process (such as loading ML models), use prewarm_fnc (lines 58–60).

When a request arrives, the worker calls AgentServer._answer_availability (lines 59–84) to evaluate these functions and respond with availability status.

Enabling Explicit Dispatch Mode

By default, LiveKit auto-dispatches agents to new rooms. To gain control over which agent runs where, disable this behavior by setting agent_name in ServerOptions:

options = ServerOptions(
    entrypoint_fnc=my_agent,
    agent_name="telephony-agent",  # Only accepts explicit dispatches

)

When agent_name is non-empty (defined at lines 90–92 in worker.py), the worker filters incoming jobs and only accepts those explicitly dispatched to that specific name.

Dispatching Jobs via the API

With explicit dispatch enabled, use the AgentDispatch endpoint to assign jobs to specific rooms. The examples/bank-ivr/dial_bank_agent.py file (lines 24–28) demonstrates this pattern:

from livekit import api

lkapi = api.LiveKitAPI()
dispatch = await lkapi.agent_dispatch.create_dispatch(
    api.CreateAgentDispatchRequest(
        agent_name="telephony-agent",
        room="support-room-123",
        metadata="Priority customer request"
    )
)

The returned Dispatch object contains a unique dispatch ID. The worker receives this ID in the job request, validates it against the configured agent_name, and launches your agent entrypoint.

Complete Implementation Example

Here is a full configuration combining custom scheduling with explicit dispatch:

from livekit.agents import ServerOptions, AgentServer, JobContext
import psutil

def system_load_calc(server) -> float:
    # Combine CPU and memory pressure

    cpu = psutil.cpu_percent(interval=0.1) / 100
    mem = psutil.virtual_memory().percent / 100
    return max(cpu, mem)

async def prewarm():
    # Load heavy models once per process

    import whisper
    whisper.load_model("base")

options = ServerOptions(
    entrypoint_fnc=my_agent,
    agent_name="support-agent",
    load_fnc=system_load_calc,
    load_threshold=0.6,
    num_idle_processes=4,
    prewarm_fnc=prewarm,
)

server = AgentServer.from_server_options(options)
await server.run()

Summary

  • Configure scheduling via ServerOptions in worker.py: use load_fnc and load_threshold for capacity management, num_idle_processes for warm pools, and prewarm_fnc for initialization.
  • Filter requests by implementing request_fnc to inspect JobRequest objects before acceptance.
  • Enable explicit dispatch by setting a non-empty agent_name, which forces the worker to ignore auto-dispatched jobs.
  • Trigger jobs programmatically using LiveKitAPI.agent_dispatch.create_dispatch to target specific rooms with metadata.

Frequently Asked Questions

How does the worker decide when to reject new jobs?

The worker evaluates the load_fnc return value against load_threshold. If the load exceeds the threshold (default 0.7 in production), AgentServer._answer_availability marks the worker as unavailable. You can provide a custom load_fnc that measures queue depth, memory usage, or GPU utilization instead of the default CPU-based calculation.

What is the difference between auto-dispatch and explicit dispatch?

Auto-dispatch automatically assigns an agent to every new room. Explicit dispatch requires you to set agent_name in ServerOptions and call create_dispatch via the API, allowing you to control exactly which agent type runs in which room and when jobs start.

Can I run different agent types on the same worker?

No. A single worker process is configured with one entrypoint_fnc and optional agent_name. To run multiple agent types, deploy separate workers with distinct agent_name values and dispatch jobs to the appropriate name based on your routing logic.

What happens if a dispatch targets an unavailable worker?

If all workers registered under the target agent_name report load above their threshold, the dispatch remains pending until a worker becomes available or the request times out. The _exceptions.py file defines timeout errors that occur when assignment fails due to capacity constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →