How to Configure Budget Limits and Spend Alerts in LiteLLM

LiteLLM provides a three-layer budget management system that lets you set spending caps, receive alerts via Slack or email, and enforce limits per-model, per-team, or per-user by combining the budget manager, limit enforcement hooks, and alert dispatchers.

The BerriAI/litellm repository includes a comprehensive budget manager designed to prevent runaway costs across OpenAI, Azure, Anthropic, and other LLM providers. Whether you need to cap spending by user session, team, or specific model, LiteLLM’s proxy server can block requests that exceed financial thresholds and notify stakeholders through configurable alert channels.

How the Budget System Works

LiteLLM implements spending controls through three coordinated layers:

When a request hits the proxy, the system executes a four-step sequence: Lookup (resolve the budget scope from budget_management_endpoints.py), Check (compare current spend against limits), Allow/Block (return HTTP 429 if over budget), and Update (record the actual cost after successful LLM calls).

Configure Budget Limits via the Management API

Define Spending Caps and Warning Thresholds

Create budgets programmatically by posting to the budget management endpoint defined in litellm/proxy/management_endpoints/budget_management_endpoints.py:

curl -X POST https://<proxy-host>/budget \
  -H "Authorization: Bearer <admin-key>" \
  -H "Content-Type: application/json" \
  -d '{
        "budget_name": "team_alpha_budget",
        "budget_type": "team",
        "identifier": "team_alpha",
        "max_budget": 500.0,
        "soft_budget": 400.0,
        "duration": "monthly",
        "alert_channels": ["slack", "email"]
      }'

Valid budget_type values include team, organization, user, tag, and model. The max_budget acts as a hard ceiling that blocks requests, while soft_budget triggers alerts without interrupting service.

Set Alert Channels and Notification Preferences

The alert_channels array maps to dispatchers in budget_alert_types.py. Supported channels include Slack webhooks, SMTP email, and custom webhooks. To add custom notification services, extend the ALERT_DISPATCHERS dictionary in the budget_alert_types.py module.

Enforce Budget Limits Programmatically

Use the Python SDK to set per-user budgets and per-session caps for granular control over individual request costs:

import litellm

# Create a per-user monthly budget (soft $20, hard $30)

litellm.set_budget(
    budget_type="user",
    identifier="user_12345",
    max_budget=30.0,
    soft_budget=20.0,
    duration="monthly",
    alert_channels=["slack"]
)

# Enforce a per-session limit of $5

litellm.set_session_budget(
    user_id="user_12345",
    session_max_budget=5.0
)

# Requests exceeding the session budget are rejected automatically

response = litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain quantum entanglement"}]
)

The session-level enforcement logic resides in litellm/proxy/hooks/max_budget_per_session_limiter.py, which checks the running total for each conversation thread before allowing the request to proceed.

Configure Global Budget Defaults

Set organization-wide fallback values using environment variables that budget_manager.py reads at startup:

Variable Purpose
MAX_GLOBAL_BUDGET Hard ceiling applied when no explicit budget is defined for a scope.
SOFT_GLOBAL_BUDGET Warning threshold used for implicit budgets.
BUDGET_RESET_CRON Cron expression defining when usage counters reset (e.g., daily).

These defaults ensure that unconfigured teams or users still operate under spending guardrails.

Manage Budgets in the LiteLLM Dashboard

The React-based administrative interface provides visual controls for non-technical stakeholders:

Key Implementation Files

Understanding the source architecture helps when debugging limit violations or extending functionality:

File Role
litellm/budget_manager.py Central persistence and spend aggregation
litellm/router_strategy/budget_limiter.py Core request-level checks and throttling logic
litellm/proxy/hooks/max_budget_limiter.py Global hard-budget enforcement
litellm/proxy/hooks/max_budget_per_session_limiter.py Session-scoped spending caps
litellm/proxy/hooks/model_max_budget_limiter.py Model-specific budget overrides
litellm/integrations/SlackAlerting/budget_alert_types.py Alert channel dispatch mapping
litellm/proxy/management_endpoints/budget_management_endpoints.py REST API for CRUD operations

Summary

  • Configure budget limits and spend alerts in LiteLLM using either the REST API (/budget endpoint), Python SDK (set_budget()), or React dashboard.
  • The system distinguishes between hard limits (max_budget) that return HTTP 429 errors and soft limits (soft_budget) that trigger Slack/email notifications.
  • Enforcement occurs through specialized hooks including max_budget_limiter.py and max_budget_per_session_limiter.py.
  • Global defaults via environment variables provide fallback protection for unconfigured scopes.

Frequently Asked Questions

What happens when a LiteLLM budget limit is exceeded?

When a request would exceed the max_budget, the proxy returns an HTTP 429 error and blocks the LLM call. If the request crosses only the soft_budget but remains under the hard limit, it proceeds while the system queues an alert notification to the configured channels.

Can I set temporary budget increases for specific teams?

Yes. According to the temporary_budget_increase.md documentation, administrators can issue temporary budget overrides for specific time windows. This is useful for handling traffic spikes or special projects without permanently increasing baseline spending caps.

Which alert channels does LiteLLM support for budget notifications?

LiteLLM supports Slack webhooks, SMTP email, and custom webhooks out of the box. The mapping between channel names and sender implementations lives in budget_alert_types.py, where you can extend the ALERT_DISPATCHERS dictionary to integrate with PagerDuty, Microsoft Teams, or internal incident management systems.

How do per-session budgets differ from per-user budgets?

Per-user budgets track cumulative spending across all sessions for that identifier over a defined duration (daily, weekly, monthly). Per-session budgets, enforced by max_budget_per_session_limiter.py, apply only to a single conversation thread or request sequence, preventing individual long-running chats from consuming the entire monthly allocation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →