How to Configure Budget Limits and Spend Alerts in LiteLLM
LiteLLM provides a three-layer budget management system that lets you set spending caps, receive alerts via Slack or email, and enforce limits per-model, per-team, or per-user by combining the budget manager, limit enforcement hooks, and alert dispatchers.
The BerriAI/litellm repository includes a comprehensive budget manager designed to prevent runaway costs across OpenAI, Azure, Anthropic, and other LLM providers. Whether you need to cap spending by user session, team, or specific model, LiteLLM’s proxy server can block requests that exceed financial thresholds and notify stakeholders through configurable alert channels.
How the Budget System Works
LiteLLM implements spending controls through three coordinated layers:
- Budget Store – Persists budget definitions in the proxy database via
litellm/budget_manager.py, storingmax_budget,soft_budget, and duration intervals. - Limit Enforcement – Intercepts each request using
litellm/router_strategy/budget_limiter.pyand specialized hook modules (max_budget_limiter.py,max_budget_per_session_limiter.py,model_max_budget_limiter.py) to approve or deny calls. - Alerting – Dispatches notifications through
litellm/integrations/SlackAlerting/budget_alert_types.pywhen spending crosses warning thresholds.
When a request hits the proxy, the system executes a four-step sequence: Lookup (resolve the budget scope from budget_management_endpoints.py), Check (compare current spend against limits), Allow/Block (return HTTP 429 if over budget), and Update (record the actual cost after successful LLM calls).
Configure Budget Limits via the Management API
Define Spending Caps and Warning Thresholds
Create budgets programmatically by posting to the budget management endpoint defined in litellm/proxy/management_endpoints/budget_management_endpoints.py:
curl -X POST https://<proxy-host>/budget \
-H "Authorization: Bearer <admin-key>" \
-H "Content-Type: application/json" \
-d '{
"budget_name": "team_alpha_budget",
"budget_type": "team",
"identifier": "team_alpha",
"max_budget": 500.0,
"soft_budget": 400.0,
"duration": "monthly",
"alert_channels": ["slack", "email"]
}'
Valid budget_type values include team, organization, user, tag, and model. The max_budget acts as a hard ceiling that blocks requests, while soft_budget triggers alerts without interrupting service.
Set Alert Channels and Notification Preferences
The alert_channels array maps to dispatchers in budget_alert_types.py. Supported channels include Slack webhooks, SMTP email, and custom webhooks. To add custom notification services, extend the ALERT_DISPATCHERS dictionary in the budget_alert_types.py module.
Enforce Budget Limits Programmatically
Use the Python SDK to set per-user budgets and per-session caps for granular control over individual request costs:
import litellm
# Create a per-user monthly budget (soft $20, hard $30)
litellm.set_budget(
budget_type="user",
identifier="user_12345",
max_budget=30.0,
soft_budget=20.0,
duration="monthly",
alert_channels=["slack"]
)
# Enforce a per-session limit of $5
litellm.set_session_budget(
user_id="user_12345",
session_max_budget=5.0
)
# Requests exceeding the session budget are rejected automatically
response = litellm.completion(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Explain quantum entanglement"}]
)
The session-level enforcement logic resides in litellm/proxy/hooks/max_budget_per_session_limiter.py, which checks the running total for each conversation thread before allowing the request to proceed.
Configure Global Budget Defaults
Set organization-wide fallback values using environment variables that budget_manager.py reads at startup:
| Variable | Purpose |
|---|---|
MAX_GLOBAL_BUDGET |
Hard ceiling applied when no explicit budget is defined for a scope. |
SOFT_GLOBAL_BUDGET |
Warning threshold used for implicit budgets. |
BUDGET_RESET_CRON |
Cron expression defining when usage counters reset (e.g., daily). |
These defaults ensure that unconfigured teams or users still operate under spending guardrails.
Manage Budgets in the LiteLLM Dashboard
The React-based administrative interface provides visual controls for non-technical stakeholders:
- Budget Panel (
ui/litellm-dashboard/src/components/budgets/budget_panel.tsx) – Displays current spend, remaining allowance, and alert status for all configured budgets. - Edit Budget Modal (
budget_modal.tsxandedit_budget_modal.tsx) – Allows admins to adjustmax_budget,soft_budget, and alert channels through a web form. - Alert History – Tracks dispatched notifications and their delivery status.
Key Implementation Files
Understanding the source architecture helps when debugging limit violations or extending functionality:
| File | Role |
|---|---|
litellm/budget_manager.py |
Central persistence and spend aggregation |
litellm/router_strategy/budget_limiter.py |
Core request-level checks and throttling logic |
litellm/proxy/hooks/max_budget_limiter.py |
Global hard-budget enforcement |
litellm/proxy/hooks/max_budget_per_session_limiter.py |
Session-scoped spending caps |
litellm/proxy/hooks/model_max_budget_limiter.py |
Model-specific budget overrides |
litellm/integrations/SlackAlerting/budget_alert_types.py |
Alert channel dispatch mapping |
litellm/proxy/management_endpoints/budget_management_endpoints.py |
REST API for CRUD operations |
Summary
- Configure budget limits and spend alerts in LiteLLM using either the REST API (
/budgetendpoint), Python SDK (set_budget()), or React dashboard. - The system distinguishes between hard limits (
max_budget) that return HTTP 429 errors and soft limits (soft_budget) that trigger Slack/email notifications. - Enforcement occurs through specialized hooks including
max_budget_limiter.pyandmax_budget_per_session_limiter.py. - Global defaults via environment variables provide fallback protection for unconfigured scopes.
Frequently Asked Questions
What happens when a LiteLLM budget limit is exceeded?
When a request would exceed the max_budget, the proxy returns an HTTP 429 error and blocks the LLM call. If the request crosses only the soft_budget but remains under the hard limit, it proceeds while the system queues an alert notification to the configured channels.
Can I set temporary budget increases for specific teams?
Yes. According to the temporary_budget_increase.md documentation, administrators can issue temporary budget overrides for specific time windows. This is useful for handling traffic spikes or special projects without permanently increasing baseline spending caps.
Which alert channels does LiteLLM support for budget notifications?
LiteLLM supports Slack webhooks, SMTP email, and custom webhooks out of the box. The mapping between channel names and sender implementations lives in budget_alert_types.py, where you can extend the ALERT_DISPATCHERS dictionary to integrate with PagerDuty, Microsoft Teams, or internal incident management systems.
How do per-session budgets differ from per-user budgets?
Per-user budgets track cumulative spending across all sessions for that identifier over a defined duration (daily, weekly, monthly). Per-session budgets, enforced by max_budget_per_session_limiter.py, apply only to a single conversation thread or request sequence, preventing individual long-running chats from consuming the entire monthly allocation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →