# litellm | Berri AI | Knowledge Base | Instagit

Python SDK, Proxy Server (AI Gateway) to call 100+ LLM APIs in OpenAI (or native) format, with cost tracking, guardrails, loadbalancing and logging. [Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic, Sagemaker, HuggingFace, VLLM, NVIDIA NIM]

GitHub Stars: 40.8k

Repository: https://github.com/BerriAI/litellm

---

## Articles

### [How to Migrate from OpenAI/Anthropic SDKs to LiteLLM: A Complete Guide](/BerriAI/litellm/migrate-openai-anthropic-sdks-to-litellm)

Effortlessly migrate from OpenAI or Anthropic SDKs to LiteLLM. Achieve zero code changes with the LiteLLM proxy or easily swap SDK calls for LiteLLM's completion function, saving time and resources.

- Tags: migration-guide
- Published: 2026-03-26

### [LiteLLM Proxy Health Check and Monitoring Endpoints: A Complete Implementation Guide](/BerriAI/litellm/litellm-proxy-health-check-monitoring)

Implement LiteLLM proxy health check and monitoring endpoints. Validate proxy availability, LLM deployments, and integrations like Datadog and Slack. Get the complete guide.

- Tags: how-to-guide
- Published: 2026-03-26

### [LiteLLM Model-Specific Parameters and Provider Transformations: A Complete Guide](/BerriAI/litellm/litellm-model-specific-parameters-provider-transformations)

Master LiteLLM model-specific parameters & provider transformations. This guide explains how LiteLLM unifies LLM providers with a two-layer architecture for seamless API integration.

- Tags: deep-dive
- Published: 2026-03-26

### [LiteLLM Embedding, Image Generation, and Audio Endpoints: Architecture and Usage Differences](/BerriAI/litellm/litellm-embedding-image-audio-endpoints)

Explore LiteLLM embedding, image generation, and audio endpoint differences. Understand how LiteLLM unifies these distinct call types within its architecture for seamless API integration.

- Tags: architecture
- Published: 2026-03-26

### [High Availability and Load Balancing LiteLLM Proxy: A Production Deployment Guide](/BerriAI/litellm/high-availability-load-balancing-litellm-proxy)

Achieve high availability and load balancing for your LiteLLM proxy. Learn how to deploy stateless, horizontally scaled, and automatically load-balanced model endpoints for production readiness.

- Tags: how-to-guide
- Published: 2026-03-26

### [Production Security Best Practices for LiteLLM Proxy: 10 Critical Measures](/BerriAI/litellm/production-security-best-practices-litellm-proxy)

Implement production security best practices for your LiteLLM proxy. Learn 10 critical measures to secure your API with env vars, encryption, and read-only filesystems.

- Tags: best-practices
- Published: 2026-03-26

### [LiteLLM Model Context Protocol (MCP) Servers and Tools: A Complete Technical Guide](/BerriAI/litellm/litellm-mcp-servers-tools)

Master LiteLLM Model Context Protocol (MCP) servers and tools. This guide explains how MCP enables OpenAI compatible clients to discover and execute external tools via a unified gateway.

- Tags: deep-dive
- Published: 2026-03-26

### [LiteLLM Proxy Custom Prompt Management: A Complete Guide to Dynamic Prompt Templates](/BerriAI/litellm/litellm-proxy-custom-prompt-management)

Master LiteLLM proxy custom prompt management. Inject, version, and retrieve dynamic prompt templates at runtime via API. Enhance your LLM applications effortlessly.

- Tags: how-to-guide
- Published: 2026-03-26

### [How to Configure Budget Limits and Spend Alerts in LiteLLM](/BerriAI/litellm/configure-budget-limits-spend-alerts-litellm)

Configure budget limits and spend alerts in LiteLLM with its three-layer system. Set caps, get Slack/email alerts, and enforce limits per model, team, or user. Optimize your LLM spending effectively.

- Tags: how-to-guide
- Published: 2026-03-26

### [LiteLLM Token Usage and Cost Tracking Across Providers: A Complete Technical Guide](/BerriAI/litellm/litellm-token-usage-cost-tracking)

Master LiteLLM token usage and cost tracking across providers. Our guide shows how LiteLLM normalizes counts and calculates costs for seamless LLM management.

- Tags: deep-dive
- Published: 2026-03-26

### [How to Invoke A2A Agents with LiteLLM: A2A Protocol Implementation Guide](/BerriAI/litellm/invoke-a2a-agents-litellm)

Learn to invoke A2A agents using LiteLLM's A2A protocol support. This guide details direct agent communication and LLM-as-agent patterns through a unified proxy.

- Tags: how-to-guide
- Published: 2026-03-26

### [How LiteLLM Virtual Keys Work for API Key Management: A Complete Technical Guide](/BerriAI/litellm/litellm-virtual-keys-api-key-management)

Learn how LiteLLM virtual keys manage API keys, authenticate requests, enforce limits, and track usage for LLM providers. A complete technical guide.

- Tags: how-to-guide
- Published: 2026-03-26

### [LiteLLM Observability Integrations: Complete Guide to Langfuse, LangSmith, DataDog, and MLflow](/BerriAI/litellm/litellm-observability-integrations)

Explore LiteLLM observability integrations with Langfuse, LangSmith, DataDog, and MLflow. Capture LLM requests and send telemetry via a unified pipeline. Learn more!

- Tags: how-to-guide
- Published: 2026-03-26

### [LiteLLM Tool Calling and Function Calling: Universal Capability Detection Across Providers](/BerriAI/litellm/litellm-tool-function-calling-support)

Explore LiteLLM tool calling and function calling. Detect capability across providers automatically, preventing API errors with unified support.

- Tags: deep-dive
- Published: 2026-03-26

### [How to Add a Custom LLM Provider to LiteLLM: Complete Implementation Guide](/BerriAI/litellm/add-custom-llm-provider-litellm)

Easily add a custom LLM provider to LiteLLM by subclassing CustomLLM and registering your provider. This guide shows you how to seamlessly integrate any proprietary or third-party language model.

- Tags: how-to-guide
- Published: 2026-03-26

### [LiteLLM Python SDK vs Proxy Server Deployment: Architecture and Use Cases](/BerriAI/litellm/litellm-python-sdk-vs-proxy)

Compare LiteLLM Python SDK vs proxy server deployment. Understand in-process calls and standalone FastAPI services for centralized auth, rate limiting, and multi-language clients.

- Tags: architecture
- Published: 2026-03-26

### [LiteLLM Streaming Responses Across LLM Providers: Architecture and Implementation](/BerriAI/litellm/litellm-streaming-responses-providers)

Master LiteLLM streaming responses from OpenAI, Anthropic, Gemini & more. Discover the unified, provider-agnostic interface and three-layer architecture for normalized real-time LLM data.

- Tags: architecture
- Published: 2026-03-26

### [How to Configure LiteLLM Caching: Redis, In-Memory, and Disk Backends Explained](/BerriAI/litellm/configure-litellm-caching)

Learn to configure LiteLLM caching with Redis, in-memory, and disk backends. Reduce LLM costs and latency using TTL and cache policies.

- Tags: how-to-guide
- Published: 2026-03-26

### [LiteLLM Caching Layer Architecture: Redis vs In-Memory Implementation Guide](/BerriAI/litellm/litellm-caching-architecture)

Explore LiteLLM's caching layer architecture. Compare Redis vs in-memory implementations and understand how to integrate them seamlessly using the BaseCache interface for optimal performance.

- Tags: architecture
- Published: 2026-03-26

### [How to Configure Custom Authentication for the LiteLLM Proxy](/BerriAI/litellm/configure-custom-authentication-litellm-proxy)

Learn to configure custom authentication for LiteLLM proxy using Python functions. Implement bespoke security policies easily without altering core code.

- Tags: how-to-guide
- Published: 2026-03-26

### [How LiteLLM Handles Rate Limiting and Cooldown Periods: A Complete Technical Guide](/BerriAI/litellm/litellm-rate-limiting-cooldown-handling)

Explore how LiteLLM manages rate limiting and cooldown periods with TPM/RPM limits, priority reservations, and automatic backend rotation. Protect your inference workloads effectively.

- Tags: how-to-guide
- Published: 2026-03-26

### [LiteLLM Router Strategies Comparison: Cost, Latency, Load, and Rate Limits Explained](/BerriAI/litellm/litellm-router-strategies-comparison)

Compare LiteLLM router strategies lowest_cost, lowest_latency, least_busy, and lowest_tpm_rpm to optimize LLM deployments for cost, speed, load, and rate limits.

- Tags: performance
- Published: 2026-03-26

### [LiteLLM Router Failover and Retry Logic: A Deep Dive into Resilient LLM Routing](/BerriAI/litellm/litellm-router-failover-retry-logic)

Explore LiteLLM router failover and retry logic. Automatically reroute failed LLM requests to backup models enhancing application resilience without code changes.

- Tags: deep-dive
- Published: 2026-03-26

