LiteLLM Python SDK vs Proxy Server Deployment: Architecture and Use Cases
The LiteLLM Python SDK runs provider calls in-process via litellm.main.completion(), while the Proxy Server deploys as a standalone FastAPI service in litellm/proxy/proxy_server.py that adds centralized auth, rate limiting, and OpenAI-compatible HTTP endpoints for multi-language clients.
The BerriAI/litellm repository provides a unified interface for calling 100+ LLM providers, offering two distinct consumption models. Understanding the architectural differences between the LiteLLM Python SDK and the LiteLLM Proxy Server helps teams choose the right deployment strategy for their latency, security, and scaling requirements.
Core Architectural Differences
The fundamental distinction lies in where the routing logic executes and how clients interact with the provider abstraction layer.
Python SDK Execution Model
The Python SDK operates through direct in-process calls. Your application imports the litellm package and invokes provider APIs within the same Python runtime, with no intermediary network hop.
The primary entry point is litellm.main.completion(), defined in [litellm/main.py](https://github.com/BerriAI/litellm/blob/main/litellm/main.py). When you call litellm.completion(), the SDK forwards arguments to Router().completion() in [litellm/router.py](https://github.com/BerriAI/litellm/blob/main/litellm/router.py), which resolves the model name to a provider handler and executes the HTTP request using httpx or aiohttp.
Proxy Server Execution Model
The Proxy Server is a separate FastAPI service that exposes OpenAI-compatible HTTP endpoints. Clients send standard HTTP requests to routes defined in [litellm/proxy/proxy_server.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py) (e.g., /v1/chat/completions), making it language-agnostic.
Under the hood, the proxy runs the same router logic inside each request handler. The completion route (around line 6938) builds a ProxyBaseLLMRequestProcessing object and calls base_process_llm_request, which invokes Router().completion—exactly the same code path as the SDK. However, the proxy adds authentication, logging, and post-processing hooks before returning the response via StreamingResponse for Server-Sent Events.
Routing and Provider Selection
Both approaches share the core routing engine, but differ in configuration persistence and request isolation.
-
SDK: Per-call routing is performed by
litellm.router.Router. The SDK reads yourmodelargument, looks it up in the in-memory config viaProviderConfigManager, and calls the appropriate provider handler (e.g.,litellm.llms.openai.OpenAIChatCompletion). Each Python process maintains its own router instance. -
Proxy: The router is instantiated once and shared across all incoming requests. This provides a single point of control for model-to-deployment mapping, retries, and fallbacks across your entire application fleet.
Authentication, Rate Limiting, and Governance
Security and access control represent the most significant functional divergence between the two deployment modes.
Python SDK does not include built-in authentication or rate limiting. You are responsible for protecting API keys before calling the SDK, and any quota management must be implemented in your application code.
Proxy Server provides enterprise-grade governance features:
- API-key authentication via
UserAPIKeyAuthdependency injection (see [litellm/proxy/auth/dependencies.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/dependencies.py)) - Team and organization-level quotas
- Per-model rate limits and budget controls
- Optional JWT/SSO integration
All enforcement occurs centrally in the proxy, allowing you to secure provider API keys behind a single controlled endpoint.
Caching, Observability, and Production Features
Operational visibility and performance optimizations differ significantly between the two models.
SDK Caching and Logging:
- Optional per-process caching via
litellm.enable_cache() - Local logging isolated to the Python process
- No built-in metrics aggregation
Proxy Observability Stack:
- Global HTTP caching with
Cache-Controlsupport - Prometheus metrics and OpenTelemetry tracing (see [
litellm/proxy/metrics.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/metrics.py)) - Centralized log aggregation and usage tracking
- Optional guardrails in
litellm/proxy/guardrails/that can intercept and modify responses before they leave the server
Code Examples: SDK vs Proxy
Using the Python SDK (In-Process)
Install the package and call providers directly without running a server:
import litellm
# Simple completion
resp = litellm.completion(
model="gpt-3.5-turbo-instruct",
prompt="Write a haiku about the sunrise.",
max_tokens=30,
)
print(resp.choices[0].text)
# Chat completion with streaming
for chunk in litellm.completion(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
stream=True,
):
print(chunk.choices[0].delta.get("content", ""), end="", flush=True)
The SDK imports the same router that the proxy uses, ensuring all provider-specific parameters behave identically.
Using the Proxy Server (HTTP API)
Start the proxy with litellm --port 4000, then send standard HTTP requests from any language:
import requests
import json
url = "http://localhost:4000/v1/chat/completions"
headers = {
"Authorization": "Bearer sk-proxy-master-key",
"Content-Type": "application/json",
}
payload = {
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "What is the capital of France?"}],
"max_tokens": 10,
}
r = requests.post(url, headers=headers, data=json.dumps(payload))
print(r.json()["choices"][0]["message"]["content"])
Streaming via Server-Sent Events:
import sseclient
import json
import requests
url = "http://localhost:4000/v1/chat/completions"
headers = {"Authorization": "Bearer sk-proxy-master-key"}
payload = {
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Tell me a joke"}],
"stream": True,
}
resp = requests.post(url, headers=headers, json=payload, stream=True)
client = sseclient.SSEClient(resp)
for event in client:
data = json.loads(event.data)
delta = data["choices"][0]["delta"].get("content", "")
print(delta, end="", flush=True)
All routing, retries, and provider-specific transformations occur inside the proxy using the same engine as the SDK, but with centralized authentication and logging hooks.
When to Use Each Approach
Choose the Python SDK for:
- Quick prototyping and single-process workloads
- Jupyter notebooks and data science pipelines
- Low-traffic services where you already control the environment and API keys
- Latency-sensitive applications where an extra network hop is unacceptable
Choose the Proxy Server for:
- Multi-tenant SaaS applications requiring centralized governance
- Multi-language client ecosystems (JavaScript, Go, Java, curl)
- Production environments requiring audit logs, budget controls, and team-based access management
- Scenarios requiring global caching or centralized observability (Prometheus, OpenTelemetry)
Summary
- The LiteLLM Python SDK provides in-process access via
litellm.main.completion()in [litellm/main.py](https://github.com/BerriAI/litellm/blob/main/litellm/main.py), using the shared [litellm/router.py](https://github.com/BerriAI/litellm/blob/main/litellm/router.py) engine for provider abstraction. - The LiteLLM Proxy Server wraps the same routing logic in a FastAPI application ([
litellm/proxy/proxy_server.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py)), adding HTTP endpoints, authentication (UserAPIKeyAuth), rate limiting, and observability hooks. - Both use identical provider handlers and response formats, but the proxy adds enterprise security and multi-language support at the cost of a network hop.
- Use the SDK for lightweight, low-latency Python applications; deploy the proxy for centralized governance and production-scale multi-tenant workloads.
Frequently Asked Questions
Can I use the LiteLLM Proxy with non-Python languages?
Yes. The proxy exposes standard OpenAI-compatible HTTP endpoints (/v1/chat/completions, /v1/completions), allowing any language that can send HTTP requests—including JavaScript, Go, Java, or curl—to access unified LLM routing. The proxy handles provider-specific translations internally.
Does the Proxy Server use the same code as the Python SDK for provider calls?
Yes. According to the source code in [litellm/proxy/proxy_server.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py), the proxy invokes Router().completion from **[litellm/router.py](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)**—the exact same class used by the SDK. The proxy is essentially a thin HTTP wrapper around the SDK's core engine, adding authentication and logging layers.
How do I handle authentication in the LiteLLM Python SDK?
The Python SDK does not provide built-in authentication. You must manage API keys and access control within your application before calling litellm.completion(). For centralized authentication, team quotas, and audit logging, deploy the Proxy Server, which validates requests via the UserAPIKeyAuth dependency in [litellm/proxy/auth/dependencies.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/dependencies.py).
Can I stream responses when using the LiteLLM Proxy Server?
Yes. The proxy supports streaming via Server-Sent Events (SSE). When you set stream: true in your request payload, the proxy returns a StreamingResponse that yields data in OpenAI-compatible SSE format. Any HTTP client that can consume text/event-stream content—including Python's sseclient or JavaScript's EventSource—can process these streams.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →