# LiteLLM Python SDK vs Proxy Server Deployment: Architecture and Use Cases

> Compare LiteLLM Python SDK vs proxy server deployment. Understand in-process calls and standalone FastAPI services for centralized auth, rate limiting, and multi-language clients.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: architecture
- Published: 2026-03-26

---

**The LiteLLM Python SDK runs provider calls in-process via `litellm.main.completion()`, while the Proxy Server deploys as a standalone FastAPI service in [`litellm/proxy/proxy_server.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py) that adds centralized auth, rate limiting, and OpenAI-compatible HTTP endpoints for multi-language clients.**

The **BerriAI/litellm** repository provides a unified interface for calling 100+ LLM providers, offering two distinct consumption models. Understanding the architectural differences between the **LiteLLM Python SDK** and the **LiteLLM Proxy Server** helps teams choose the right deployment strategy for their latency, security, and scaling requirements.

## Core Architectural Differences

The fundamental distinction lies in where the routing logic executes and how clients interact with the provider abstraction layer.

### Python SDK Execution Model

The **Python SDK** operates through direct in-process calls. Your application imports the `litellm` package and invokes provider APIs within the same Python runtime, with no intermediary network hop.

The primary entry point is **`litellm.main.completion()`**, defined in **[[`litellm/main.py`](https://github.com/BerriAI/litellm/blob/main/litellm/main.py)](https://github.com/BerriAI/litellm/blob/main/litellm/main.py)**. When you call `litellm.completion()`, the SDK forwards arguments to **`Router().completion()`** in **[[`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)**, which resolves the model name to a provider handler and executes the HTTP request using `httpx` or `aiohttp`.

### Proxy Server Execution Model

The **Proxy Server** is a separate FastAPI service that exposes OpenAI-compatible HTTP endpoints. Clients send standard HTTP requests to routes defined in **[[`litellm/proxy/proxy_server.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py)** (e.g., `/v1/chat/completions`), making it language-agnostic.

Under the hood, the proxy runs the **same router logic** inside each request handler. The `completion` route (around line 6938) builds a `ProxyBaseLLMRequestProcessing` object and calls `base_process_llm_request`, which invokes `Router().completion`—**exactly the same code path** as the SDK. However, the proxy adds authentication, logging, and post-processing hooks before returning the response via `StreamingResponse` for Server-Sent Events.

## Routing and Provider Selection

Both approaches share the core routing engine, but differ in configuration persistence and request isolation.

- **SDK**: Per-call routing is performed by **`litellm.router.Router`**. The SDK reads your `model` argument, looks it up in the in-memory config via `ProviderConfigManager`, and calls the appropriate provider handler (e.g., `litellm.llms.openai.OpenAIChatCompletion`). Each Python process maintains its own router instance.

- **Proxy**: The router is instantiated once and shared across all incoming requests. This provides a **single point of control** for model-to-deployment mapping, retries, and fallbacks across your entire application fleet.

## Authentication, Rate Limiting, and Governance

Security and access control represent the most significant functional divergence between the two deployment modes.

**Python SDK** does not include built-in authentication or rate limiting. You are responsible for protecting API keys before calling the SDK, and any quota management must be implemented in your application code.

**Proxy Server** provides enterprise-grade governance features:
- **API-key authentication** via `UserAPIKeyAuth` dependency injection (see **[[`litellm/proxy/auth/dependencies.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/dependencies.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/dependencies.py)**)
- Team and organization-level quotas
- Per-model rate limits and budget controls
- Optional JWT/SSO integration

All enforcement occurs centrally in the proxy, allowing you to secure provider API keys behind a single controlled endpoint.

## Caching, Observability, and Production Features

Operational visibility and performance optimizations differ significantly between the two models.

**SDK Caching and Logging**:
- Optional per-process caching via `litellm.enable_cache()`
- Local logging isolated to the Python process
- No built-in metrics aggregation

**Proxy Observability Stack**:
- Global HTTP caching with `Cache-Control` support
- Prometheus metrics and OpenTelemetry tracing (see **[[`litellm/proxy/metrics.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/metrics.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/metrics.py)**)
- Centralized log aggregation and usage tracking
- Optional guardrails in **[`litellm/proxy/guardrails/`](https://github.com/BerriAI/litellm/tree/main/litellm/proxy/guardrails)** that can intercept and modify responses before they leave the server

## Code Examples: SDK vs Proxy

### Using the Python SDK (In-Process)

Install the package and call providers directly without running a server:

```python
import litellm

# Simple completion

resp = litellm.completion(
    model="gpt-3.5-turbo-instruct",
    prompt="Write a haiku about the sunrise.",
    max_tokens=30,
)

print(resp.choices[0].text)

# Chat completion with streaming

for chunk in litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
    stream=True,
):
    print(chunk.choices[0].delta.get("content", ""), end="", flush=True)

```

The SDK imports the same router that the proxy uses, ensuring all provider-specific parameters behave identically.

### Using the Proxy Server (HTTP API)

Start the proxy with `litellm --port 4000`, then send standard HTTP requests from any language:

```python
import requests
import json

url = "http://localhost:4000/v1/chat/completions"
headers = {
    "Authorization": "Bearer sk-proxy-master-key",
    "Content-Type": "application/json",
}
payload = {
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "What is the capital of France?"}],
    "max_tokens": 10,
}

r = requests.post(url, headers=headers, data=json.dumps(payload))
print(r.json()["choices"][0]["message"]["content"])

```

**Streaming via Server-Sent Events:**

```python
import sseclient
import json
import requests

url = "http://localhost:4000/v1/chat/completions"
headers = {"Authorization": "Bearer sk-proxy-master-key"}
payload = {
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Tell me a joke"}],
    "stream": True,
}

resp = requests.post(url, headers=headers, json=payload, stream=True)
client = sseclient.SSEClient(resp)

for event in client:
    data = json.loads(event.data)
    delta = data["choices"][0]["delta"].get("content", "")
    print(delta, end="", flush=True)

```

All routing, retries, and provider-specific transformations occur inside the proxy using the same engine as the SDK, but with centralized authentication and logging hooks.

## When to Use Each Approach

Choose the **Python SDK** for:
- Quick prototyping and single-process workloads
- Jupyter notebooks and data science pipelines
- Low-traffic services where you already control the environment and API keys
- Latency-sensitive applications where an extra network hop is unacceptable

Choose the **Proxy Server** for:
- Multi-tenant SaaS applications requiring centralized governance
- Multi-language client ecosystems (JavaScript, Go, Java, curl)
- Production environments requiring audit logs, budget controls, and team-based access management
- Scenarios requiring global caching or centralized observability (Prometheus, OpenTelemetry)

## Summary

- The **LiteLLM Python SDK** provides in-process access via `litellm.main.completion()` in **[[`litellm/main.py`](https://github.com/BerriAI/litellm/blob/main/litellm/main.py)](https://github.com/BerriAI/litellm/blob/main/litellm/main.py)**, using the shared **[[`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)** engine for provider abstraction.
- The **LiteLLM Proxy Server** wraps the same routing logic in a FastAPI application (**[[`litellm/proxy/proxy_server.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py)**), adding HTTP endpoints, **authentication** (`UserAPIKeyAuth`), **rate limiting**, and **observability** hooks.
- Both use identical provider handlers and response formats, but the proxy adds enterprise security and multi-language support at the cost of a network hop.
- Use the SDK for lightweight, low-latency Python applications; deploy the proxy for centralized governance and production-scale multi-tenant workloads.

## Frequently Asked Questions

### Can I use the LiteLLM Proxy with non-Python languages?

Yes. The proxy exposes standard OpenAI-compatible HTTP endpoints (`/v1/chat/completions`, `/v1/completions`), allowing any language that can send HTTP requests—including JavaScript, Go, Java, or curl—to access unified LLM routing. The proxy handles provider-specific translations internally.

### Does the Proxy Server use the same code as the Python SDK for provider calls?

Yes. According to the source code in **[[`litellm/proxy/proxy_server.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py)**, the proxy invokes `Router().completion` from **[[`litellm/router.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)**—the exact same class used by the SDK. The proxy is essentially a thin HTTP wrapper around the SDK's core engine, adding authentication and logging layers.

### How do I handle authentication in the LiteLLM Python SDK?

The Python SDK does not provide built-in authentication. You must manage API keys and access control within your application before calling `litellm.completion()`. For centralized authentication, team quotas, and audit logging, deploy the **Proxy Server**, which validates requests via the `UserAPIKeyAuth` dependency in **[[`litellm/proxy/auth/dependencies.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/dependencies.py)](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/dependencies.py)**.

### Can I stream responses when using the LiteLLM Proxy Server?

Yes. The proxy supports streaming via Server-Sent Events (SSE). When you set `stream: true` in your request payload, the proxy returns a `StreamingResponse` that yields data in OpenAI-compatible SSE format. Any HTTP client that can consume `text/event-stream` content—including Python's `sseclient` or JavaScript's `EventSource`—can process these streams.