LiteLLM Python SDK vs Proxy Server Deployment: Architecture and Use Cases

The LiteLLM Python SDK runs provider calls in-process via litellm.main.completion(), while the Proxy Server deploys as a standalone FastAPI service in litellm/proxy/proxy_server.py that adds centralized auth, rate limiting, and OpenAI-compatible HTTP endpoints for multi-language clients.

The BerriAI/litellm repository provides a unified interface for calling 100+ LLM providers, offering two distinct consumption models. Understanding the architectural differences between the LiteLLM Python SDK and the LiteLLM Proxy Server helps teams choose the right deployment strategy for their latency, security, and scaling requirements.

Core Architectural Differences

The fundamental distinction lies in where the routing logic executes and how clients interact with the provider abstraction layer.

Python SDK Execution Model

The Python SDK operates through direct in-process calls. Your application imports the litellm package and invokes provider APIs within the same Python runtime, with no intermediary network hop.

The primary entry point is litellm.main.completion(), defined in [litellm/main.py](https://github.com/BerriAI/litellm/blob/main/litellm/main.py). When you call litellm.completion(), the SDK forwards arguments to Router().completion() in [litellm/router.py](https://github.com/BerriAI/litellm/blob/main/litellm/router.py), which resolves the model name to a provider handler and executes the HTTP request using httpx or aiohttp.

Proxy Server Execution Model

The Proxy Server is a separate FastAPI service that exposes OpenAI-compatible HTTP endpoints. Clients send standard HTTP requests to routes defined in [litellm/proxy/proxy_server.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py) (e.g., /v1/chat/completions), making it language-agnostic.

Under the hood, the proxy runs the same router logic inside each request handler. The completion route (around line 6938) builds a ProxyBaseLLMRequestProcessing object and calls base_process_llm_request, which invokes Router().completion—exactly the same code path as the SDK. However, the proxy adds authentication, logging, and post-processing hooks before returning the response via StreamingResponse for Server-Sent Events.

Routing and Provider Selection

Both approaches share the core routing engine, but differ in configuration persistence and request isolation.

  • SDK: Per-call routing is performed by litellm.router.Router. The SDK reads your model argument, looks it up in the in-memory config via ProviderConfigManager, and calls the appropriate provider handler (e.g., litellm.llms.openai.OpenAIChatCompletion). Each Python process maintains its own router instance.

  • Proxy: The router is instantiated once and shared across all incoming requests. This provides a single point of control for model-to-deployment mapping, retries, and fallbacks across your entire application fleet.

Authentication, Rate Limiting, and Governance

Security and access control represent the most significant functional divergence between the two deployment modes.

Python SDK does not include built-in authentication or rate limiting. You are responsible for protecting API keys before calling the SDK, and any quota management must be implemented in your application code.

Proxy Server provides enterprise-grade governance features:

All enforcement occurs centrally in the proxy, allowing you to secure provider API keys behind a single controlled endpoint.

Caching, Observability, and Production Features

Operational visibility and performance optimizations differ significantly between the two models.

SDK Caching and Logging:

  • Optional per-process caching via litellm.enable_cache()
  • Local logging isolated to the Python process
  • No built-in metrics aggregation

Proxy Observability Stack:

Code Examples: SDK vs Proxy

Using the Python SDK (In-Process)

Install the package and call providers directly without running a server:

import litellm

# Simple completion

resp = litellm.completion(
    model="gpt-3.5-turbo-instruct",
    prompt="Write a haiku about the sunrise.",
    max_tokens=30,
)

print(resp.choices[0].text)

# Chat completion with streaming

for chunk in litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain quantum tunnelling"}],
    stream=True,
):
    print(chunk.choices[0].delta.get("content", ""), end="", flush=True)

The SDK imports the same router that the proxy uses, ensuring all provider-specific parameters behave identically.

Using the Proxy Server (HTTP API)

Start the proxy with litellm --port 4000, then send standard HTTP requests from any language:

import requests
import json

url = "http://localhost:4000/v1/chat/completions"
headers = {
    "Authorization": "Bearer sk-proxy-master-key",
    "Content-Type": "application/json",
}
payload = {
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "What is the capital of France?"}],
    "max_tokens": 10,
}

r = requests.post(url, headers=headers, data=json.dumps(payload))
print(r.json()["choices"][0]["message"]["content"])

Streaming via Server-Sent Events:

import sseclient
import json
import requests

url = "http://localhost:4000/v1/chat/completions"
headers = {"Authorization": "Bearer sk-proxy-master-key"}
payload = {
    "model": "gpt-4o-mini",
    "messages": [{"role": "user", "content": "Tell me a joke"}],
    "stream": True,
}

resp = requests.post(url, headers=headers, json=payload, stream=True)
client = sseclient.SSEClient(resp)

for event in client:
    data = json.loads(event.data)
    delta = data["choices"][0]["delta"].get("content", "")
    print(delta, end="", flush=True)

All routing, retries, and provider-specific transformations occur inside the proxy using the same engine as the SDK, but with centralized authentication and logging hooks.

When to Use Each Approach

Choose the Python SDK for:

  • Quick prototyping and single-process workloads
  • Jupyter notebooks and data science pipelines
  • Low-traffic services where you already control the environment and API keys
  • Latency-sensitive applications where an extra network hop is unacceptable

Choose the Proxy Server for:

  • Multi-tenant SaaS applications requiring centralized governance
  • Multi-language client ecosystems (JavaScript, Go, Java, curl)
  • Production environments requiring audit logs, budget controls, and team-based access management
  • Scenarios requiring global caching or centralized observability (Prometheus, OpenTelemetry)

Summary

Frequently Asked Questions

Can I use the LiteLLM Proxy with non-Python languages?

Yes. The proxy exposes standard OpenAI-compatible HTTP endpoints (/v1/chat/completions, /v1/completions), allowing any language that can send HTTP requests—including JavaScript, Go, Java, or curl—to access unified LLM routing. The proxy handles provider-specific translations internally.

Does the Proxy Server use the same code as the Python SDK for provider calls?

Yes. According to the source code in [litellm/proxy/proxy_server.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py), the proxy invokes Router().completion from **[litellm/router.py](https://github.com/BerriAI/litellm/blob/main/litellm/router.py)**—the exact same class used by the SDK. The proxy is essentially a thin HTTP wrapper around the SDK's core engine, adding authentication and logging layers.

How do I handle authentication in the LiteLLM Python SDK?

The Python SDK does not provide built-in authentication. You must manage API keys and access control within your application before calling litellm.completion(). For centralized authentication, team quotas, and audit logging, deploy the Proxy Server, which validates requests via the UserAPIKeyAuth dependency in [litellm/proxy/auth/dependencies.py](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/auth/dependencies.py).

Can I stream responses when using the LiteLLM Proxy Server?

Yes. The proxy supports streaming via Server-Sent Events (SSE). When you set stream: true in your request payload, the proxy returns a StreamingResponse that yields data in OpenAI-compatible SSE format. Any HTTP client that can consume text/event-stream content—including Python's sseclient or JavaScript's EventSource—can process these streams.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →