LiteLLM Embedding, Image Generation, and Audio Endpoints: Architecture and Usage Differences
LiteLLM treats embeddings, image generation, and audio transcription as three distinct internal call-type families with separate provider configurations, request handlers, and response transformers, despite exposing them through a unified OpenAI-compatible API.
LiteLLM (BerriAI/litellm) abstracts dozens of LLM providers behind a single Python interface and HTTP proxy. While the public endpoints for LiteLLM embedding, image generation, and audio operations look similar to OpenAI's specification, the library internally wires each modality through specialized pipelines that handle provider-specific request shapes, authentication patterns, and response formats.
Routing and Call-Type Dispatch
All incoming HTTP requests hit the FastAPI router defined in litellm/proxy/proxy_server.py. Each endpoint registers a specific CallTypes enumeration in litellm/types/utils.py that identifies the modality:
- Embeddings:
/v1/embeddingsmaps toCallTypes.embeddingorCallTypes.aembedding - Image Generation:
/v1/images/generationsmaps toCallTypes.image_generationorCallTypes.aimage_generation - Audio Transcription:
/v1/audio/transcriptionsmaps toCallTypes.transcriptionorCallTypes.atranscription
During request processing, ProxyBaseLLMRequestProcessing.base_process_llm_request (located in litellm/proxy/route_llm_request.py) inspects the call type and delegates to the appropriate handler. This dispatch mechanism ensures that an embedding request never traverses the image generation code path, even when both target the same underlying provider.
Provider-Specific Configuration
The heavy lifting for provider resolution lives in litellm/utils.py via the ProviderConfigManager class. Each modality uses a distinct lookup method:
- Embeddings:
ProviderConfigManager.get_provider_embedding_config(model, provider)returns a subclass ofBaseEmbeddingConfig(e.g.,CohereEmbeddingConfig,VoyageEmbeddingConfig) that defines the exact request shape, required headers, and cost metadata. - Image Generation:
get_provider_image_generation_configresolves configurations for models like OpenAI DALL-E, Stability, or Gemini, injecting model-specific fields such assize,quality, andstyle. - Audio Transcription:
get_provider_audio_transcription_configmaps to transformers like OpenAI Whisper, Mistral Voxtral, or IBM WatsonX, handling multipart form-data requirements.
These lookups are O(1) dictionary accesses after the lazy initialization of _PROVIDER_CONFIG_MAP, ensuring the router remains performant even with dozens of providers configured.
Request Normalization
Before issuing a provider call, LiteLLM normalizes the request payload according to the modality:
Embeddings handle token-array inputs for routers that support them (e.g., litellm.open_ai_embedding_models). If a provider does not accept token arrays, the tokens are decoded back to text using litellm.decode.
Image Generation merges client-supplied image_generation_optional_params with provider-specific defaults. It also resolves model-specific endpoint URLs, such as Azure's .../images/generations:submit path.
Audio operations parse multipart form-data, validate max_file_size_mb (a premium-only feature), and map optional fields like language and prompt into the provider's native request format.
Response Transformation
After the provider returns raw HTTP payloads, LiteLLM runs modality-specific transformers to ensure OpenAI-compatible JSON output:
- Embedding responses return vectors under
data[i].embeddingwith usage metadata includingprompt_tokensandtotal_tokens. - Image Generation wraps URLs inside an
ImageResponseobject (type: "image_generation_call"), exposingcreated,data[0].url, and cost fields. - Audio returns a
TranscriptionResponsecontainingtext,language, and optionaldurationfields.
These transformations are defined in provider-specific modules such as litellm/llms/openai/transcriptions/whisper_transformation.py and litellm/images/main.py.
Practical Implementation Examples
Synchronous Python API
All three modalities share the same top-level interface while using distinct underlying handlers:
import litellm
# Embeddings
emb = litellm.embedding(
model="text-embedding-3-large",
input="The quick brown fox jumps over the lazy dog",
)
# Image Generation
img = litellm.image_generation(
model="dall-e-3",
prompt="A futuristic cityscape at sunset, painted in oil",
size="1024x1024",
)
# Audio Transcription
transcript = litellm.audio_transcriptions(
model="whisper-1",
file=open("speech.wav", "rb"),
language="en",
)
Async Operations
Use the async variants (aembedding, aimage_generation, atranscription) to leverage the same routing, caching, and budget-tracking mechanisms without blocking:
import litellm
import asyncio
async def demo():
# Parallel execution across three different modalities
emb, img, audio = await asyncio.gather(
litellm.aembedding(
model="text-embedding-3-large",
input=["Sentence one", "Sentence two"]
),
litellm.aimage_generation(
model="dall-e-3",
prompt="A cyberpunk street market at night",
size="1024x1024"
),
litellm.atranscription(
model="whisper-1",
file=open("speech.wav", "rb"),
language="en"
)
)
return emb, img, audio
asyncio.run(demo())
Proxy Server Usage
When running the LiteLLM proxy (litellm --port 4000), each endpoint accepts standard OpenAI-formatted requests:
import httpx
import json
# Embeddings via proxy
resp = httpx.post(
"http://localhost:4000/v1/embeddings",
json={"model": "text-embedding-3-large", "input": "Hello world"},
headers={"Authorization": "Bearer sk-proxy-key"}
)
embedding = json.loads(resp.text)["data"][0]["embedding"]
# Image generation via proxy
resp = httpx.post(
"http://localhost:4000/v1/images/generations",
json={"model": "dall-e-3", "prompt": "A cat wearing a tuxedo", "size": "1024x1024"},
headers={"Authorization": "Bearer sk-proxy-key"}
)
image_url = json.loads(resp.text)["data"][0]["url"]
# Audio transcription via proxy
with open("speech.wav", "rb") as f:
files = {"file": ("speech.wav", f, "audio/wav")}
resp = httpx.post(
"http://localhost:4000/v1/audio/transcriptions",
data={"model": "whisper-1"},
files=files,
headers={"Authorization": "Bearer sk-proxy-key"}
)
text = json.loads(resp.text)["text"]
Summary
- LiteLLM embedding, image generation, and audio endpoints are routed through distinct
CallTypesenumerations inlitellm/types/utils.py. - Each modality uses dedicated ProviderConfigManager methods in
litellm/utils.pyto resolve provider-specific request shapes and authentication. - Request normalization differs by type: embeddings handle token arrays, images merge generation parameters, and audio processes multipart form-data.
- Response transformers in modality-specific modules ensure all providers return OpenAI-compatible JSON structures.
- The unified Python API (
litellm.embedding,litellm.image_generation,litellm.audio_transcriptions) abstracts these differences while maintaining provider-specific optimizations under the hood.
Frequently Asked Questions
How does LiteLLM route embedding requests differently from image generation requests?
LiteLLM routes requests through the ProxyBaseLLMRequestProcessing.base_process_llm_request method in litellm/proxy/route_llm_request.py, which inspects the CallTypes enum to determine the modality. Embeddings trigger ProviderConfigManager.get_provider_embedding_config, while image generation calls invoke get_provider_image_generation_config, ensuring each request type loads the correct transformation logic and endpoint URLs.
Can I use the same provider client for embeddings and audio transcription in LiteLLM?
No. While the Python API surface is unified, LiteLLM instantiates separate provider configurations for each modality. For example, OpenAI embedding requests utilize litellm/llms/openai_like/embedding/handler.py, whereas audio transcription requests use litellm/llms/openai/transcriptions/whisper_transformation.py (or equivalent provider-specific modules), because each modality requires distinct request serialization and response parsing.
What file size limits apply to LiteLLM audio endpoints compared to text embeddings?
Audio endpoints process multipart form-data and enforce max_file_size_mb validation (a premium feature) within the proxy_server.audio_transcriptions handler. Embeddings accept only text or token arrays and have no file upload constraints, operating entirely within JSON request bodies handled by the embedding provider configs.
Does LiteLLM cache embedding and image generation responses the same way?
Yes. All three modalities share the same caching, rate-limiting, and budget-tracking infrastructure defined in the base request processing layer. However, cache keys are computed differently for each CallTypes enum to prevent collisions between an embedding vector and an image URL, ensuring modality-specific storage and retrieval logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →