LiteLLM Embedding, Image Generation, and Audio Endpoints: Architecture and Usage Differences

LiteLLM treats embeddings, image generation, and audio transcription as three distinct internal call-type families with separate provider configurations, request handlers, and response transformers, despite exposing them through a unified OpenAI-compatible API.

LiteLLM (BerriAI/litellm) abstracts dozens of LLM providers behind a single Python interface and HTTP proxy. While the public endpoints for LiteLLM embedding, image generation, and audio operations look similar to OpenAI's specification, the library internally wires each modality through specialized pipelines that handle provider-specific request shapes, authentication patterns, and response formats.

Routing and Call-Type Dispatch

All incoming HTTP requests hit the FastAPI router defined in litellm/proxy/proxy_server.py. Each endpoint registers a specific CallTypes enumeration in litellm/types/utils.py that identifies the modality:

  • Embeddings: /v1/embeddings maps to CallTypes.embedding or CallTypes.aembedding
  • Image Generation: /v1/images/generations maps to CallTypes.image_generation or CallTypes.aimage_generation
  • Audio Transcription: /v1/audio/transcriptions maps to CallTypes.transcription or CallTypes.atranscription

During request processing, ProxyBaseLLMRequestProcessing.base_process_llm_request (located in litellm/proxy/route_llm_request.py) inspects the call type and delegates to the appropriate handler. This dispatch mechanism ensures that an embedding request never traverses the image generation code path, even when both target the same underlying provider.

Provider-Specific Configuration

The heavy lifting for provider resolution lives in litellm/utils.py via the ProviderConfigManager class. Each modality uses a distinct lookup method:

  • Embeddings: ProviderConfigManager.get_provider_embedding_config(model, provider) returns a subclass of BaseEmbeddingConfig (e.g., CohereEmbeddingConfig, VoyageEmbeddingConfig) that defines the exact request shape, required headers, and cost metadata.
  • Image Generation: get_provider_image_generation_config resolves configurations for models like OpenAI DALL-E, Stability, or Gemini, injecting model-specific fields such as size, quality, and style.
  • Audio Transcription: get_provider_audio_transcription_config maps to transformers like OpenAI Whisper, Mistral Voxtral, or IBM WatsonX, handling multipart form-data requirements.

These lookups are O(1) dictionary accesses after the lazy initialization of _PROVIDER_CONFIG_MAP, ensuring the router remains performant even with dozens of providers configured.

Request Normalization

Before issuing a provider call, LiteLLM normalizes the request payload according to the modality:

Embeddings handle token-array inputs for routers that support them (e.g., litellm.open_ai_embedding_models). If a provider does not accept token arrays, the tokens are decoded back to text using litellm.decode.

Image Generation merges client-supplied image_generation_optional_params with provider-specific defaults. It also resolves model-specific endpoint URLs, such as Azure's .../images/generations:submit path.

Audio operations parse multipart form-data, validate max_file_size_mb (a premium-only feature), and map optional fields like language and prompt into the provider's native request format.

Response Transformation

After the provider returns raw HTTP payloads, LiteLLM runs modality-specific transformers to ensure OpenAI-compatible JSON output:

  • Embedding responses return vectors under data[i].embedding with usage metadata including prompt_tokens and total_tokens.
  • Image Generation wraps URLs inside an ImageResponse object (type: "image_generation_call"), exposing created, data[0].url, and cost fields.
  • Audio returns a TranscriptionResponse containing text, language, and optional duration fields.

These transformations are defined in provider-specific modules such as litellm/llms/openai/transcriptions/whisper_transformation.py and litellm/images/main.py.

Practical Implementation Examples

Synchronous Python API

All three modalities share the same top-level interface while using distinct underlying handlers:

import litellm

# Embeddings

emb = litellm.embedding(
    model="text-embedding-3-large",
    input="The quick brown fox jumps over the lazy dog",
)

# Image Generation

img = litellm.image_generation(
    model="dall-e-3",
    prompt="A futuristic cityscape at sunset, painted in oil",
    size="1024x1024",
)

# Audio Transcription

transcript = litellm.audio_transcriptions(
    model="whisper-1",
    file=open("speech.wav", "rb"),
    language="en",
)

Async Operations

Use the async variants (aembedding, aimage_generation, atranscription) to leverage the same routing, caching, and budget-tracking mechanisms without blocking:

import litellm
import asyncio

async def demo():
    # Parallel execution across three different modalities

    emb, img, audio = await asyncio.gather(
        litellm.aembedding(
            model="text-embedding-3-large",
            input=["Sentence one", "Sentence two"]
        ),
        litellm.aimage_generation(
            model="dall-e-3",
            prompt="A cyberpunk street market at night",
            size="1024x1024"
        ),
        litellm.atranscription(
            model="whisper-1",
            file=open("speech.wav", "rb"),
            language="en"
        )
    )
    return emb, img, audio

asyncio.run(demo())

Proxy Server Usage

When running the LiteLLM proxy (litellm --port 4000), each endpoint accepts standard OpenAI-formatted requests:

import httpx
import json

# Embeddings via proxy

resp = httpx.post(
    "http://localhost:4000/v1/embeddings",
    json={"model": "text-embedding-3-large", "input": "Hello world"},
    headers={"Authorization": "Bearer sk-proxy-key"}
)
embedding = json.loads(resp.text)["data"][0]["embedding"]

# Image generation via proxy

resp = httpx.post(
    "http://localhost:4000/v1/images/generations",
    json={"model": "dall-e-3", "prompt": "A cat wearing a tuxedo", "size": "1024x1024"},
    headers={"Authorization": "Bearer sk-proxy-key"}
)
image_url = json.loads(resp.text)["data"][0]["url"]

# Audio transcription via proxy

with open("speech.wav", "rb") as f:
    files = {"file": ("speech.wav", f, "audio/wav")}
    resp = httpx.post(
        "http://localhost:4000/v1/audio/transcriptions",
        data={"model": "whisper-1"},
        files=files,
        headers={"Authorization": "Bearer sk-proxy-key"}
    )
text = json.loads(resp.text)["text"]

Summary

  • LiteLLM embedding, image generation, and audio endpoints are routed through distinct CallTypes enumerations in litellm/types/utils.py.
  • Each modality uses dedicated ProviderConfigManager methods in litellm/utils.py to resolve provider-specific request shapes and authentication.
  • Request normalization differs by type: embeddings handle token arrays, images merge generation parameters, and audio processes multipart form-data.
  • Response transformers in modality-specific modules ensure all providers return OpenAI-compatible JSON structures.
  • The unified Python API (litellm.embedding, litellm.image_generation, litellm.audio_transcriptions) abstracts these differences while maintaining provider-specific optimizations under the hood.

Frequently Asked Questions

How does LiteLLM route embedding requests differently from image generation requests?

LiteLLM routes requests through the ProxyBaseLLMRequestProcessing.base_process_llm_request method in litellm/proxy/route_llm_request.py, which inspects the CallTypes enum to determine the modality. Embeddings trigger ProviderConfigManager.get_provider_embedding_config, while image generation calls invoke get_provider_image_generation_config, ensuring each request type loads the correct transformation logic and endpoint URLs.

Can I use the same provider client for embeddings and audio transcription in LiteLLM?

No. While the Python API surface is unified, LiteLLM instantiates separate provider configurations for each modality. For example, OpenAI embedding requests utilize litellm/llms/openai_like/embedding/handler.py, whereas audio transcription requests use litellm/llms/openai/transcriptions/whisper_transformation.py (or equivalent provider-specific modules), because each modality requires distinct request serialization and response parsing.

What file size limits apply to LiteLLM audio endpoints compared to text embeddings?

Audio endpoints process multipart form-data and enforce max_file_size_mb validation (a premium feature) within the proxy_server.audio_transcriptions handler. Embeddings accept only text or token arrays and have no file upload constraints, operating entirely within JSON request bodies handled by the embedding provider configs.

Does LiteLLM cache embedding and image generation responses the same way?

Yes. All three modalities share the same caching, rate-limiting, and budget-tracking infrastructure defined in the base request processing layer. However, cache keys are computed differently for each CallTypes enum to prevent collisions between an embedding vector and an image URL, ensuring modality-specific storage and retrieval logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →