# LiteLLM Embedding, Image Generation, and Audio Endpoints: Architecture and Usage Differences

> Explore LiteLLM embedding, image generation, and audio endpoint differences. Understand how LiteLLM unifies these distinct call types within its architecture for seamless API integration.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: architecture
- Published: 2026-03-26

---

**LiteLLM treats embeddings, image generation, and audio transcription as three distinct internal call-type families with separate provider configurations, request handlers, and response transformers, despite exposing them through a unified OpenAI-compatible API.**

LiteLLM (BerriAI/litellm) abstracts dozens of LLM providers behind a single Python interface and HTTP proxy. While the public endpoints for **LiteLLM embedding**, **image generation**, and **audio** operations look similar to OpenAI's specification, the library internally wires each modality through specialized pipelines that handle provider-specific request shapes, authentication patterns, and response formats.

## Routing and Call-Type Dispatch

All incoming HTTP requests hit the **FastAPI router** defined in [`litellm/proxy/proxy_server.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/proxy_server.py). Each endpoint registers a specific **CallTypes** enumeration in [`litellm/types/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/utils.py) that identifies the modality:

- **Embeddings**: `/v1/embeddings` maps to `CallTypes.embedding` or `CallTypes.aembedding`
- **Image Generation**: `/v1/images/generations` maps to `CallTypes.image_generation` or `CallTypes.aimage_generation`
- **Audio Transcription**: `/v1/audio/transcriptions` maps to `CallTypes.transcription` or `CallTypes.atranscription`

During request processing, `ProxyBaseLLMRequestProcessing.base_process_llm_request` (located in [`litellm/proxy/route_llm_request.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/route_llm_request.py)) inspects the call type and delegates to the appropriate handler. This dispatch mechanism ensures that an embedding request never traverses the image generation code path, even when both target the same underlying provider.

## Provider-Specific Configuration

The heavy lifting for provider resolution lives in [`litellm/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/utils.py) via the **ProviderConfigManager** class. Each modality uses a distinct lookup method:

- **Embeddings**: `ProviderConfigManager.get_provider_embedding_config(model, provider)` returns a subclass of `BaseEmbeddingConfig` (e.g., `CohereEmbeddingConfig`, `VoyageEmbeddingConfig`) that defines the exact request shape, required headers, and cost metadata.
- **Image Generation**: `get_provider_image_generation_config` resolves configurations for models like OpenAI DALL-E, Stability, or Gemini, injecting model-specific fields such as `size`, `quality`, and `style`.
- **Audio Transcription**: `get_provider_audio_transcription_config` maps to transformers like OpenAI Whisper, Mistral Voxtral, or IBM WatsonX, handling multipart form-data requirements.

These lookups are **O(1)** dictionary accesses after the lazy initialization of `_PROVIDER_CONFIG_MAP`, ensuring the router remains performant even with dozens of providers configured.

## Request Normalization

Before issuing a provider call, LiteLLM normalizes the request payload according to the modality:

**Embeddings** handle token-array inputs for routers that support them (e.g., `litellm.open_ai_embedding_models`). If a provider does not accept token arrays, the tokens are decoded back to text using `litellm.decode`.

**Image Generation** merges client-supplied `image_generation_optional_params` with provider-specific defaults. It also resolves model-specific endpoint URLs, such as Azure's `.../images/generations:submit` path.

**Audio** operations parse multipart form-data, validate `max_file_size_mb` (a premium-only feature), and map optional fields like `language` and `prompt` into the provider's native request format.

## Response Transformation

After the provider returns raw HTTP payloads, LiteLLM runs modality-specific transformers to ensure OpenAI-compatible JSON output:

- **Embedding responses** return vectors under `data[i].embedding` with usage metadata including `prompt_tokens` and `total_tokens`.
- **Image Generation** wraps URLs inside an `ImageResponse` object (`type: "image_generation_call"`), exposing `created`, `data[0].url`, and cost fields.
- **Audio** returns a `TranscriptionResponse` containing `text`, `language`, and optional `duration` fields.

These transformations are defined in provider-specific modules such as [`litellm/llms/openai/transcriptions/whisper_transformation.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/openai/transcriptions/whisper_transformation.py) and [`litellm/images/main.py`](https://github.com/BerriAI/litellm/blob/main/litellm/images/main.py).

## Practical Implementation Examples

### Synchronous Python API

All three modalities share the same top-level interface while using distinct underlying handlers:

```python
import litellm

# Embeddings

emb = litellm.embedding(
    model="text-embedding-3-large",
    input="The quick brown fox jumps over the lazy dog",
)

# Image Generation

img = litellm.image_generation(
    model="dall-e-3",
    prompt="A futuristic cityscape at sunset, painted in oil",
    size="1024x1024",
)

# Audio Transcription

transcript = litellm.audio_transcriptions(
    model="whisper-1",
    file=open("speech.wav", "rb"),
    language="en",
)

```

### Async Operations

Use the async variants (`aembedding`, `aimage_generation`, `atranscription`) to leverage the same routing, caching, and budget-tracking mechanisms without blocking:

```python
import litellm
import asyncio

async def demo():
    # Parallel execution across three different modalities

    emb, img, audio = await asyncio.gather(
        litellm.aembedding(
            model="text-embedding-3-large",
            input=["Sentence one", "Sentence two"]
        ),
        litellm.aimage_generation(
            model="dall-e-3",
            prompt="A cyberpunk street market at night",
            size="1024x1024"
        ),
        litellm.atranscription(
            model="whisper-1",
            file=open("speech.wav", "rb"),
            language="en"
        )
    )
    return emb, img, audio

asyncio.run(demo())

```

### Proxy Server Usage

When running the LiteLLM proxy (`litellm --port 4000`), each endpoint accepts standard OpenAI-formatted requests:

```python
import httpx
import json

# Embeddings via proxy

resp = httpx.post(
    "http://localhost:4000/v1/embeddings",
    json={"model": "text-embedding-3-large", "input": "Hello world"},
    headers={"Authorization": "Bearer sk-proxy-key"}
)
embedding = json.loads(resp.text)["data"][0]["embedding"]

# Image generation via proxy

resp = httpx.post(
    "http://localhost:4000/v1/images/generations",
    json={"model": "dall-e-3", "prompt": "A cat wearing a tuxedo", "size": "1024x1024"},
    headers={"Authorization": "Bearer sk-proxy-key"}
)
image_url = json.loads(resp.text)["data"][0]["url"]

# Audio transcription via proxy

with open("speech.wav", "rb") as f:
    files = {"file": ("speech.wav", f, "audio/wav")}
    resp = httpx.post(
        "http://localhost:4000/v1/audio/transcriptions",
        data={"model": "whisper-1"},
        files=files,
        headers={"Authorization": "Bearer sk-proxy-key"}
    )
text = json.loads(resp.text)["text"]

```

## Summary

- **LiteLLM embedding**, **image generation**, and **audio endpoints** are routed through distinct `CallTypes` enumerations in [`litellm/types/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/types/utils.py).
- Each modality uses dedicated **ProviderConfigManager** methods in [`litellm/utils.py`](https://github.com/BerriAI/litellm/blob/main/litellm/utils.py) to resolve provider-specific request shapes and authentication.
- **Request normalization** differs by type: embeddings handle token arrays, images merge generation parameters, and audio processes multipart form-data.
- **Response transformers** in modality-specific modules ensure all providers return OpenAI-compatible JSON structures.
- The **unified Python API** (`litellm.embedding`, `litellm.image_generation`, `litellm.audio_transcriptions`) abstracts these differences while maintaining provider-specific optimizations under the hood.

## Frequently Asked Questions

### How does LiteLLM route embedding requests differently from image generation requests?

LiteLLM routes requests through the `ProxyBaseLLMRequestProcessing.base_process_llm_request` method in [`litellm/proxy/route_llm_request.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/route_llm_request.py), which inspects the `CallTypes` enum to determine the modality. Embeddings trigger `ProviderConfigManager.get_provider_embedding_config`, while image generation calls invoke `get_provider_image_generation_config`, ensuring each request type loads the correct transformation logic and endpoint URLs.

### Can I use the same provider client for embeddings and audio transcription in LiteLLM?

No. While the Python API surface is unified, LiteLLM instantiates separate provider configurations for each modality. For example, OpenAI embedding requests utilize [`litellm/llms/openai_like/embedding/handler.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/openai_like/embedding/handler.py), whereas audio transcription requests use [`litellm/llms/openai/transcriptions/whisper_transformation.py`](https://github.com/BerriAI/litellm/blob/main/litellm/llms/openai/transcriptions/whisper_transformation.py) (or equivalent provider-specific modules), because each modality requires distinct request serialization and response parsing.

### What file size limits apply to LiteLLM audio endpoints compared to text embeddings?

Audio endpoints process multipart form-data and enforce `max_file_size_mb` validation (a premium feature) within the `proxy_server.audio_transcriptions` handler. Embeddings accept only text or token arrays and have no file upload constraints, operating entirely within JSON request bodies handled by the embedding provider configs.

### Does LiteLLM cache embedding and image generation responses the same way?

Yes. All three modalities share the same **caching**, **rate-limiting**, and **budget-tracking** infrastructure defined in the base request processing layer. However, cache keys are computed differently for each `CallTypes` enum to prevent collisions between an embedding vector and an image URL, ensuring modality-specific storage and retrieval logic.