# OpenAI and Anthropic API Compatibility in MLX-Omni-Server: Architecture and Key Differences

> Explore OpenAI vs Anthropic API compatibility in MLX Omni Server. Understand key differences in routing, schemas, streaming, and extended thinking for local MLX models.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: deep-dive
- Published: 2026-03-06

---

**MLX-Omni-Server implements parallel OpenAI and Anthropic API compatibility layers that allow you to run local MLX models using standard client SDKs, differing primarily in endpoint routing, request schemas, streaming event formats, and native support for extended thinking modes.**

MLX-Omni-Server enables developers to serve local MLX models through familiar HTTP interfaces by providing dual OpenAI and Anthropic API compatibility layers. While both interfaces leverage the same underlying inference engine, they expose distinct routing paths, data schemas, and protocol behaviors that mirror their respective cloud APIs. Understanding these architectural differences ensures you select the optimal integration strategy for your specific use case.

## High-Level Architecture

MLX-Omni-Server maintains two separate compatibility stacks built on top of a shared core. This design allows you to switch between OpenAI and Anthropic client libraries simply by changing the base URL, without modifying your local model deployment.

### Shared Core Inference Engine

Both compatibility layers obtain a cached **`ChatGenerator`** instance via `ChatGenerator.get_or_create(...)` as implemented in [`src/mlx_omni_server/chat/mlx/chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/mlx/chat_generator.py). This caching mechanism prevents costly model reloads and ensures that both the OpenAI and Anthropic adapters utilize the same underlying MLX inference engine. The adapters access this core through helper methods `_create_text_model` and `_create_anthropic_model`, which manage model lifecycle and tokenization.

### OpenAI Compatibility Layer

The OpenAI compatibility layer resides in [`src/mlx_omni_server/chat/openai/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/router.py), mounting endpoints at `/v1/chat/completions` and `/chat/completions`. The **`OpenAIAdapter`** class in [`src/mlx_omni_server/chat/openai/openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/openai_adapter.py) handles translation between OpenAI's `ChatCompletionRequest` schema (defined in [`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py)) and the internal `ChatGenerator` calls. This layer supports additional capabilities like audio transcription (STT), text-to-speech (TTS), image generation, and embeddings through separate routers in `src/mlx_omni_server/tts`, `src/mlx_omni_server/stt`, and related modules.

### Anthropic Compatibility Layer

The Anthropic compatibility layer operates through [`src/mlx_omni_server/chat/anthropic/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/router.py), exposing the Messages API at `/anthropic/v1/messages` and `/v1/messages`. The **`AnthropicMessagesAdapter`** in [`src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py) manages conversion between Anthropic's `MessagesRequest` schema (defined in [`src/mlx_omni_server/chat/anthropic/anthropic_schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_schema.py)) and the shared generator. This layer natively supports Claude-specific features like extended thinking blocks and distinct tool-use content formats.

## Key Differences in API Implementation

### Endpoint Routing and URL Structure

The routing architectures differ in their base path conventions and endpoint organization:

- **OpenAI**: Uses standard OpenAI-style paths including `/v1/chat/completions`, `/v1/models`, `/v1/audio/speech`, and `/v1/images/generations`. The router supports both the standard `/v1/` prefix and an internal `/chat/` prefix for flexibility.

- **Anthropic**: Implements Claude's Messages API structure with `/anthropic/v1/messages` as the primary chat endpoint, along with `/anthropic/v1/models`. The routing explicitly separates Anthropic-compatible endpoints under the `/anthropic/` prefix to avoid collisions with OpenAI routes.

### Request and Response Schemas

Each compatibility layer maintains distinct Pydantic models that mirror their respective cloud API contracts:

**OpenAI schemas** ([`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py)):
- **`ChatCompletionRequest`**: Accepts `messages` as an array of objects with `role` and `content`, plus optional `tools`, `tool_choice`, and `response_format` fields.
- **`ChatCompletionResponse`**: Returns `choices` containing `ChatMessage` objects with optional `tool_calls` arrays.
- **Stop reasons**: Maps directly to OpenAI's `finish_reason` values (`stop`, `length`, `tool_calls`, etc.).

**Anthropic schemas** ([`src/mlx_omni_server/chat/anthropic/anthropic_schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_schema.py)):
- **`MessagesRequest`**: Structures input as a `messages` array but uses `max_tokens` as a required parameter and supports a native `thinking` field for reasoning budgets.
- **`MessagesResponse`**: Returns `content` as an array of `ContentBlock` objects that can represent text, `tool_use`, or thinking blocks.
- **Stop reasons**: Uses Anthropic-specific enum values (`END_TURN`, `MAX_TOKENS`, `STOP_SEQUENCE`, `TOOL_USE`).

### Streaming Implementations

The streaming protocols differ significantly in their event structures and chunk formats:

**OpenAI streaming** produces `ChatCompletionChunk` objects where each chunk contains a `delta` field with incremental content updates. You can optionally set `stream_options.include_usage` to receive a final usage chunk at the end of the stream.

**Anthropic streaming** emits `MessageStreamEvent` objects with distinct event types including `message_start`, `content_block_start`, `content_block_delta`, and `message_delta`. Usage statistics appear in the final `MESSAGE_DELTA` event rather than as a separate chunk, following Claude's native streaming protocol.

### Tool Calling Formats

Both layers support function calling but serialize tool interactions differently:

- **OpenAI**: Represents tool calls within the `ChatMessage` model using a `tool_calls` array containing `id`, `type`, and `function` objects with `name` and `arguments` strings.

- **Anthropic**: Embeds tool interactions as `tool_use` content blocks within the response's `content` array, each containing `id`, `name`, and `input` fields. The adapter converts between these representations while maintaining the internal tool execution logic.

### Thinking and Reasoning Modes

The handling of extended reasoning or "thinking" capabilities represents a major functional difference:

**OpenAI compatibility** exposes thinking mode indirectly through the generic `response_format` parameter or via `extra_body` parameters like `enable_thinking`. This approach treats reasoning as an extension of the standard completion flow.

**Anthropic compatibility** provides first-class support through the native `thinking` field in `MessagesRequest`, which accepts a budget token count. The response includes dedicated thinking content blocks that are separate from text outputs, allowing applications to distinguish between reasoning traces and final answers.

### Usage Tracking and Metadata

Usage reporting conventions vary between the two implementations:

- **OpenAI**: Optional usage reporting controlled by `stream_options.include_usage`. When enabled, emits a final chunk containing `prompt_tokens` and `completion_tokens`.

- **Anthropic**: Always includes usage data in the final `MessagesResponse` and streams usage statistics via the `usage` field on the final `MESSAGE_DELTA` event, providing consistent visibility into token consumption.

## Feature Availability Comparison

While both APIs share core chat capabilities, the OpenAI compatibility layer includes additional endpoints not present in the Anthropic implementation:

| Feature | OpenAI Layer | Anthropic Layer |
|---------|--------------|-----------------|
| Chat completions | ✅ `/v1/chat/completions` | ✅ `/anthropic/v1/messages` |
| Audio (TTS/STT) | ✅ `/v1/audio/*` | ❌ Not supported |
| Image generation | ✅ `/v1/images/generations` | ❌ Not supported |
| Embeddings | ✅ `/v1/embeddings` | ❌ Not supported |
| Tool calling | ✅ `tool_calls` format | ✅ `tool_use` blocks |
| Streaming | ✅ `ChatCompletionChunk` | ✅ `MessageStreamEvent` |
| Extended thinking | ⚠️ Via `extra_body` | ✅ Native `thinking` field |
| Model pagination | Simple list | Cursor-based (`before_id`, `after_id`) |

## Implementation Examples

### OpenAI-Compatible Client Usage

To interact with the OpenAI compatibility layer, point your client to the `/v1` endpoint:

```python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:10240/v1",
    api_key="not-needed"
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Fetch current weather for a location",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string"},
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
                },
                "required": ["location"]
            },
        },
    }
]

response = client.chat.completions.create(
    model="mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=tools,
    tool_choice="auto"
)

print(response.choices[0].message.content)
print(response.choices[0].message.tool_calls)

```

The [`openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/openai_adapter.py) file handles the conversion between these SDK calls and the internal `ChatGenerator` interface, mapping `tool_calls` responses appropriately.

### Anthropic-Compatible Client Usage

For Anthropic compatibility, configure your client to use the `/anthropic` prefix:

```python
import anthropic

client = anthropic.Anthropic(
    base_url="http://localhost:10240/anthropic",
    api_key="not-needed"
)

tools = [
    {
        "name": "get_weather",
        "description": "Fetch current weather for a location",
        "input_schema": {
            "type": "object",
            "properties": {
                "location": {"type": "string"},
                "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
            },
            "required": ["location"]
        },
    }
]

msg = client.messages.create(
    model="mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit",
    max_tokens=1024,
    tools=tools,
    messages=[
        {"role": "user", "content": "Check the weather in Tokyo and send an email."}
    ],
)

for block in msg.content:
    if block.type == "text":
        print("Assistant:", block.text)
    elif block.type == "tool_use":
        print("Tool call:", block.name, block.input)

```

The [`anthropic_messages_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/anthropic_messages_adapter.py) translates these requests into `ChatGenerator` calls and reformats the output into Anthropic's expected `ContentBlock` structure.

### Streaming with OpenAI

OpenAI-style streaming yields incremental text updates through the `delta` field:

```python
for chunk in client.chat.completions.create(
    model="mlx-community/Llama-3.2-3B-Instruct-4bit",
    messages=[{"role": "user", "content": "Tell me a story"}],
    stream=True,
):
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

```

### Streaming with Anthropic

Anthropic streaming requires handling specific event types during iteration:

```python
with client.messages.stream(
    model="mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit",
    max_tokens=1000,
    messages=[{"role": "user", "content": "Tell me a joke"}],
) as stream:
    for event in stream:
        if event.type == "content_block_delta" and hasattr(event.delta, "text"):
            print(event.delta.text, end="", flush=True)

```

## Core Source Files

Understanding the repository structure helps when debugging or extending the compatibility layers:

- **[`src/mlx_omni_server/chat/openai/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/router.py)**: FastAPI routes for OpenAI-compatible endpoints including `/v1/chat/completions` and `/v1/models`.
- **[`src/mlx_omni_server/chat/openai/openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/openai_adapter.py)**: Translates OpenAI request objects to `ChatGenerator` calls; handles `tool_calls` conversion, streaming chunk generation, and usage aggregation.
- **[`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py)**: Pydantic models for OpenAI requests/responses including `ChatCompletionRequest`, `ChatMessage`, and `ToolCall`.
- **[`src/mlx_omni_server/chat/anthropic/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/router.py)**: FastAPI routes for Anthropic-compatible endpoints at `/anthropic/v1/messages` and `/anthropic/v1/models`.
- **[`src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py)**: Maps Anthropic Messages API to the shared `ChatGenerator`; implements thinking mode, `tool_use` blocks, and streaming event formatting.
- **[`src/mlx_omni_server/chat/anthropic/anthropic_schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_schema.py)**: Pydantic models for Anthropic requests/responses including `MessagesRequest`, `MessagesResponse`, and `ContentBlock`.
- **[`src/mlx_omni_server/chat/mlx/chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/mlx/chat_generator.py)**: Core generator that loads MLX models, manages caching via `get_or_create()`, and provides `generate()` and `generate_stream()` methods used by both adapters.

## Summary

- **MLX-Omni-Server provides dual OpenAI and Anthropic API compatibility layers** that share the same `ChatGenerator` inference engine but expose different routing paths and schemas.
- **OpenAI compatibility** supports additional modalities including audio, images, and embeddings, while using `tool_calls` arrays and optional streaming usage flags.
- **Anthropic compatibility** offers native extended thinking blocks, cursor-based model pagination, and `tool_use` content blocks within a Messages API structure.
- **Both adapters cache model instances** through `ChatGenerator.get_or_create()` to prevent reload overhead when switching between API styles.
- **Streaming implementations differ** in event types: OpenAI uses `ChatCompletionChunk` with delta updates, while Anthropic uses `MessageStreamEvent` with explicit event type discrimination.
- **Tool calling formats vary** between the `tool_calls` field (OpenAI) and `tool_use` content blocks (Anthropic), with adapters handling bidirectional conversion.

## Frequently Asked Questions

### Can I use both OpenAI and Anthropic clients simultaneously against the same server instance?

Yes. MLX-Omni-Server runs both compatibility layers concurrently on different routes. You can point an OpenAI client to `http://localhost:10240/v1` and an Anthropic client to `http://localhost:10240/anthropic` simultaneously. Both requests will utilize the same cached model instance through the shared `ChatGenerator` core, though each client must use its respective API schema.

### Why does the Anthropic layer support thinking mode while OpenAI requires extra parameters?

The Anthropic Messages API specification includes a native `thinking` field for extended reasoning budgets, which MLX-Omni-Server implements directly in [`anthropic_schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/anthropic_schema.py). The OpenAI API specification does not define a standard thinking parameter, so the adapter exposes this capability through `extra_body` parameters or `response_format` extensions rather than the core schema. This reflects the actual differences between the official cloud APIs.

### Are tool calls interoperable between the two API styles?

While both APIs support function calling, the request and response formats differ. The OpenAI adapter expects `tools` with `function` objects and returns `tool_calls` arrays, while the Anthropic adapter expects `tools` with `input_schema` and returns `tool_use` content blocks. You cannot mix formats—requests must conform to the specific API style you are using, though the underlying tool execution logic remains consistent.

### Which API should I use for multimodal features like audio or image generation?

Use the **OpenAI compatibility layer** for multimodal features. The OpenAI router exposes additional endpoints for text-to-speech (`/v1/audio/speech`), speech-to-text (`/v1/audio/transcriptions`), and image generation (`/v1/images/generations`) that are not part of the Anthropic API specification. The Anthropic compatibility layer focuses exclusively on text-based chat completions and tool use.