OpenAI and Anthropic API Compatibility in MLX-Omni-Server: Architecture and Key Differences
MLX-Omni-Server implements parallel OpenAI and Anthropic API compatibility layers that allow you to run local MLX models using standard client SDKs, differing primarily in endpoint routing, request schemas, streaming event formats, and native support for extended thinking modes.
MLX-Omni-Server enables developers to serve local MLX models through familiar HTTP interfaces by providing dual OpenAI and Anthropic API compatibility layers. While both interfaces leverage the same underlying inference engine, they expose distinct routing paths, data schemas, and protocol behaviors that mirror their respective cloud APIs. Understanding these architectural differences ensures you select the optimal integration strategy for your specific use case.
High-Level Architecture
MLX-Omni-Server maintains two separate compatibility stacks built on top of a shared core. This design allows you to switch between OpenAI and Anthropic client libraries simply by changing the base URL, without modifying your local model deployment.
Shared Core Inference Engine
Both compatibility layers obtain a cached ChatGenerator instance via ChatGenerator.get_or_create(...) as implemented in src/mlx_omni_server/chat/mlx/chat_generator.py. This caching mechanism prevents costly model reloads and ensures that both the OpenAI and Anthropic adapters utilize the same underlying MLX inference engine. The adapters access this core through helper methods _create_text_model and _create_anthropic_model, which manage model lifecycle and tokenization.
OpenAI Compatibility Layer
The OpenAI compatibility layer resides in src/mlx_omni_server/chat/openai/router.py, mounting endpoints at /v1/chat/completions and /chat/completions. The OpenAIAdapter class in src/mlx_omni_server/chat/openai/openai_adapter.py handles translation between OpenAI's ChatCompletionRequest schema (defined in src/mlx_omni_server/chat/openai/schema.py) and the internal ChatGenerator calls. This layer supports additional capabilities like audio transcription (STT), text-to-speech (TTS), image generation, and embeddings through separate routers in src/mlx_omni_server/tts, src/mlx_omni_server/stt, and related modules.
Anthropic Compatibility Layer
The Anthropic compatibility layer operates through src/mlx_omni_server/chat/anthropic/router.py, exposing the Messages API at /anthropic/v1/messages and /v1/messages. The AnthropicMessagesAdapter in src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py manages conversion between Anthropic's MessagesRequest schema (defined in src/mlx_omni_server/chat/anthropic/anthropic_schema.py) and the shared generator. This layer natively supports Claude-specific features like extended thinking blocks and distinct tool-use content formats.
Key Differences in API Implementation
Endpoint Routing and URL Structure
The routing architectures differ in their base path conventions and endpoint organization:
-
OpenAI: Uses standard OpenAI-style paths including
/v1/chat/completions,/v1/models,/v1/audio/speech, and/v1/images/generations. The router supports both the standard/v1/prefix and an internal/chat/prefix for flexibility. -
Anthropic: Implements Claude's Messages API structure with
/anthropic/v1/messagesas the primary chat endpoint, along with/anthropic/v1/models. The routing explicitly separates Anthropic-compatible endpoints under the/anthropic/prefix to avoid collisions with OpenAI routes.
Request and Response Schemas
Each compatibility layer maintains distinct Pydantic models that mirror their respective cloud API contracts:
OpenAI schemas (src/mlx_omni_server/chat/openai/schema.py):
ChatCompletionRequest: Acceptsmessagesas an array of objects withroleandcontent, plus optionaltools,tool_choice, andresponse_formatfields.ChatCompletionResponse: ReturnschoicescontainingChatMessageobjects with optionaltool_callsarrays.- Stop reasons: Maps directly to OpenAI's
finish_reasonvalues (stop,length,tool_calls, etc.).
Anthropic schemas (src/mlx_omni_server/chat/anthropic/anthropic_schema.py):
MessagesRequest: Structures input as amessagesarray but usesmax_tokensas a required parameter and supports a nativethinkingfield for reasoning budgets.MessagesResponse: Returnscontentas an array ofContentBlockobjects that can represent text,tool_use, or thinking blocks.- Stop reasons: Uses Anthropic-specific enum values (
END_TURN,MAX_TOKENS,STOP_SEQUENCE,TOOL_USE).
Streaming Implementations
The streaming protocols differ significantly in their event structures and chunk formats:
OpenAI streaming produces ChatCompletionChunk objects where each chunk contains a delta field with incremental content updates. You can optionally set stream_options.include_usage to receive a final usage chunk at the end of the stream.
Anthropic streaming emits MessageStreamEvent objects with distinct event types including message_start, content_block_start, content_block_delta, and message_delta. Usage statistics appear in the final MESSAGE_DELTA event rather than as a separate chunk, following Claude's native streaming protocol.
Tool Calling Formats
Both layers support function calling but serialize tool interactions differently:
-
OpenAI: Represents tool calls within the
ChatMessagemodel using atool_callsarray containingid,type, andfunctionobjects withnameandargumentsstrings. -
Anthropic: Embeds tool interactions as
tool_usecontent blocks within the response'scontentarray, each containingid,name, andinputfields. The adapter converts between these representations while maintaining the internal tool execution logic.
Thinking and Reasoning Modes
The handling of extended reasoning or "thinking" capabilities represents a major functional difference:
OpenAI compatibility exposes thinking mode indirectly through the generic response_format parameter or via extra_body parameters like enable_thinking. This approach treats reasoning as an extension of the standard completion flow.
Anthropic compatibility provides first-class support through the native thinking field in MessagesRequest, which accepts a budget token count. The response includes dedicated thinking content blocks that are separate from text outputs, allowing applications to distinguish between reasoning traces and final answers.
Usage Tracking and Metadata
Usage reporting conventions vary between the two implementations:
-
OpenAI: Optional usage reporting controlled by
stream_options.include_usage. When enabled, emits a final chunk containingprompt_tokensandcompletion_tokens. -
Anthropic: Always includes usage data in the final
MessagesResponseand streams usage statistics via theusagefield on the finalMESSAGE_DELTAevent, providing consistent visibility into token consumption.
Feature Availability Comparison
While both APIs share core chat capabilities, the OpenAI compatibility layer includes additional endpoints not present in the Anthropic implementation:
| Feature | OpenAI Layer | Anthropic Layer |
|---|---|---|
| Chat completions | ✅ /v1/chat/completions |
✅ /anthropic/v1/messages |
| Audio (TTS/STT) | ✅ /v1/audio/* |
❌ Not supported |
| Image generation | ✅ /v1/images/generations |
❌ Not supported |
| Embeddings | ✅ /v1/embeddings |
❌ Not supported |
| Tool calling | ✅ tool_calls format |
✅ tool_use blocks |
| Streaming | ✅ ChatCompletionChunk |
✅ MessageStreamEvent |
| Extended thinking | ⚠️ Via extra_body |
✅ Native thinking field |
| Model pagination | Simple list | Cursor-based (before_id, after_id) |
Implementation Examples
OpenAI-Compatible Client Usage
To interact with the OpenAI compatibility layer, point your client to the /v1 endpoint:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:10240/v1",
api_key="not-needed"
)
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Fetch current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"]
},
},
}
]
response = client.chat.completions.create(
model="mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit",
messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
tools=tools,
tool_choice="auto"
)
print(response.choices[0].message.content)
print(response.choices[0].message.tool_calls)
The openai_adapter.py file handles the conversion between these SDK calls and the internal ChatGenerator interface, mapping tool_calls responses appropriately.
Anthropic-Compatible Client Usage
For Anthropic compatibility, configure your client to use the /anthropic prefix:
import anthropic
client = anthropic.Anthropic(
base_url="http://localhost:10240/anthropic",
api_key="not-needed"
)
tools = [
{
"name": "get_weather",
"description": "Fetch current weather for a location",
"input_schema": {
"type": "object",
"properties": {
"location": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["location"]
},
}
]
msg = client.messages.create(
model="mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit",
max_tokens=1024,
tools=tools,
messages=[
{"role": "user", "content": "Check the weather in Tokyo and send an email."}
],
)
for block in msg.content:
if block.type == "text":
print("Assistant:", block.text)
elif block.type == "tool_use":
print("Tool call:", block.name, block.input)
The anthropic_messages_adapter.py translates these requests into ChatGenerator calls and reformats the output into Anthropic's expected ContentBlock structure.
Streaming with OpenAI
OpenAI-style streaming yields incremental text updates through the delta field:
for chunk in client.chat.completions.create(
model="mlx-community/Llama-3.2-3B-Instruct-4bit",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True,
):
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Streaming with Anthropic
Anthropic streaming requires handling specific event types during iteration:
with client.messages.stream(
model="mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit",
max_tokens=1000,
messages=[{"role": "user", "content": "Tell me a joke"}],
) as stream:
for event in stream:
if event.type == "content_block_delta" and hasattr(event.delta, "text"):
print(event.delta.text, end="", flush=True)
Core Source Files
Understanding the repository structure helps when debugging or extending the compatibility layers:
src/mlx_omni_server/chat/openai/router.py: FastAPI routes for OpenAI-compatible endpoints including/v1/chat/completionsand/v1/models.src/mlx_omni_server/chat/openai/openai_adapter.py: Translates OpenAI request objects toChatGeneratorcalls; handlestool_callsconversion, streaming chunk generation, and usage aggregation.src/mlx_omni_server/chat/openai/schema.py: Pydantic models for OpenAI requests/responses includingChatCompletionRequest,ChatMessage, andToolCall.src/mlx_omni_server/chat/anthropic/router.py: FastAPI routes for Anthropic-compatible endpoints at/anthropic/v1/messagesand/anthropic/v1/models.src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py: Maps Anthropic Messages API to the sharedChatGenerator; implements thinking mode,tool_useblocks, and streaming event formatting.src/mlx_omni_server/chat/anthropic/anthropic_schema.py: Pydantic models for Anthropic requests/responses includingMessagesRequest,MessagesResponse, andContentBlock.src/mlx_omni_server/chat/mlx/chat_generator.py: Core generator that loads MLX models, manages caching viaget_or_create(), and providesgenerate()andgenerate_stream()methods used by both adapters.
Summary
- MLX-Omni-Server provides dual OpenAI and Anthropic API compatibility layers that share the same
ChatGeneratorinference engine but expose different routing paths and schemas. - OpenAI compatibility supports additional modalities including audio, images, and embeddings, while using
tool_callsarrays and optional streaming usage flags. - Anthropic compatibility offers native extended thinking blocks, cursor-based model pagination, and
tool_usecontent blocks within a Messages API structure. - Both adapters cache model instances through
ChatGenerator.get_or_create()to prevent reload overhead when switching between API styles. - Streaming implementations differ in event types: OpenAI uses
ChatCompletionChunkwith delta updates, while Anthropic usesMessageStreamEventwith explicit event type discrimination. - Tool calling formats vary between the
tool_callsfield (OpenAI) andtool_usecontent blocks (Anthropic), with adapters handling bidirectional conversion.
Frequently Asked Questions
Can I use both OpenAI and Anthropic clients simultaneously against the same server instance?
Yes. MLX-Omni-Server runs both compatibility layers concurrently on different routes. You can point an OpenAI client to http://localhost:10240/v1 and an Anthropic client to http://localhost:10240/anthropic simultaneously. Both requests will utilize the same cached model instance through the shared ChatGenerator core, though each client must use its respective API schema.
Why does the Anthropic layer support thinking mode while OpenAI requires extra parameters?
The Anthropic Messages API specification includes a native thinking field for extended reasoning budgets, which MLX-Omni-Server implements directly in anthropic_schema.py. The OpenAI API specification does not define a standard thinking parameter, so the adapter exposes this capability through extra_body parameters or response_format extensions rather than the core schema. This reflects the actual differences between the official cloud APIs.
Are tool calls interoperable between the two API styles?
While both APIs support function calling, the request and response formats differ. The OpenAI adapter expects tools with function objects and returns tool_calls arrays, while the Anthropic adapter expects tools with input_schema and returns tool_use content blocks. You cannot mix formats—requests must conform to the specific API style you are using, though the underlying tool execution logic remains consistent.
Which API should I use for multimodal features like audio or image generation?
Use the OpenAI compatibility layer for multimodal features. The OpenAI router exposes additional endpoints for text-to-speech (/v1/audio/speech), speech-to-text (/v1/audio/transcriptions), and image generation (/v1/images/generations) that are not part of the Anthropic API specification. The Anthropic compatibility layer focuses exclusively on text-based chat completions and tool use.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →