How to Integrate mlx-omni-server with Existing OpenAI and Anthropic SDK Clients

The mlx-omni-server provides thin adapter layers in mlx_omni_server.chat.openai and mlx_omni_server.chat.anthropic that wrap the official Python SDKs, enabling seamless integration through FastAPI routers at /openai and /anthropic endpoints.

This guide explains how to leverage the mlx-omni-server repository to integrate with existing OpenAI and Anthropic SDK clients without modifying your application code. The architecture uses adapter patterns to bridge the Core-MLX backend with vendor-specific APIs, handling authentication, streaming, and tool-call conversion automatically.

Architecture Overview

The integration relies on specialized adapter classes that instantiate and manage official SDK clients within the FastAPI application lifecycle.

OpenAI Adapter Components

The OpenAIAdapter class in src/mlx_omni_server/chat/openai/openai_adapter.py serves as the primary wrapper for the openai.OpenAI client. It implements server-side handling for chat.completions.create, audio.speech.create, and embeddings.create operations. The adapter performs critical conversions through the _convert_tool_calls method, transforming Core-MLX tool-call formats into OpenAI's ToolCall JSON schema. It also manages optional response caching and streaming support.

The OpenAI Router in src/mlx_omni_server/chat/openai/router.py exposes all /openai/* endpoints and injects a singleton OpenAIAdapter instance into request handlers.

Anthropic Adapter Components

The AnthropicMessagesAdapter in src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py mirrors the OpenAI functionality for Anthropic's messages.create endpoint. It wraps anthropic.Anthropic instances to handle model selection, streaming responses, and tool-call conversion while maintaining API compatibility with Anthropic's native SDK contract.

The Anthropic Router in src/mlx_omni_server/chat/anthropic/router.py lazily creates and manages a singleton AnthropicMessagesAdapter, exposing endpoints under the /anthropic/* path.

Schema Definitions

Request and response models are defined in src/mlx_omni_server/chat/openai/schema.py and src/mlx_omni_server/chat/anthropic/anthropic_schema.py. These Pydantic models ensure type safety when translating between FastAPI request bodies and SDK-native payloads.

How the Adapters Work

The integration follows a five-step pipeline that maintains full compatibility with existing SDK client code:

  1. Client instantiation – Adapters lazily create SDK clients using environment variables (OPENAI_API_KEY, ANTHROPIC_API_KEY) or explicitly supplied base_url parameters.

  2. Request mapping – Incoming FastAPI request bodies defined in the schema modules are converted to the SDK's native request shape before forwarding.

  3. Tool-call handling – The _convert_tool_calls method rewrites MLX-style tool calls into OpenAI-compatible function objects, enabling function calling through the standard SDK interface.

  4. Streaming support – When stream=True is specified, the adapters yield incremental chunks through OpenAIAdapter._handle_stream, preserving the original SDK streaming contract.

  5. Unified response – Adapters return standardized Pydantic models that mirror SDK responses, allowing downstream services like Chainlit or Phidata to consume outputs without protocol changes.

Implementation Examples

Basic OpenAI SDK Integration

import os
from openai import OpenAI

# The SDK reads OPENAI_API_KEY automatically; you can also pass it directly.

client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

# Simple chat completion (the adapter is used under the hood by the router)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain the difference between HTTP and HTTPS."},
    ],
    temperature=0.7,
)

print(response.choices[0].message.content)

The request routes through OpenAIAdapter in src/mlx_omni_server/chat/openai/openai_adapter.py, which performs any necessary tool-call conversion before delegating to the underlying SDK.

Streaming Responses

from openai import OpenAI

client = OpenAI()

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Tell me a short story about a robot."}],
    stream=True,
)

for chunk in stream:
    if delta := chunk.choices[0].delta.content:
        print(delta, end="", flush=True)

The streaming logic resides in OpenAIAdapter._handle_stream and yields chunks matching the official SDK structure exactly.

Anthropic SDK Integration

import os
import anthropic

client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))

response = client.messages.create(
    model="claude-3-5-sonnet-20240620",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Write a Python function to compute the Fibonacci sequence."},
    ],
)

print(response.content[0].text)

This call processes through AnthropicMessagesAdapter in src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py, which adds tool-call and streaming support identical to the OpenAI implementation.

Function Calling with OpenAI

from openai import OpenAI

client = OpenAI()

functions = [
    {
        "name": "get_weather",
        "description": "Fetches the current weather for a location.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    }
]

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What's the weather in Paris?"}],
    tools=functions,
)

tool_calls = response.choices[0].message.tool_calls
print(tool_calls)

The conversion from MLX-style tool calls to the tools schema occurs in _convert_tool_calls within src/mlx_omni_server/chat/openai/openai_adapter.py.

Key Source Files

File Role
src/mlx_omni_server/chat/openai/openai_adapter.py Core wrapper for the OpenAI SDK; handles chat, audio, embeddings, tool conversion, streaming, and caching.
src/mlx_omni_server/chat/openai/router.py FastAPI router exposing /openai/* endpoints; manages singleton OpenAIAdapter injection.
src/mlx_omni_server/chat/openai/schema.py Pydantic models for OpenAI request/response validation.
src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py Wrapper around anthropic.Anthropic; implements messages API with streaming and tool support.
src/mlx_omni_server/chat/anthropic/router.py FastAPI router exposing /anthropic/* endpoints with lazy adapter initialization.
src/mlx_omni_server/chat/anthropic/anthropic_schema.py Request/response models for Anthropic messages API.
src/mlx_omni_server/routers.py Aggregates both OpenAI and Anthropic routers into the main FastAPI application.

Summary

  • mlx-omni-server integrates with existing OpenAI and Anthropic SDKs through adapter classes that wrap official clients in mlx_omni_server.chat.openai and mlx_omni_server.chat.anthropic packages.
  • The OpenAIAdapter and AnthropicMessagesAdapter handle authentication, request mapping, tool-call conversion, and streaming while maintaining full SDK compatibility.
  • FastAPI routers in router.py files expose endpoints at /openai and /anthropic, injecting singleton adapter instances to handle requests.
  • Environment variables OPENAI_API_KEY and ANTHROPIC_API_KEY drive client authentication, with support for custom base_url configurations.
  • Tool calls undergo automatic conversion through _convert_tool_calls to bridge MLX internal formats with vendor-specific JSON schemas.

Frequently Asked Questions

How does mlx-omni-server handle authentication for the OpenAI and Anthropic SDKs?

The adapters lazily instantiate SDK clients using standard environment variables (OPENAI_API_KEY, ANTHROPIC_API_KEY) or explicit base_url parameters passed during initialization. This allows seamless integration with existing credential management systems without requiring code changes.

Can I use streaming with both OpenAI and Anthropic integrations?

Yes. Both the OpenAIAdapter and AnthropicMessagesAdapter support streaming through their respective _handle_stream implementations. When stream=True is passed in requests, the adapters yield incremental chunks that match the official SDK streaming contracts exactly.

What tool-call format does mlx-omni-server use internally?

The server uses a Core-MLX internal format for tool calls, which the _convert_tool_calls method in src/mlx_omni_server/chat/openai/openai_adapter.py transforms into OpenAI-compatible function objects or Anthropic-equivalent schemas before forwarding to the underlying SDK clients.

Where are the request and response schemas defined for SDK integrations?

Schema definitions reside in src/mlx_omni_server/chat/openai/schema.py for OpenAI and src/mlx_omni_server/chat/anthropic/anthropic_schema.py for Anthropic. These Pydantic models validate FastAPI requests and ensure type-safe translation to SDK-native payloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →