# How to Integrate mlx-omni-server with Existing OpenAI and Anthropic SDK Clients

> Easily integrate mlx-omni-server with your existing OpenAI and Anthropic SDK clients. Discover how our adapter layers enable seamless integration via FastAPI routers for efficient development.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: how-to-guide
- Published: 2026-03-06

---

**The mlx-omni-server provides thin adapter layers in `mlx_omni_server.chat.openai` and `mlx_omni_server.chat.anthropic` that wrap the official Python SDKs, enabling seamless integration through FastAPI routers at `/openai` and `/anthropic` endpoints.**

This guide explains how to leverage the **mlx-omni-server** repository to integrate with existing OpenAI and Anthropic SDK clients without modifying your application code. The architecture uses adapter patterns to bridge the Core-MLX backend with vendor-specific APIs, handling authentication, streaming, and tool-call conversion automatically.

## Architecture Overview

The integration relies on specialized adapter classes that instantiate and manage official SDK clients within the FastAPI application lifecycle.

### OpenAI Adapter Components

The **`OpenAIAdapter`** class in [`src/mlx_omni_server/chat/openai/openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/openai_adapter.py) serves as the primary wrapper for the `openai.OpenAI` client. It implements server-side handling for `chat.completions.create`, `audio.speech.create`, and `embeddings.create` operations. The adapter performs critical conversions through the `_convert_tool_calls` method, transforming Core-MLX tool-call formats into OpenAI's `ToolCall` JSON schema. It also manages optional response caching and streaming support.

The **`OpenAI Router`** in [`src/mlx_omni_server/chat/openai/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/router.py) exposes all `/openai/*` endpoints and injects a singleton `OpenAIAdapter` instance into request handlers.

### Anthropic Adapter Components

The **`AnthropicMessagesAdapter`** in [`src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py) mirrors the OpenAI functionality for Anthropic's `messages.create` endpoint. It wraps `anthropic.Anthropic` instances to handle model selection, streaming responses, and tool-call conversion while maintaining API compatibility with Anthropic's native SDK contract.

The **`Anthropic Router`** in [`src/mlx_omni_server/chat/anthropic/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/router.py) lazily creates and manages a singleton `AnthropicMessagesAdapter`, exposing endpoints under the `/anthropic/*` path.

### Schema Definitions

Request and response models are defined in [`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py) and [`src/mlx_omni_server/chat/anthropic/anthropic_schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_schema.py). These Pydantic models ensure type safety when translating between FastAPI request bodies and SDK-native payloads.

## How the Adapters Work

The integration follows a five-step pipeline that maintains full compatibility with existing SDK client code:

1. **Client instantiation** – Adapters lazily create SDK clients using environment variables (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`) or explicitly supplied `base_url` parameters.

2. **Request mapping** – Incoming FastAPI request bodies defined in the schema modules are converted to the SDK's native request shape before forwarding.

3. **Tool-call handling** – The `_convert_tool_calls` method rewrites MLX-style tool calls into OpenAI-compatible `function` objects, enabling function calling through the standard SDK interface.

4. **Streaming support** – When `stream=True` is specified, the adapters yield incremental chunks through `OpenAIAdapter._handle_stream`, preserving the original SDK streaming contract.

5. **Unified response** – Adapters return standardized Pydantic models that mirror SDK responses, allowing downstream services like Chainlit or Phidata to consume outputs without protocol changes.

## Implementation Examples

### Basic OpenAI SDK Integration

```python
import os
from openai import OpenAI

# The SDK reads OPENAI_API_KEY automatically; you can also pass it directly.

client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

# Simple chat completion (the adapter is used under the hood by the router)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain the difference between HTTP and HTTPS."},
    ],
    temperature=0.7,
)

print(response.choices[0].message.content)

```

The request routes through `OpenAIAdapter` in [`src/mlx_omni_server/chat/openai/openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/openai_adapter.py), which performs any necessary tool-call conversion before delegating to the underlying SDK.

### Streaming Responses

```python
from openai import OpenAI

client = OpenAI()

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Tell me a short story about a robot."}],
    stream=True,
)

for chunk in stream:
    if delta := chunk.choices[0].delta.content:
        print(delta, end="", flush=True)

```

The streaming logic resides in `OpenAIAdapter._handle_stream` and yields chunks matching the official SDK structure exactly.

### Anthropic SDK Integration

```python
import os
import anthropic

client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))

response = client.messages.create(
    model="claude-3-5-sonnet-20240620",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Write a Python function to compute the Fibonacci sequence."},
    ],
)

print(response.content[0].text)

```

This call processes through `AnthropicMessagesAdapter` in [`src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py), which adds tool-call and streaming support identical to the OpenAI implementation.

### Function Calling with OpenAI

```python
from openai import OpenAI

client = OpenAI()

functions = [
    {
        "name": "get_weather",
        "description": "Fetches the current weather for a location.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    }
]

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What's the weather in Paris?"}],
    tools=functions,
)

tool_calls = response.choices[0].message.tool_calls
print(tool_calls)

```

The conversion from MLX-style tool calls to the `tools` schema occurs in `_convert_tool_calls` within [`src/mlx_omni_server/chat/openai/openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/openai_adapter.py).

## Key Source Files

| File | Role |
|------|------|
| [`src/mlx_omni_server/chat/openai/openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/openai_adapter.py) | Core wrapper for the OpenAI SDK; handles chat, audio, embeddings, tool conversion, streaming, and caching. |
| [`src/mlx_omni_server/chat/openai/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/router.py) | FastAPI router exposing `/openai/*` endpoints; manages singleton `OpenAIAdapter` injection. |
| [`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py) | Pydantic models for OpenAI request/response validation. |
| [`src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_messages_adapter.py) | Wrapper around `anthropic.Anthropic`; implements messages API with streaming and tool support. |
| [`src/mlx_omni_server/chat/anthropic/router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/router.py) | FastAPI router exposing `/anthropic/*` endpoints with lazy adapter initialization. |
| [`src/mlx_omni_server/chat/anthropic/anthropic_schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_schema.py) | Request/response models for Anthropic messages API. |
| [`src/mlx_omni_server/routers.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/routers.py) | Aggregates both OpenAI and Anthropic routers into the main FastAPI application. |

## Summary

- **mlx-omni-server** integrates with existing OpenAI and Anthropic SDKs through adapter classes that wrap official clients in `mlx_omni_server.chat.openai` and `mlx_omni_server.chat.anthropic` packages.
- The **`OpenAIAdapter`** and **`AnthropicMessagesAdapter`** handle authentication, request mapping, tool-call conversion, and streaming while maintaining full SDK compatibility.
- FastAPI routers in [`router.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/router.py) files expose endpoints at `/openai` and `/anthropic`, injecting singleton adapter instances to handle requests.
- Environment variables `OPENAI_API_KEY` and `ANTHROPIC_API_KEY` drive client authentication, with support for custom `base_url` configurations.
- Tool calls undergo automatic conversion through `_convert_tool_calls` to bridge MLX internal formats with vendor-specific JSON schemas.

## Frequently Asked Questions

### How does mlx-omni-server handle authentication for the OpenAI and Anthropic SDKs?

The adapters lazily instantiate SDK clients using standard environment variables (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`) or explicit `base_url` parameters passed during initialization. This allows seamless integration with existing credential management systems without requiring code changes.

### Can I use streaming with both OpenAI and Anthropic integrations?

Yes. Both the `OpenAIAdapter` and `AnthropicMessagesAdapter` support streaming through their respective `_handle_stream` implementations. When `stream=True` is passed in requests, the adapters yield incremental chunks that match the official SDK streaming contracts exactly.

### What tool-call format does mlx-omni-server use internally?

The server uses a Core-MLX internal format for tool calls, which the `_convert_tool_calls` method in [`src/mlx_omni_server/chat/openai/openai_adapter.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/openai_adapter.py) transforms into OpenAI-compatible `function` objects or Anthropic-equivalent schemas before forwarding to the underlying SDK clients.

### Where are the request and response schemas defined for SDK integrations?

Schema definitions reside in [`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py) for OpenAI and [`src/mlx_omni_server/chat/anthropic/anthropic_schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/anthropic/anthropic_schema.py) for Anthropic. These Pydantic models validate FastAPI requests and ensure type-safe translation to SDK-native payloads.