# How to Enable and Use Thinking Mode for Reasoning Models in mlx-omni-server

> Learn how to enable and use thinking mode for reasoning models in mlx-omni-server by setting thinking_mode=True. Enhance your model's reasoning capabilities.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Set `thinking_mode=True` in your chat request to activate chain-of-thought reasoning, which injects a system prompt and uses the `ThinkingDecoder` to separate intermediate reasoning from the final answer.**

The **mlx-omni-server** repository provides a unified API for multiple LLM providers, and its thinking mode feature allows any supported model to generate explicit reasoning steps before producing tool calls or final responses. This capability is implemented through a middleware layer and a specialized decoder that processes the model's output stream.

## What Is Thinking Mode and How It Works

Thinking mode transforms standard chat completions into structured reasoning workflows. When enabled, the server:

1. **Injects a chain-of-thought system prompt** via the `ThinkingMiddleware` located in [[`src/mlx_omni_server/middleware/thinking.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/middleware/thinking.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/middleware/thinking.py)
2. **Intercepts the raw token stream** using the `ThinkingDecoder` class in [[`src/mlx_omni_server/chat/thinking_decoder.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py)
3. **Splits output into segments** by looking for the delimiter `\n===` (configurable) to separate the "thoughts" block from the final "answer" or action

The request schema in [[`src/mlx_omni_server/chat/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/schema.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/schema.py) defines the boolean field `thinking_mode` that controls this behavior. When `true`, the router in [[`src/mlx_omni_server/routers.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/routers.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/routers.py) forwards the payload through the thinking pipeline.

## Enabling Thinking Mode in Your Requests

You can activate thinking mode either through the high-level Python client or by calling the HTTP API directly.

### Using the Python Client

The [`ChatClient`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/client.py) class accepts a `thinking_mode` parameter that defaults to `False`:

```python
from mlx_omni_server.chat.client import ChatClient

client = ChatClient(base_url="http://localhost:8000")

response = client.chat(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain the quicksort algorithm step by step."}],
    thinking_mode=True,  # Enable chain-of-thought reasoning

)

print(response["content"])

```

When `thinking_mode=True`, the client automatically adds the flag to the JSON body sent to the `/chat/completions` endpoint, triggering the full thinking pipeline on the server.

### Using the HTTP API Directly

Send a POST request with the `thinking_mode` field set to `true`:

```bash
curl -X POST http://localhost:8000/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4o-mini",
        "messages": [{"role": "user", "content": "How does binary search work?"}],
        "thinking_mode": true
      }'

```

The server will return a response containing both the model's internal reasoning and the final formatted answer, separated by the configured delimiter.

## Customizing the Thinking Decoder

Advanced users can modify how the server separates thoughts from actions. The `ThinkingDecoder` class uses a configurable delimiter constant defined in [[`src/mlx_omni_server/chat/thinking_decoder.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py):

```python

# In thinking_decoder.py

THINKING_DELIMITER = "\n==="  # Change this to customize the split marker

```

After modifying the delimiter and restarting the server, the decoder will automatically respect the new boundary when parsing model outputs. This allows compatibility with models that use different formatting conventions for their reasoning steps.

## Testing and Validation

The repository includes comprehensive tests for the thinking functionality. The test file [[`tests/chat/mlx/test_thinking_decoder.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/tests/chat/mlx/test_thinking_decoder.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/tests/chat/mlx/test_thinking_decoder.py) validates that the decoder correctly identifies and separates thought segments:

```python
from mlx_omni_server.chat.thinking_decoder import ThinkingDecoder

def test_thinking_decoder_split():
    raw_output = "Let me think step by step...\n===\nThe answer is 42."
    decoded = ThinkingDecoder().decode(raw_output)
    
    assert "thoughts" in decoded
    assert "answer" in decoded
    assert decoded["thoughts"] == "Let me think step by step..."

```

Additionally, the Jupyter notebook [`examples/chat_thinking.ipynb`](https://github.com/madroidmaq/mlx-omni-server/blob/main/examples/chat_thinking.ipynb) provides an interactive demonstration that starts the FastAPI server and visualizes the before/after effect of enabling thinking mode on real queries.

## Summary

- **Enable thinking mode** by setting `thinking_mode=True` in API requests or Python client calls
- **Architecture** relies on `ThinkingMiddleware` for prompt injection and `ThinkingDecoder` for output parsing
- **Key files**: [[`src/mlx_omni_server/middleware/thinking.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/middleware/thinking.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/middleware/thinking.py), [[`src/mlx_omni_server/chat/thinking_decoder.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py), and [[`src/mlx_omni_server/chat/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/schema.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/schema.py)
- **Default delimiter** is `\n===` but can be customized in the decoder source
- **Works uniformly** across OpenAI, Anthropic, Mistral, LLaMA 3, Qwen 3, and other supported providers

## Frequently Asked Questions

### How do I disable thinking mode after enabling it?

Simply omit the `thinking_mode` parameter or explicitly set it to `false` in your request. The default behavior in [[`src/mlx_omni_server/chat/client.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/client.py)](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/client.py) uses `thinking_mode=False`, so standard chat calls will not trigger the reasoning pipeline unless you opt in.

### Can I use thinking mode with streaming responses?

Yes, the `ThinkingDecoder` processes token streams in real-time. As the model generates content, the decoder buffers tokens until it encounters the delimiter, then yields the structured response with separate "thoughts" and "answer" fields. This works identically for both streaming and non-streaming requests to the `/chat/completions` endpoint.

### Which models support thinking mode?

Thinking mode is provider-agnostic and works with any model supported by mlx-omni-server, including GPT-4o, Claude 3, Mistral Large, LLaMA 3, and Qwen 3. The feature relies on prompting techniques rather than model-specific capabilities, though reasoning quality varies by model size and training.