How to Enable and Use Thinking Mode for Reasoning Models in mlx-omni-server
Set thinking_mode=True in your chat request to activate chain-of-thought reasoning, which injects a system prompt and uses the ThinkingDecoder to separate intermediate reasoning from the final answer.
The mlx-omni-server repository provides a unified API for multiple LLM providers, and its thinking mode feature allows any supported model to generate explicit reasoning steps before producing tool calls or final responses. This capability is implemented through a middleware layer and a specialized decoder that processes the model's output stream.
What Is Thinking Mode and How It Works
Thinking mode transforms standard chat completions into structured reasoning workflows. When enabled, the server:
- Injects a chain-of-thought system prompt via the
ThinkingMiddlewarelocated in [src/mlx_omni_server/middleware/thinking.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/middleware/thinking.py) - Intercepts the raw token stream using the
ThinkingDecoderclass in [src/mlx_omni_server/chat/thinking_decoder.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py) - Splits output into segments by looking for the delimiter
\n===(configurable) to separate the "thoughts" block from the final "answer" or action
The request schema in [src/mlx_omni_server/chat/schema.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/schema.py) defines the boolean field thinking_mode that controls this behavior. When true, the router in [src/mlx_omni_server/routers.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/routers.py) forwards the payload through the thinking pipeline.
Enabling Thinking Mode in Your Requests
You can activate thinking mode either through the high-level Python client or by calling the HTTP API directly.
Using the Python Client
The ChatClient class accepts a thinking_mode parameter that defaults to False:
from mlx_omni_server.chat.client import ChatClient
client = ChatClient(base_url="http://localhost:8000")
response = client.chat(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Explain the quicksort algorithm step by step."}],
thinking_mode=True, # Enable chain-of-thought reasoning
)
print(response["content"])
When thinking_mode=True, the client automatically adds the flag to the JSON body sent to the /chat/completions endpoint, triggering the full thinking pipeline on the server.
Using the HTTP API Directly
Send a POST request with the thinking_mode field set to true:
curl -X POST http://localhost:8000/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "How does binary search work?"}],
"thinking_mode": true
}'
The server will return a response containing both the model's internal reasoning and the final formatted answer, separated by the configured delimiter.
Customizing the Thinking Decoder
Advanced users can modify how the server separates thoughts from actions. The ThinkingDecoder class uses a configurable delimiter constant defined in [src/mlx_omni_server/chat/thinking_decoder.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py):
# In thinking_decoder.py
THINKING_DELIMITER = "\n===" # Change this to customize the split marker
After modifying the delimiter and restarting the server, the decoder will automatically respect the new boundary when parsing model outputs. This allows compatibility with models that use different formatting conventions for their reasoning steps.
Testing and Validation
The repository includes comprehensive tests for the thinking functionality. The test file [tests/chat/mlx/test_thinking_decoder.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/tests/chat/mlx/test_thinking_decoder.py) validates that the decoder correctly identifies and separates thought segments:
from mlx_omni_server.chat.thinking_decoder import ThinkingDecoder
def test_thinking_decoder_split():
raw_output = "Let me think step by step...\n===\nThe answer is 42."
decoded = ThinkingDecoder().decode(raw_output)
assert "thoughts" in decoded
assert "answer" in decoded
assert decoded["thoughts"] == "Let me think step by step..."
Additionally, the Jupyter notebook examples/chat_thinking.ipynb provides an interactive demonstration that starts the FastAPI server and visualizes the before/after effect of enabling thinking mode on real queries.
Summary
- Enable thinking mode by setting
thinking_mode=Truein API requests or Python client calls - Architecture relies on
ThinkingMiddlewarefor prompt injection andThinkingDecoderfor output parsing - Key files: [
src/mlx_omni_server/middleware/thinking.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/middleware/thinking.py), [src/mlx_omni_server/chat/thinking_decoder.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/thinking_decoder.py), and [src/mlx_omni_server/chat/schema.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/schema.py) - Default delimiter is
\n===but can be customized in the decoder source - Works uniformly across OpenAI, Anthropic, Mistral, LLaMA 3, Qwen 3, and other supported providers
Frequently Asked Questions
How do I disable thinking mode after enabling it?
Simply omit the thinking_mode parameter or explicitly set it to false in your request. The default behavior in [src/mlx_omni_server/chat/client.py](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/client.py) uses thinking_mode=False, so standard chat calls will not trigger the reasoning pipeline unless you opt in.
Can I use thinking mode with streaming responses?
Yes, the ThinkingDecoder processes token streams in real-time. As the model generates content, the decoder buffers tokens until it encounters the delimiter, then yields the structured response with separate "thoughts" and "answer" fields. This works identically for both streaming and non-streaming requests to the /chat/completions endpoint.
Which models support thinking mode?
Thinking mode is provider-agnostic and works with any model supported by mlx-omni-server, including GPT-4o, Claude 3, Mistral Large, LLaMA 3, and Qwen 3. The feature relies on prompting techniques rather than model-specific capabilities, though reasoning quality varies by model size and training.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →