# How Structured Output with JSON Schema Validation Works in MLX‑Omni‑Server

> Learn how MLX-Omni-Server ensures structured output with JSON schema validation. Discover Pydantic models, prompt preparation, and logits-level enforcement for compliant JSON generation.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: deep-dive
- Published: 2026-03-06

---

**MLX‑Omni‑Server implements structured output by validating JSON schemas through Pydantic models at the API layer, coordinating schema constraints with thinking modes during prompt preparation, and applying logits-level enforcement during generation to guarantee compliant JSON output.**

MLX‑Omni‑Server extends Apple's MLX framework with OpenAI-compatible chat completions, including robust support for **structured output with JSON schema validation**. When you submit a request with `response_format` set to `type: "json_schema"`, the server validates the schema definition upfront, prepares the generation context to handle structural constraints alongside reasoning modes, and enforces the schema during token sampling. This ensures that language model outputs conform exactly to your specified JSON structure without post-processing.

## Schema Definition and Request Validation

The validation layer ensures that JSON schemas are well-formed and properly attached to requests before generation begins.

### The JsonSchemaFormat Model

In [`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py), the `JsonSchemaFormat` model (lines 152-166) defines the structure for schema-based output requests. This Pydantic model validates that the client-provided schema is a valid JSON object and extracts the underlying structure that guides the generation process.

### ResponseFormat Validation Logic

Surrounding the schema definition, the `ResponseFormat` class (lines 79-91 in the same file) enforces strict business rules: it validates that a schema is present **if and only if** `type == "json_schema"`, and rejects any stray schema objects when other response types like `"text"` or `"json_object"` are selected. This prevents schema pollution in non-structured generation requests and ensures API compatibility with OpenAI's specification.

## Prompt Preparation and Schema Coordination

Once validated, the schema flows into the generation pipeline where it coordinates with model-specific features like thinking or reasoning modes.

### Handling Thinking Mode and Schema Constraints

In [`src/mlx_omni_server/chat/mlx/chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/mlx/chat_generator.py), the `ChatGenerator._prepare_prompt` method (lines 59-67) manages the intersection of reasoning capabilities and structured generation. When `template_kwargs` contains both `"enable_thinking"` and a `json_schema` argument (which may be a dictionary, JSON string, or Pydantic class), the generator sets `skip_thinking_prefill`. This flag signals the downstream logits processor to handle the coordination between reasoning tokens and schema enforcement, ensuring that structured output constraints take precedence during the generation phase.

### Logits Processor Integration

The preparation stage configures the logits processor to enforce the JSON schema at the token level. By setting `skip_thinking_prefill` when both features are active, the system prevents the thinking mode's prefill tokens from interfering with schema validation. This allows the processor to match generated tokens against the schema structure and ensure only compliant JSON is produced.

## Practical API Usage

To request structured output, include a `response_format` object with `type: "json_schema"` and provide your schema definition:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:10240/v1", api_key="mlx-omni-server")

response = client.chat.completions.create(
    model="mlx-community/Llama-3.2-3B-Instruct",
    messages=[{"role": "user", "content": "List three machine learning frameworks"}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "framework_list",
            "schema": {
                "type": "object",
                "properties": {
                    "frameworks": {
                        "type": "array",
                        "items": {
                            "type": "object",
                            "properties": {
                                "name": {"type": "string"},
                                "language": {"type": "string"}
                            },
                            "required": ["name", "language"]
                        }
                    }
                },
                "required": ["frameworks"]
            }
        }
    }
)

result = response.choices[0].message.content
print(result)  # Valid JSON matching the schema

```

This request validates the schema against the `JsonSchemaFormat` model in [`schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/schema.py), then the `ChatGenerator` enforces the structure during token generation through the coordinated logits processor.

## Summary

- **Schema Validation:** MLX‑Omni‑Server validates JSON schemas using the `JsonSchemaFormat` Pydantic model in [`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py), ensuring only well-formed schemas reach the generation layer.
- **Request Integrity:** The `ResponseFormat` validators (lines 79-91) enforce that schemas accompany only `json_schema` type requests, preventing invalid API configurations.
- **Generation Coordination:** The `ChatGenerator._prepare_prompt` method coordinates between thinking modes and schema constraints by setting `skip_thinking_prefill`, enabling the downstream logits processor to enforce JSON compliance during token generation.

## Frequently Asked Questions

### What JSON Schema versions does MLX‑Omni‑Server support?

MLX‑Omni‑Server supports standard JSON Schema Draft 7 and above through its Pydantic-based validation layer in [`src/mlx_omni_server/chat/openai/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/openai/schema.py). Complex schemas with nested objects, arrays, and enumerated values are fully supported as defined in the `JsonSchemaFormat` model.

### Can I use Pydantic models directly instead of raw JSON schemas?

Yes. According to the implementation in `ChatGenerator._prepare_prompt`, the `json_schema` argument accepts a Pydantic class, dictionary, or JSON string. When a Pydantic model is provided, the system extracts the schema automatically using the model's JSON schema export functionality before validation.

### How does structured output interact with reasoning or thinking modes?

When both `enable_thinking` and a JSON schema are present in the request, the generator sets `skip_thinking_prefill` to prevent reasoning tokens from interfering with schema enforcement. This ensures the logits processor can strictly validate the final output against your schema without contamination from intermediate reasoning steps.

### Where is the schema enforced during token generation?

The schema is enforced at the logits level by a downstream processor configured during the `_prepare_prompt` stage in [`src/mlx_omni_server/chat/mlx/chat_generator.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/chat/mlx/chat_generator.py). This processor matches generated tokens against the validated schema structure in real-time, ensuring only valid JSON that conforms to the `JsonSchemaFormat` definition is produced.