How Structured Output with JSON Schema Validation Works in MLX‑Omni‑Server
MLX‑Omni‑Server implements structured output by validating JSON schemas through Pydantic models at the API layer, coordinating schema constraints with thinking modes during prompt preparation, and applying logits-level enforcement during generation to guarantee compliant JSON output.
MLX‑Omni‑Server extends Apple's MLX framework with OpenAI-compatible chat completions, including robust support for structured output with JSON schema validation. When you submit a request with response_format set to type: "json_schema", the server validates the schema definition upfront, prepares the generation context to handle structural constraints alongside reasoning modes, and enforces the schema during token sampling. This ensures that language model outputs conform exactly to your specified JSON structure without post-processing.
Schema Definition and Request Validation
The validation layer ensures that JSON schemas are well-formed and properly attached to requests before generation begins.
The JsonSchemaFormat Model
In src/mlx_omni_server/chat/openai/schema.py, the JsonSchemaFormat model (lines 152-166) defines the structure for schema-based output requests. This Pydantic model validates that the client-provided schema is a valid JSON object and extracts the underlying structure that guides the generation process.
ResponseFormat Validation Logic
Surrounding the schema definition, the ResponseFormat class (lines 79-91 in the same file) enforces strict business rules: it validates that a schema is present if and only if type == "json_schema", and rejects any stray schema objects when other response types like "text" or "json_object" are selected. This prevents schema pollution in non-structured generation requests and ensures API compatibility with OpenAI's specification.
Prompt Preparation and Schema Coordination
Once validated, the schema flows into the generation pipeline where it coordinates with model-specific features like thinking or reasoning modes.
Handling Thinking Mode and Schema Constraints
In src/mlx_omni_server/chat/mlx/chat_generator.py, the ChatGenerator._prepare_prompt method (lines 59-67) manages the intersection of reasoning capabilities and structured generation. When template_kwargs contains both "enable_thinking" and a json_schema argument (which may be a dictionary, JSON string, or Pydantic class), the generator sets skip_thinking_prefill. This flag signals the downstream logits processor to handle the coordination between reasoning tokens and schema enforcement, ensuring that structured output constraints take precedence during the generation phase.
Logits Processor Integration
The preparation stage configures the logits processor to enforce the JSON schema at the token level. By setting skip_thinking_prefill when both features are active, the system prevents the thinking mode's prefill tokens from interfering with schema validation. This allows the processor to match generated tokens against the schema structure and ensure only compliant JSON is produced.
Practical API Usage
To request structured output, include a response_format object with type: "json_schema" and provide your schema definition:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:10240/v1", api_key="mlx-omni-server")
response = client.chat.completions.create(
model="mlx-community/Llama-3.2-3B-Instruct",
messages=[{"role": "user", "content": "List three machine learning frameworks"}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "framework_list",
"schema": {
"type": "object",
"properties": {
"frameworks": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"language": {"type": "string"}
},
"required": ["name", "language"]
}
}
},
"required": ["frameworks"]
}
}
}
)
result = response.choices[0].message.content
print(result) # Valid JSON matching the schema
This request validates the schema against the JsonSchemaFormat model in schema.py, then the ChatGenerator enforces the structure during token generation through the coordinated logits processor.
Summary
- Schema Validation: MLX‑Omni‑Server validates JSON schemas using the
JsonSchemaFormatPydantic model insrc/mlx_omni_server/chat/openai/schema.py, ensuring only well-formed schemas reach the generation layer. - Request Integrity: The
ResponseFormatvalidators (lines 79-91) enforce that schemas accompany onlyjson_schematype requests, preventing invalid API configurations. - Generation Coordination: The
ChatGenerator._prepare_promptmethod coordinates between thinking modes and schema constraints by settingskip_thinking_prefill, enabling the downstream logits processor to enforce JSON compliance during token generation.
Frequently Asked Questions
What JSON Schema versions does MLX‑Omni‑Server support?
MLX‑Omni‑Server supports standard JSON Schema Draft 7 and above through its Pydantic-based validation layer in src/mlx_omni_server/chat/openai/schema.py. Complex schemas with nested objects, arrays, and enumerated values are fully supported as defined in the JsonSchemaFormat model.
Can I use Pydantic models directly instead of raw JSON schemas?
Yes. According to the implementation in ChatGenerator._prepare_prompt, the json_schema argument accepts a Pydantic class, dictionary, or JSON string. When a Pydantic model is provided, the system extracts the schema automatically using the model's JSON schema export functionality before validation.
How does structured output interact with reasoning or thinking modes?
When both enable_thinking and a JSON schema are present in the request, the generator sets skip_thinking_prefill to prevent reasoning tokens from interfering with schema enforcement. This ensures the logits processor can strictly validate the final output against your schema without contamination from intermediate reasoning steps.
Where is the schema enforced during token generation?
The schema is enforced at the logits level by a downstream processor configured during the _prepare_prompt stage in src/mlx_omni_server/chat/mlx/chat_generator.py. This processor matches generated tokens against the validated schema structure in real-time, ensuring only valid JSON that conforms to the JsonSchemaFormat definition is produced.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →