How to Use Structured Outputs (JSON Schema, Regex, Grammar) in vLLM
vLLM supports structured output constraints that force language models to generate text matching JSON schemas, regular expressions, or custom grammars by applying token-level bitmasks during decoding.
vLLM enables precise control over model generations through structured outputs in vLLM, allowing you to constrain responses to validate against JSON schemas, match regular expressions, or follow custom grammars. According to the vllm-project/vllm source code, this feature works by compiling constraints into token-level finite state machines (FSMs) that the inference engine enforces at every generation step via StructuredOutputManager in vllm/v1/structured_output/__init__.py.
How Structured Outputs Work in vLLM
The structured output pipeline operates in four stages as implemented in vllm/v1/structured_output/__init__.py:
-
Request Parsing – The engine extracts the
structured_outputsfield from OpenAI-compatible requests invllm/v1/structured_output/request.py, or theresponse_formatfield for JSON schemas invllm/entrypoints/openai/chat_completion/protocol.py. -
Backend Selection –
StructuredOutputManager.grammar_initcreates the appropriate backend (XgrammarBackend,GuidanceBackend,OutlinesBackend, orLMFormatEnforcerBackend) based on the_backendparameter, defaulting to "auto". -
Grammar Compilation – The backend compiles the constraint into a deterministic FSM using
compile_grammar(request_type, grammar_spec). For example,XgrammarBackendinvllm/v1/structured_output/backend_xgrammar.pyhandles JSON schemas, regex, and custom grammars. -
Bitmask Enforcement –
StructuredOutputManager.grammar_bitmaskgenerates a token-level mask that GPU workers apply viaapply_grammar_bitmask()invllm/v1/worker/gpu/model_runner.pyto mask out invalid tokens during sampling.
Types of Structured Output Constraints
vLLM supports five constraint categories through either the structured_outputs parameter or the response_format field.
JSON Schema Validation
Use response_format with type json_schema to enforce JSON output matching a specific schema. The schema is processed in vllm/entrypoints/openai/chat_completion/protocol.py and compiled by XgrammarBackend by default.
response_format = {
"type": "json_schema",
"json_schema": {
"name": "car-description",
"schema": {
"type": "object",
"properties": {
"brand": {"type": "string"},
"model": {"type": "string"},
"car_type": {"type": "string", "enum": ["SEDAN", "SUV", "TRUCK", "COUPE"]}
},
"required": ["brand", "model", "car_type"]
}
}
}
Regular Expression Constraints
Force outputs to match regex patterns by passing a regex string in the structured_outputs field via extra_body. The regex is interpreted in vllm/v1/structured_output/request.py.
Custom Grammar Definitions
Define formal grammars using EBNF-like syntax in the grammar field. The grammar string is compiled into an FSM by XgrammarBackend in vllm/v1/structured_output/backend_xgrammar.py.
Choice Constraints
Restrict output to specific strings using the choice field, which accepts a list of allowed values.
Structural Tags
Use the structural_tag field in structured_outputs to constrain outputs according to predefined structural templates.
Practical Implementation Examples
Enforcing JSON Schema via OpenAI-Compatible Client
import asyncio
import openai
async def generate_json():
client = openai.AsyncOpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY"
)
model = (await client.models.list()).data[0].id
resp = await client.chat.completions.create(
model=model,
max_tokens=256,
messages=[{
"role": "user",
"content": "Generate a JSON with brand, model, and car_type of the most iconic 90s car"
}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "car",
"schema": {
"type": "object",
"properties": {
"brand": {"type": "string"},
"model": {"type": "string"},
"car_type": {"type": "string", "enum": ["SEDAN", "COUPE"]}
},
"required": ["brand", "model", "car_type"]
}
}
}
)
print(resp.choices[0].message.content)
asyncio.run(generate_json())
Under the hood, SamplingParams._validate_structured_outputs in vllm/sampling_params.py validates the schema and initializes the backend, which creates the constraining bitmask during generation.
Applying Regex Constraints
await client.chat.completions.create(
model=model,
max_tokens=64,
messages=[{"role": "user", "content": "Generate an email for Alan Turing"}],
extra_body={
"structured_outputs": {
"regex": r"[a-z0-9.]{1,20}@\w{6,10}\.com\n"
}
}
)
Defining Custom SQL Grammars
sql_grammar = """
root ::= select_statement
select_statement ::= "SELECT " column " FROM " table " WHERE " condition
column ::= "col_1 " | "col_2 "
table ::= "table_1 " | "table_2 "
condition ::= column "= " number
number ::= "1 " | "2 "
"""
await client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Write a SQL query selecting col_1 from table_1 where col_1 = 1"}],
extra_body={"structured_outputs": {"grammar": sql_grammar}}
)
Running the Complete Demo
Test all constraint types using the built-in example script at examples/online_serving/structured_outputs/structured_outputs.py:
python -m vllm.examples.online_serving.structured_outputs.structured_outputs \
--constraint regex json grammar choice structural_tag \
--stream
Backend Selection and Configuration
vLLM defaults to xgrammar for optimal performance, but supports multiple backends selectable via the _backend parameter in structured_outputs:
- XgrammarBackend: Default, handles JSON schemas, regex, and grammars efficiently in
vllm/v1/structured_output/backend_xgrammar.py - GuidanceBackend: Alternative using the
guidancelibrary (vllm/v1/structured_output/backend_guidance.py) - OutlinesBackend: Supports advanced parsing scenarios (
vllm/v1/structured_output/backend_outlines.py) - LMFormatEnforcerBackend: Additional compatibility option
The backend compiles constraints into token-level FSMs that StructuredOutputManager.grammar_bitmask converts into GPU-compatible masks applied during sampling.
Integration with Reasoning Models
When using reasoning models, StructuredOutputManager.should_fill_bitmask determines whether to apply constraints during the reasoning phase or only to the final output. This allows the model to generate unstructured reasoning tokens while ensuring the final answer adheres to your structured format.
Summary
- Structured outputs in vLLM constrain generations to match JSON schemas, regex patterns, or custom grammars at the token level using GPU-accelerated bitmasks.
- The system uses
StructuredOutputManagerinvllm/v1/structured_output/__init__.pyto orchestrate backend selection, grammar compilation, and bitmask generation. - Declare constraints via either
response_format(JSON schema) orextra_body["structured_outputs"](regex, grammar, choice). - Available backends include
xgrammar(default),guidance,outlines, andlm-format-enforcer, each implementingcompile_grammar()to build enforcement FSMs. - GPU workers in
vllm/v1/worker/gpu/model_runner.pyapply token-level bitmasks during each decoding step to guarantee output validity.
Frequently Asked Questions
What backends does vLLM support for structured outputs?
vLLM supports four backends: xgrammar (default), guidance, outlines, and lm-format-enforcer. The backend is selected via the _backend field in structured_outputs or defaults to "auto", which typically selects XgrammarBackend for its performance and broad format support as implemented in vllm/v1/structured_output/backend_xgrammar.py.
How does vLLM enforce structured outputs during generation?
The engine compiles constraints into finite state machines using the selected backend's compile_grammar() method. StructuredOutputManager.grammar_bitmask() creates a token-level mask that GPU workers apply via apply_grammar_bitmask() in vllm/v1/worker/gpu/model_runner.py before each token is sampled, effectively masking out any token that would violate the constraint.
Can I use structured outputs with reasoning models?
Yes. When a reasoning parser is configured, StructuredOutputManager.should_fill_bitmask() controls whether constraints apply during the reasoning phase or only to the final output. This allows the model to generate unstructured reasoning tokens while ensuring the final answer adheres to your schema or grammar.
Is there a performance penalty when using structured outputs?
There is minimal overhead because constraints are compiled once per request into efficient FSMs, and bitmask application occurs on the GPU during the existing sampling step. The xgrammar backend is optimized for this workflow and is generally faster than post-processing or rejection sampling approaches.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →