How to Use Structured Outputs (JSON Schema, Regex, Grammar) in vLLM

vLLM supports structured output constraints that force language models to generate text matching JSON schemas, regular expressions, or custom grammars by applying token-level bitmasks during decoding.

vLLM enables precise control over model generations through structured outputs in vLLM, allowing you to constrain responses to validate against JSON schemas, match regular expressions, or follow custom grammars. According to the vllm-project/vllm source code, this feature works by compiling constraints into token-level finite state machines (FSMs) that the inference engine enforces at every generation step via StructuredOutputManager in vllm/v1/structured_output/__init__.py.

How Structured Outputs Work in vLLM

The structured output pipeline operates in four stages as implemented in vllm/v1/structured_output/__init__.py:

  1. Request Parsing – The engine extracts the structured_outputs field from OpenAI-compatible requests in vllm/v1/structured_output/request.py, or the response_format field for JSON schemas in vllm/entrypoints/openai/chat_completion/protocol.py.

  2. Backend Selection – StructuredOutputManager.grammar_init creates the appropriate backend (XgrammarBackend, GuidanceBackend, OutlinesBackend, or LMFormatEnforcerBackend) based on the _backend parameter, defaulting to "auto".

  3. Grammar Compilation – The backend compiles the constraint into a deterministic FSM using compile_grammar(request_type, grammar_spec). For example, XgrammarBackend in vllm/v1/structured_output/backend_xgrammar.py handles JSON schemas, regex, and custom grammars.

  4. Bitmask Enforcement – StructuredOutputManager.grammar_bitmask generates a token-level mask that GPU workers apply via apply_grammar_bitmask() in vllm/v1/worker/gpu/model_runner.py to mask out invalid tokens during sampling.

Types of Structured Output Constraints

vLLM supports five constraint categories through either the structured_outputs parameter or the response_format field.

JSON Schema Validation

Use response_format with type json_schema to enforce JSON output matching a specific schema. The schema is processed in vllm/entrypoints/openai/chat_completion/protocol.py and compiled by XgrammarBackend by default.

response_format = {
    "type": "json_schema",
    "json_schema": {
        "name": "car-description",
        "schema": {
            "type": "object",
            "properties": {
                "brand": {"type": "string"},
                "model": {"type": "string"},
                "car_type": {"type": "string", "enum": ["SEDAN", "SUV", "TRUCK", "COUPE"]}
            },
            "required": ["brand", "model", "car_type"]
        }
    }
}

Regular Expression Constraints

Force outputs to match regex patterns by passing a regex string in the structured_outputs field via extra_body. The regex is interpreted in vllm/v1/structured_output/request.py.

Custom Grammar Definitions

Define formal grammars using EBNF-like syntax in the grammar field. The grammar string is compiled into an FSM by XgrammarBackend in vllm/v1/structured_output/backend_xgrammar.py.

Choice Constraints

Restrict output to specific strings using the choice field, which accepts a list of allowed values.

Structural Tags

Use the structural_tag field in structured_outputs to constrain outputs according to predefined structural templates.

Practical Implementation Examples

Enforcing JSON Schema via OpenAI-Compatible Client

import asyncio
import openai

async def generate_json():
    client = openai.AsyncOpenAI(
        base_url="http://localhost:8000/v1",
        api_key="EMPTY"
    )
    model = (await client.models.list()).data[0].id
    
    resp = await client.chat.completions.create(
        model=model,
        max_tokens=256,
        messages=[{
            "role": "user",
            "content": "Generate a JSON with brand, model, and car_type of the most iconic 90s car"
        }],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "car",
                "schema": {
                    "type": "object",
                    "properties": {
                        "brand": {"type": "string"},
                        "model": {"type": "string"},
                        "car_type": {"type": "string", "enum": ["SEDAN", "COUPE"]}
                    },
                    "required": ["brand", "model", "car_type"]
                }
            }
        }
    )
    print(resp.choices[0].message.content)

asyncio.run(generate_json())

Under the hood, SamplingParams._validate_structured_outputs in vllm/sampling_params.py validates the schema and initializes the backend, which creates the constraining bitmask during generation.

Applying Regex Constraints

await client.chat.completions.create(
    model=model,
    max_tokens=64,
    messages=[{"role": "user", "content": "Generate an email for Alan Turing"}],
    extra_body={
        "structured_outputs": {
            "regex": r"[a-z0-9.]{1,20}@\w{6,10}\.com\n"
        }
    }
)

Defining Custom SQL Grammars

sql_grammar = """
root ::= select_statement
select_statement ::= "SELECT " column " FROM " table " WHERE " condition
column ::= "col_1 " | "col_2 "
table ::= "table_1 " | "table_2 "
condition ::= column "= " number
number ::= "1 " | "2 "
"""

await client.chat.completions.create(
    model=model,
    messages=[{"role": "user", "content": "Write a SQL query selecting col_1 from table_1 where col_1 = 1"}],
    extra_body={"structured_outputs": {"grammar": sql_grammar}}
)

Running the Complete Demo

Test all constraint types using the built-in example script at examples/online_serving/structured_outputs/structured_outputs.py:

python -m vllm.examples.online_serving.structured_outputs.structured_outputs \
    --constraint regex json grammar choice structural_tag \
    --stream

Backend Selection and Configuration

vLLM defaults to xgrammar for optimal performance, but supports multiple backends selectable via the _backend parameter in structured_outputs:

The backend compiles constraints into token-level FSMs that StructuredOutputManager.grammar_bitmask converts into GPU-compatible masks applied during sampling.

Integration with Reasoning Models

When using reasoning models, StructuredOutputManager.should_fill_bitmask determines whether to apply constraints during the reasoning phase or only to the final output. This allows the model to generate unstructured reasoning tokens while ensuring the final answer adheres to your structured format.

Summary

  • Structured outputs in vLLM constrain generations to match JSON schemas, regex patterns, or custom grammars at the token level using GPU-accelerated bitmasks.
  • The system uses StructuredOutputManager in vllm/v1/structured_output/__init__.py to orchestrate backend selection, grammar compilation, and bitmask generation.
  • Declare constraints via either response_format (JSON schema) or extra_body["structured_outputs"] (regex, grammar, choice).
  • Available backends include xgrammar (default), guidance, outlines, and lm-format-enforcer, each implementing compile_grammar() to build enforcement FSMs.
  • GPU workers in vllm/v1/worker/gpu/model_runner.py apply token-level bitmasks during each decoding step to guarantee output validity.

Frequently Asked Questions

What backends does vLLM support for structured outputs?

vLLM supports four backends: xgrammar (default), guidance, outlines, and lm-format-enforcer. The backend is selected via the _backend field in structured_outputs or defaults to "auto", which typically selects XgrammarBackend for its performance and broad format support as implemented in vllm/v1/structured_output/backend_xgrammar.py.

How does vLLM enforce structured outputs during generation?

The engine compiles constraints into finite state machines using the selected backend's compile_grammar() method. StructuredOutputManager.grammar_bitmask() creates a token-level mask that GPU workers apply via apply_grammar_bitmask() in vllm/v1/worker/gpu/model_runner.py before each token is sampled, effectively masking out any token that would violate the constraint.

Can I use structured outputs with reasoning models?

Yes. When a reasoning parser is configured, StructuredOutputManager.should_fill_bitmask() controls whether constraints apply during the reasoning phase or only to the final output. This allows the model to generate unstructured reasoning tokens while ensuring the final answer adheres to your schema or grammar.

Is there a performance penalty when using structured outputs?

There is minimal overhead because constraints are compiled once per request into efficient FSMs, and bitmask application occurs on the GPU during the existing sampling step. The xgrammar backend is optimized for this workflow and is generally faster than post-processing or rejection sampling approaches.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →