# How to Use Structured Outputs (JSON Schema, Regex, Grammar) in vLLM

> Learn how vLLM enforces structured outputs using JSON schema, regex, or grammars with token-level bitmasks for precise text generation.

- Repository: [vLLM/vllm](https://github.com/vllm-project/vllm)
- Tags: how-to-guide
- Published: 2026-03-03

---

**vLLM supports structured output constraints that force language models to generate text matching JSON schemas, regular expressions, or custom grammars by applying token-level bitmasks during decoding.**

vLLM enables precise control over model generations through **structured outputs in vLLM**, allowing you to constrain responses to validate against JSON schemas, match regular expressions, or follow custom grammars. According to the vllm-project/vllm source code, this feature works by compiling constraints into token-level finite state machines (FSMs) that the inference engine enforces at every generation step via `StructuredOutputManager` in [`vllm/v1/structured_output/__init__.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/__init__.py).

## How Structured Outputs Work in vLLM

The structured output pipeline operates in four stages as implemented in [`vllm/v1/structured_output/__init__.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/__init__.py):

1. **Request Parsing** – The engine extracts the `structured_outputs` field from OpenAI-compatible requests in [`vllm/v1/structured_output/request.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/request.py), or the `response_format` field for JSON schemas in [`vllm/entrypoints/openai/chat_completion/protocol.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/chat_completion/protocol.py).

2. **Backend Selection** – `StructuredOutputManager.grammar_init` creates the appropriate backend (`XgrammarBackend`, `GuidanceBackend`, `OutlinesBackend`, or `LMFormatEnforcerBackend`) based on the `_backend` parameter, defaulting to "auto".

3. **Grammar Compilation** – The backend compiles the constraint into a deterministic FSM using `compile_grammar(request_type, grammar_spec)`. For example, `XgrammarBackend` in [`vllm/v1/structured_output/backend_xgrammar.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/backend_xgrammar.py) handles JSON schemas, regex, and custom grammars.

4. **Bitmask Enforcement** – `StructuredOutputManager.grammar_bitmask` generates a token-level mask that GPU workers apply via `apply_grammar_bitmask()` in [`vllm/v1/worker/gpu/model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu/model_runner.py) to mask out invalid tokens during sampling.

## Types of Structured Output Constraints

vLLM supports five constraint categories through either the `structured_outputs` parameter or the `response_format` field.

### JSON Schema Validation

Use `response_format` with type `json_schema` to enforce JSON output matching a specific schema. The schema is processed in [`vllm/entrypoints/openai/chat_completion/protocol.py`](https://github.com/vllm-project/vllm/blob/main/vllm/entrypoints/openai/chat_completion/protocol.py) and compiled by `XgrammarBackend` by default.

```python
response_format = {
    "type": "json_schema",
    "json_schema": {
        "name": "car-description",
        "schema": {
            "type": "object",
            "properties": {
                "brand": {"type": "string"},
                "model": {"type": "string"},
                "car_type": {"type": "string", "enum": ["SEDAN", "SUV", "TRUCK", "COUPE"]}
            },
            "required": ["brand", "model", "car_type"]
        }
    }
}

```

### Regular Expression Constraints

Force outputs to match regex patterns by passing a `regex` string in the `structured_outputs` field via `extra_body`. The regex is interpreted in [`vllm/v1/structured_output/request.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/request.py).

### Custom Grammar Definitions

Define formal grammars using EBNF-like syntax in the `grammar` field. The grammar string is compiled into an FSM by `XgrammarBackend` in [`vllm/v1/structured_output/backend_xgrammar.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/backend_xgrammar.py).

### Choice Constraints

Restrict output to specific strings using the `choice` field, which accepts a list of allowed values.

### Structural Tags

Use the `structural_tag` field in `structured_outputs` to constrain outputs according to predefined structural templates.

## Practical Implementation Examples

### Enforcing JSON Schema via OpenAI-Compatible Client

```python
import asyncio
import openai

async def generate_json():
    client = openai.AsyncOpenAI(
        base_url="http://localhost:8000/v1",
        api_key="EMPTY"
    )
    model = (await client.models.list()).data[0].id
    
    resp = await client.chat.completions.create(
        model=model,
        max_tokens=256,
        messages=[{
            "role": "user",
            "content": "Generate a JSON with brand, model, and car_type of the most iconic 90s car"
        }],
        response_format={
            "type": "json_schema",
            "json_schema": {
                "name": "car",
                "schema": {
                    "type": "object",
                    "properties": {
                        "brand": {"type": "string"},
                        "model": {"type": "string"},
                        "car_type": {"type": "string", "enum": ["SEDAN", "COUPE"]}
                    },
                    "required": ["brand", "model", "car_type"]
                }
            }
        }
    )
    print(resp.choices[0].message.content)

asyncio.run(generate_json())

```

Under the hood, `SamplingParams._validate_structured_outputs` in [`vllm/sampling_params.py`](https://github.com/vllm-project/vllm/blob/main/vllm/sampling_params.py) validates the schema and initializes the backend, which creates the constraining bitmask during generation.

### Applying Regex Constraints

```python
await client.chat.completions.create(
    model=model,
    max_tokens=64,
    messages=[{"role": "user", "content": "Generate an email for Alan Turing"}],
    extra_body={
        "structured_outputs": {
            "regex": r"[a-z0-9.]{1,20}@\w{6,10}\.com\n"
        }
    }
)

```

### Defining Custom SQL Grammars

```python
sql_grammar = """
root ::= select_statement
select_statement ::= "SELECT " column " FROM " table " WHERE " condition
column ::= "col_1 " | "col_2 "
table ::= "table_1 " | "table_2 "
condition ::= column "= " number
number ::= "1 " | "2 "
"""

await client.chat.completions.create(
    model=model,
    messages=[{"role": "user", "content": "Write a SQL query selecting col_1 from table_1 where col_1 = 1"}],
    extra_body={"structured_outputs": {"grammar": sql_grammar}}
)

```

### Running the Complete Demo

Test all constraint types using the built-in example script at [`examples/online_serving/structured_outputs/structured_outputs.py`](https://github.com/vllm-project/vllm/blob/main/examples/online_serving/structured_outputs/structured_outputs.py):

```bash
python -m vllm.examples.online_serving.structured_outputs.structured_outputs \
    --constraint regex json grammar choice structural_tag \
    --stream

```

## Backend Selection and Configuration

vLLM defaults to `xgrammar` for optimal performance, but supports multiple backends selectable via the `_backend` parameter in `structured_outputs`:

- **XgrammarBackend**: Default, handles JSON schemas, regex, and grammars efficiently in [`vllm/v1/structured_output/backend_xgrammar.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/backend_xgrammar.py)
- **GuidanceBackend**: Alternative using the `guidance` library ([`vllm/v1/structured_output/backend_guidance.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/backend_guidance.py))
- **OutlinesBackend**: Supports advanced parsing scenarios ([`vllm/v1/structured_output/backend_outlines.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/backend_outlines.py))
- **LMFormatEnforcerBackend**: Additional compatibility option

The backend compiles constraints into token-level FSMs that `StructuredOutputManager.grammar_bitmask` converts into GPU-compatible masks applied during sampling.

## Integration with Reasoning Models

When using reasoning models, `StructuredOutputManager.should_fill_bitmask` determines whether to apply constraints during the reasoning phase or only to the final output. This allows the model to generate unstructured reasoning tokens while ensuring the final answer adheres to your structured format.

## Summary

- **Structured outputs in vLLM** constrain generations to match JSON schemas, regex patterns, or custom grammars at the token level using GPU-accelerated bitmasks.
- The system uses `StructuredOutputManager` in [`vllm/v1/structured_output/__init__.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/__init__.py) to orchestrate backend selection, grammar compilation, and bitmask generation.
- Declare constraints via either `response_format` (JSON schema) or `extra_body["structured_outputs"]` (regex, grammar, choice).
- Available backends include `xgrammar` (default), `guidance`, `outlines`, and `lm-format-enforcer`, each implementing `compile_grammar()` to build enforcement FSMs.
- GPU workers in [`vllm/v1/worker/gpu/model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu/model_runner.py) apply token-level bitmasks during each decoding step to guarantee output validity.

## Frequently Asked Questions

### What backends does vLLM support for structured outputs?

vLLM supports four backends: `xgrammar` (default), `guidance`, `outlines`, and `lm-format-enforcer`. The backend is selected via the `_backend` field in `structured_outputs` or defaults to "auto", which typically selects `XgrammarBackend` for its performance and broad format support as implemented in [`vllm/v1/structured_output/backend_xgrammar.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/structured_output/backend_xgrammar.py).

### How does vLLM enforce structured outputs during generation?

The engine compiles constraints into finite state machines using the selected backend's `compile_grammar()` method. `StructuredOutputManager.grammar_bitmask()` creates a token-level mask that GPU workers apply via `apply_grammar_bitmask()` in [`vllm/v1/worker/gpu/model_runner.py`](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu/model_runner.py) before each token is sampled, effectively masking out any token that would violate the constraint.

### Can I use structured outputs with reasoning models?

Yes. When a reasoning parser is configured, `StructuredOutputManager.should_fill_bitmask()` controls whether constraints apply during the reasoning phase or only to the final output. This allows the model to generate unstructured reasoning tokens while ensuring the final answer adheres to your schema or grammar.

### Is there a performance penalty when using structured outputs?

There is minimal overhead because constraints are compiled once per request into efficient FSMs, and bitmask application occurs on the GPU during the existing sampling step. The `xgrammar` backend is optimized for this workflow and is generally faster than post-processing or rejection sampling approaches.