# High-Performance Serving of Sana with SGLang and OpenAI-Compatible API

> Serve NVlabs Sana diffusion models at scale with SGLang. Enjoy unified runtime, command-line, Python SDK, and OpenAI-compatible API for high performance.

- Repository: [NVIDIA Research Projects/Sana](https://github.com/NVlabs/Sana)
- Tags: performance
- Published: 2026-05-19

---

**You can serve the NVlabs/Sana diffusion model at production scale using SGLang, which provides a unified runtime supporting command-line, Python SDK, and OpenAI-compatible HTTP server interfaces backed by chunked autoregressive inference and intelligent KV-caching.**

Sana is a diffusion-based video synthesis model developed by NVIDIA Labs that generates high-quality visual content through efficient transformer architectures. When deploying Sana for high-throughput applications, **high-performance serving with SGLang and an OpenAI-compatible API** eliminates boilerplate code while leveraging optimized caching mechanisms in the inference pipeline. This integration enables you to drop Sana into existing OpenAI client workflows or run it efficiently via CLI and SDK interfaces, all while maximizing GPU utilization through temporal chunking and memory offloading strategies.

## Three Methods to Serve Sana with SGLang

SGLang exposes Sana through three distinct entry points, each utilizing the same underlying `DiffGenerator` class and `SanaInferencePipeline` but catering to different deployment scenarios.

### Command-Line Interface for Rapid Prototyping

The `sglang generate` command provides the fastest path to testing Sana inference without writing code. When executed, SGLang loads the Diffusers-compatible checkpoint from HuggingFace (`Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers`), instantiates the `DiffGenerator` (located in `sglang.multimodal_gen`), and executes the diffusion process defined in [`diffusion/longsana/pipeline/sana_inference_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/pipeline/sana_inference_pipeline.py). The CLI automatically handles model downloading, scheduler initialization, and output serialization to the `outputs/` directory.

### Python SDK for Programmatic Control

For production applications requiring dynamic parameter tuning, import `DiffGenerator` from `sglang.multimodal_gen` to construct generation pipelines programmatically. The SDK method `DiffGenerator.from_pretrained()` loads the Sana architecture, while `generate()` accepts a `sampling_params_kwargs` dictionary containing prompts, dimensions, inference steps, and guidance scales. This approach provides direct access to the `SanaInferencePipeline` internals, including cache configuration and LoRA loading, without the overhead of HTTP transport.

### OpenAI-Compatible HTTP Server

The `sglang serve` command launches a FastAPI/uvicorn server exposing the `/v1/images/generations` endpoint, which accepts the same JSON payload schema as OpenAI's image generation API. Incoming requests are routed internally to `DiffGenerator.generate()`, with outputs serialized as base64-encoded JSON or raw pixel tensors. The server supports standard parameters including `prompt`, `size`, `num_inference_steps`, `guidance_scale`, and `seed`, enabling drop-in replacement for existing OpenAI client libraries.

## Performance Architecture and Optimizations

High-throughput serving relies on specific architectural features implemented in the Sana codebase that minimize memory bandwidth and maximize GPU parallelism.

### Chunked Autoregressive Inference

To generate long videos without exhausting VRAM, the pipeline implemented in [`diffusion/longsana/pipeline/sana_inference_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/pipeline/sana_inference_pipeline.py) splits the temporal dimension into blocks controlled by `num_frame_per_block`. The method `_create_autoregressive_segments()` manages the splitting logic, while `_initialize_kv_cache()` prepares cache structures for each segment. This chunked approach allows the model to process extended sequences that exceed native context limits by recomputing attention only for new segments while preserving cached states from previous chunks.

### KV-Cache Management with Cached Blocks

The Sana architecture utilizes `CachedCausalAttention` and `CachedGLUMBConvTemp` modules (defined in [`diffusion/model/nets/sana_blocks.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/model/nets/sana_blocks.py)) to store intermediate activations across generation steps. During inference, `_accumulate_kv_cache()` updates these cached blocks, while `num_cached_blocks` limits the maximum cache size to balance speed against memory consumption. This caching strategy eliminates redundant computation when processing chunked video segments or iterative denoising steps.

### Memory Optimization via CPU Offloading

For deployment on memory-constrained GPUs, SGLang supports aggressive offloading strategies through command-line flags. Specifying `--text-encoder-cpu-offload`, `--vae-cpu-offload`, or `--dit-cpu-offload` moves the corresponding modules to system RAM, streaming weights to GPU only during active computation. The `--pin-cpu-memory` flag optimizes this transfer by preventing CPU memory paging. These options, combined with the caching mechanisms, enable Sana to run on consumer-grade hardware while maintaining acceptable latency.

### Dynamic LoRA Adapter Loading

SGLang supports runtime style customization through `--lora-path`, which loads Low-Rank Adaptation weights and merges them on-the-fly into the base Sana model. This occurs during `DiffGenerator` initialization, allowing multiple specialized endpoints to share the same base checkpoint while serving different visual domains without full model replication.

## Implementation Examples

### Generating Content via CLI

Execute a single generation request from the terminal:

```bash
sglang generate \
    --model-path Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers \
    --prompt "a cyberpunk cat with a neon sign that says Sana" \
    --save-output

```

This invokes the full pipeline, automatically saving outputs to the `outputs/` folder.

### Programmatic Generation with Python

Control sampling parameters precisely using the SDK:

```python
from sglang.multimodal_gen import DiffGenerator

generator = DiffGenerator.from_pretrained(
    model_path="Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers",
    num_gpus=1,
)

image = generator.generate(
    sampling_params_kwargs=dict(
        prompt='a cyberpunk cat with a neon sign that says "Sana"',
        height=1024,
        width=1024,
        num_inference_steps=20,
        guidance_scale=4.5,
        seed=42,
        save_output=True,
        output_path="outputs/",
    )
)

```

The `DiffGenerator` constructs the same `SanaInferencePipeline` used by the CLI, exposing methods like `_initialize_cached_modules` for advanced cache tuning.

### Deploying the OpenAI-Compatible Server

Start the HTTP server with optimized memory settings:

```bash
sglang serve --model-path Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers \
    --host 0.0.0.0 --port 30000 \
    --dit-cpu-offload

```

### Client Integration Using Standard OpenAI Patterns

Once the server is running, consume it with any HTTP client:

```python
import requests

resp = requests.post(
    "http://127.0.0.1:30000/v1/images/generations",
    json={
        "prompt": 'a cyberpunk cat with a neon sign that says "Sana"',
        "size": "1024x1024",
        "num_inference_steps": 20,
        "guidance_scale": 4.5,
        "seed": 42,
        "response_format": "b64_json",
        "n": 1,
    },
)

result = resp.json()

# result['data'][0]['b64_json'] contains the base64-encoded PNG

```

## Summary

- **SGLang provides three interfaces** for Sana: CLI (`sglang generate`), Python SDK (`DiffGenerator`), and OpenAI-compatible HTTP server (`sglang serve`), all utilizing the core `SanaInferencePipeline`.
- **Chunked autoregressive inference** in [`diffusion/longsana/pipeline/sana_inference_pipeline.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/pipeline/sana_inference_pipeline.py) enables long-form video generation through `_create_autoregressive_segments` and `_accumulate_kv_cache`.
- **KV-cache optimization** via `CachedCausalAttention` and `CachedGLUMBConvTemp` blocks (defined in [`diffusion/model/nets/sana_blocks.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/model/nets/sana_blocks.py)) minimizes redundant computation across generation steps.
- **CPU offloading flags** (`--text-encoder-cpu-offload`, `--vae-cpu-offload`, `--dit-cpu-offload`) reduce VRAM requirements for deployment on constrained hardware.
- **LoRA support** allows runtime customization without model retraining through the `--lora-path` parameter.
- **OpenAI API compatibility** enables immediate integration with existing client libraries via the `/v1/images/generations` endpoint.

## Frequently Asked Questions

### What is the difference between using the SGLang CLI and the OpenAI-compatible API?

The **CLI** (`sglang generate`) runs inference in a single process and writes outputs directly to disk, ideal for batch processing or testing. The **OpenAI-compatible API** (`sglang serve`) maintains a persistent HTTP server that handles concurrent requests from multiple clients, making it suitable for production microservices architectures. Both ultimately invoke the same `DiffGenerator.generate()` method, but the API adds request queuing and JSON serialization overhead.

### How does chunked autoregressive inference improve generation performance?

Chunked autoregressive inference splits long video sequences into manageable `num_frame_per_block` segments processed sequentially in `SanaInferencePipeline._create_autoregressive_segments()`. By reusing cached attention states across chunks through `_accumulate_kv_cache()`, the pipeline avoids recomputing attention for already-generated frames, reducing computational complexity from quadratic to near-linear relative to sequence length. This enables generation of videos that would otherwise exceed GPU memory capacity.

### Can I use custom LoRA adapters with the SGLang server?

Yes. Launch the server with the `--lora-path` flag pointing to your adapter weights: `sglang serve --lora-path /path/to/adapter.safetensors`. SGLang merges these weights into the base Sana model at runtime, allowing you to serve multiple stylistic variations from a single base checkpoint by running separate server instances with different LoRA paths, or by reloading adapters between requests in SDK mode.

### Which scheduler implementations does SGLang use for Sana inference?

SGLang utilizes the scheduler implementations located in `diffusion/scheduler/` within the NVlabs/Sana repository, including LCM and DPM-Solver variants. These schedulers are wrapped by `DiffGenerator.generate()` and exposed through the `num_inference_steps` and `guidance_scale` parameters, ensuring that SGLang serving produces outputs identical to native Sana inference while benefiting from the runtime's optimized memory management.