# How to Implement and Disable Guardrails for Content Filtering in Cosmos 3 Generation

> Implement and disable Cosmos 3 guardrails for content filtering using the cosmos_guardrail package. Control safety features per-request or server-wide for generated media.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-14

---

**Cosmos 3 ships with built-in safety guardrails that automatically screen prompts for disallowed content and blur faces in generated media, which you can disable per-request or server-wide via the `cosmos_guardrail` package configuration.**

The **NVIDIA Cosmos** repository provides a comprehensive video generation platform that includes an optional safety system called *guardrails*. This system is implemented in a separate `cosmos_guardrail` package and integrated into the generation pipelines (Diffusers and vLLM-Omni) to filter prompts and anonymize faces. Understanding how to control these guardrails is essential for developers who need flexibility in content filtering policies or want to optimize inference performance by skipping safety checks.

## What Are Cosmos 3 Guardrails?

Guardrails in Cosmos 3 are a two-component safety system designed to prevent the generation of harmful content and protect privacy. According to the NVIDIA Cosmos source code, the system consists of:

- **Text Filter Model**: Screens input prompts for disallowed content before generation begins
- **Vision Model**: Detects and blurs faces in the generated video output

These models are packaged in the `cosmos_guardrail` dependency and loaded automatically when you instantiate a Cosmos 3 generator pipeline. The guardrails are enabled by default in both the Diffusers pipeline and the vLLM-Omni server, though they can be overridden at multiple levels.

## Installing the Guardrail Dependencies

Before you can use or disable guardrails, you must install the `cosmos_guardrail` package alongside the main Cosmos 3 dependencies. As documented in [[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)](https://github.com/NVIDIA/cosmos/blob/main/README.md#L219-L225), include the guardrail package in your installation:

```bash
pip install cosmos-guardrail

```

This package contains the safety models that the pipelines reference at runtime. Without this package, the generation pipelines may fail to load or operate in a degraded mode when guardrail checks are requested.

## How to Disable Guardrails in Cosmos 3

The NVIDIA Cosmos repository provides three distinct methods for disabling content filtering, depending on whether you need to disable safety checks for a single request, a specific pipeline instance, or an entire server deployment.

### Per-Request Disabling (vLLM-Omni)

For the vLLM-Omni server interface, you can disable guardrails for individual API calls by setting the **`guardrails`** parameter to `false` inside the `extra_params` JSON object. This override is documented in the vLLM-Omni cookbook at [[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)](https://github.com/NVIDIA/cosmos/blob/main/README.md#L94-L100).

The following curl command demonstrates how to generate video without safety filtering for a single request:

```bash
curl -sS -X POST http://localhost:8000/v1/videos/sync \
  --form-string "prompt=A robot paints a mural on a wall." \
  --form-string 'extra_params={"guardrails":false,"use_resolution_template":false,"use_duration_template":false}' \
  -o cosmos3_no_guardrails.mp4

```

Setting `"guardrails":false` instructs the server to skip both prompt filtering and face-blurring for this specific generation, while leaving the guardrail models loaded in memory for subsequent requests.

### Pipeline-Level Disabling (Diffusers)

When using the Diffusers-based `Cosmos3OmniPipeline`, you can disable guardrails for a specific pipeline instance. The pipeline exposes a setter method that controls safety checks:

```python
import torch
from diffusers import Cosmos3OmniPipeline

# Load the model (guardrails load by default)

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

# Disable guardrails for this pipeline instance

pipe.set_guardrails(False)

# Generate without content filtering

result = pipe(
    prompt="A futuristic cityscape at sunset.",
    num_frames=189,
    height=720,
    width=1280,
    fps=24,
    num_inference_steps=35,
    guidance_scale=6.0,
    generator=torch.Generator(device="cuda").manual_seed(1234),
)

```

Note that `set_guardrails(False)` disables safety checks for all subsequent calls using that pipeline instance, but the guardrail models remain loaded in GPU memory.

### Server-Wide Disabling (Deployment Config)

To completely eliminate guardrail overhead and prevent the safety models from loading entirely, you must provide a **deployment configuration YAML** file when starting the vLLM-Omni server. As shown in [[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)](https://github.com/NVIDIA/cosmos/blob/main/README.md#L103-L118), set `guardrails: false` and optionally `offload_guardrail_models: false` in the model configuration:

Create a file named [`no_guardrails.yaml`](https://github.com/NVIDIA/cosmos/blob/main/no_guardrails.yaml):

```yaml
async_chunk: false
stages:
  - stage_id: 0
    max_num_seqs: 1
    enforce_eager: true
    trust_remote_code: true
    model_class_name: Cosmos3OmniDiffusersPipeline
    model_config:
      guardrails: false          # Disables guardrails globally

      offload_guardrail_models: false

```

Start the server with this configuration:

```bash
vllm serve nvidia/Cosmos3-Nano \
  --omni \
  --model-class-name Cosmos3OmniDiffusersPipeline \
  --deploy-config no_guardrails.yaml \
  --port 8000 \
  --init-timeout 1800

```

With this server-wide configuration, the `cosmos_guardrail` models are never loaded into memory, eliminating both the GPU memory overhead and the inference latency associated with safety checks.

## Key Implementation Files

When working with Cosmos 3 guardrails, these specific files in the NVIDIA Cosmos repository contain the authoritative implementation details:

- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** (lines 219-225): Documents the `cosmos_guardrail` package installation requirements
- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** (lines 94-100): Contains the `extra_params` JSON schema for per-request guardrail control
- **[`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md)** (lines 103-118): Provides the deployment configuration YAML structure for server-wide disabling
- **[`cookbooks/cosmos3/generator/transfer/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/transfer/README.md)**: Defines the API contract for vLLM-Omni guardrail parameters

## Summary

- **Cosmos 3 guardrails** are implemented in the `cosmos_guardrail` package and provide automated prompt filtering and face blurring.
- **Install** the safety system via `pip install cosmos-guardrail` before running generation pipelines.
- **Disable per-request** in vLLM-Omni by setting `"guardrails": false` in the `extra_params` JSON payload.
- **Disable per-pipeline** in Diffusers by calling `pipe.set_guardrails(False)` on the pipeline instance.
- **Disable server-wide** by providing a deployment config YAML with `guardrails: false` in the `model_config` section, which prevents the guardrail models from loading entirely.

## Frequently Asked Questions

### What happens if I disable guardrails but the `cosmos_guardrail` package is not installed?

The generation pipeline will attempt to load the guardrail models and fail with an import error or runtime exception, as the pipelines expect the safety components to be available even when disabled at the configuration level. Always install `cosmos_guardrail` before attempting to use Cosmos 3 generation features.

### Does disabling guardrails improve generation speed?

Yes. When you disable guardrails server-wide via the deployment configuration, the safety models are never loaded into GPU memory, eliminating the inference overhead associated with prompt classification and face detection. Per-request disabling skips the safety inference steps but maintains the memory overhead since the models remain loaded.

### Can I disable only face blurring but keep prompt filtering?

The current Cosmos 3 implementation exposes guardrails as a single boolean flag that controls both the text filter and vision filter simultaneously. You cannot selectively disable only face blurring or only prompt filtering through the public API; both safety mechanisms are toggled together via the `guardrails` parameter.

### Is there a performance penalty for keeping guardrails enabled?

Yes, there is a small latency penalty for each generation request when guardrails are enabled, as the system must run the text classification model on the prompt and potentially process the vision model on generated frames. Additionally, the guardrail models consume GPU memory that could otherwise be allocated to the diffusion process.