Cosmos 3 Action Policy Models for Robot Manipulation: Architecture and Deployment
Cosmos 3 action policy models convert vision-language prompts into executable robot trajectories, enabling the Cosmos3-Nano-Policy-DROID model to generate 10-dimensional DROID arm joint commands from scene images and text instructions.
Cosmos 3 introduces a specialized family of action-policy models that transform the world-model framework into a robot controller. According to the NVIDIA Cosmos repository, these models leverage a unified Mixture-of-Transformers architecture to bridge high-level language instructions with low-level motor execution. The flagship Cosmos3-Nano-Policy-DROID model represents a 16-billion-parameter vision-language system specifically trained for manipulation tasks.
Mixture-of-Transformers Architecture for Robot Control
The policy models inherit the same Mixture-of-Transformers (MoT) architecture used throughout Cosmos 3, as listed in README.md lines 80-86. This design enables the model to process multimodal inputs and generate coherent action trajectories rather than purely visual content.
Core Components in Policy Mode
When operating in policy mode, the model functions as a generator conditioned on visual and textual inputs. The architecture consists of four key components:
- Autoregressive transformer: Processes concatenated text-vision token streams to understand task instructions and scene context
- Diffusion transformer (DM): Denoises latent action sequences while producing visualization frames, yielding coherent 10-dimensional DROID arm joint commands
- Multimodal attention layers: Share weights across vision, text, and action tokens, enabling joint reasoning about visual state and motor commands
- 3-D multi-dimensional RoPE: Encodes spatial-temporal relationships through 3-D rotary position embeddings, ensuring generated actions respect the geometry of the scene
Safety Guardrails
Built-in guardrails filter harmful prompts and blur faces in generated videos. These safety checks operate by default during inference but can be disabled per-request if required for specific research applications.
Cosmos3-Nano-Policy-DROID Specifications
The flagship Cosmos3-Nano-Policy-DROID model contains approximately 16 billion parameters and is highlighted in the repository's model table as a "Vision-language robot policy for DROID manipulation and control." Unlike standard video generation models, this system predicts action trajectories—specifically sequences of 10-dimensional DROID arm joint commands—while simultaneously producing short rollout videos that visualize the predicted motion.
The model operates as a conditional generator: it takes an initial scene image and a textual instruction (e.g., "Pick up the red block and place it on the blue platform"), then denoises a latent representation of the action trajectory through the diffusion transformer pipeline. The action_chunk_size parameter (typically set to 20) controls the temporal horizon of predicted motion.
Deployment and Integration
Developers can deploy Cosmos 3 action policy models through two primary interfaces: the Cosmos Framework policy server or the vLLM-Omni API.
Running the Policy Server
The policy server runs as a lightweight Python script within the Cosmos Framework Docker container. According to cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md lines 57-64, launch the server with:
python -m cosmos_framework.scripts.action_policy_server_robolab \
--port 8000
Once the server is active on port 8000, RoboLab clients can stream actions to simulated or real robots.
RoboLab Client Integration
Connect a RoboLab task to the policy server using the client script described in run_policy_with_cosmos_framework.md lines 64-90:
python policies/cosmos3/run.py --task BananaInBowlTask
This setup enables real-time action streaming from the Cosmos 3 policy model to robot hardware.
vLLM-Omni API Access
For programmatic access, the policy surface exposes an OpenAI-compatible endpoint through vLLM-Omni. As documented in cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md lines 84-102, send requests using curl:
curl -sS -X POST http://localhost:8000/v1/videos/sync \
--form-string "prompt=Pick up the red block and place it on the blue platform." \
--form-string "size=1280x720" \
--form-string "num_frames=189" \
--form-string "fps=24" \
--form-string "num_inference_steps=35" \
--form-string "guidance_scale=6.0" \
--form-string 'extra_params={"action_mode":"policy","domain_name":"droid","raw_action_dim":10,"action_chunk_size":20}' \
-o policy_output.mp4
The extra_params JSON payload configures the policy mode, selects the DROID embodiment, and specifies the raw action dimensionality and chunk size.
Python Client Example
Alternatively, use the OpenAI SDK to interact with the policy endpoint:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
resp = client.videos.create(
model="nvidia/Cosmos3-Nano",
prompt="Pick up the red block and place it on the blue platform.",
size="1280x720",
num_frames=189,
fps=24,
num_inference_steps=35,
guidance_scale=6.0,
extra_params={"action_mode": "policy", "domain_name": "droid",
"raw_action_dim": 10, "action_chunk_size": 20},
)
with open("policy_output.mp4", "wb") as f:
f.write(resp.video)
The response contains an MP4 video visualizing the predicted motion. When using the asynchronous /v1/videos endpoint, a companion JSON file provides the raw action sequence.
Performance Characteristics
Inference benchmarks documented in inference_benchmarks.md lines 352-356 provide latency and throughput expectations for production deployments. The action-policy mode maintains competitive performance when serving through vLLM-Omni, supporting scalable robot control pipelines.
Summary
- Cosmos3-Nano-Policy-DROID is a 16-billion-parameter vision-language model that generates robot actions from text and image inputs
- The Mixture-of-Transformers architecture processes multimodal tokens and denoises latent action trajectories using a diffusion transformer
- Action outputs follow the DROID embodiment format with 10-dimensional joint commands and configurable action chunk sizes
- Deployment options include the Cosmos Framework policy server for RoboLab integration and vLLM-Omni for REST API access
- Built-in guardrails provide safety filtering and face blurring capabilities
Frequently Asked Questions
What is the parameter count of Cosmos 3 policy models?
The flagship policy model, Cosmos3-Nano-Policy-DROID, contains approximately 16 billion parameters. This makes it the largest model in the Cosmos 3 action-policy family while maintaining efficient inference through the Mixture-of-Transformers architecture.
How does the policy model generate robot actions?
The model uses a diffusion transformer to denoise a latent representation of the action trajectory. Conditioned on the initial scene image and text instruction, it predicts sequences of 10-dimensional DROID arm joint commands while simultaneously generating a visualization video that shows the predicted robot motion.
Can I run the policy model without the Cosmos Framework?
Yes. While the Cosmos Framework provides the recommended policy server (cosmos_framework.scripts.action_policy_server_robolab), you can also access the model through the vLLM-Omni OpenAI-compatible API. This allows direct HTTP requests using standard tools like curl or the OpenAI Python SDK, making it possible to integrate with custom robotics stacks outside the official framework.
What safety mechanisms protect against harmful outputs?
Cosmos 3 policy models include built-in guardrails that filter harmful prompts and automatically blur faces in generated videos. These safety checks operate by default during inference but can be disabled per-request when necessary for specific research applications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →