Forward Dynamics, Inverse Dynamics, and Policy Action Modes in Cosmos 3: Architecture and API Differences

Forward dynamics generates future video from planned actions, inverse dynamics extracts action vectors from observed video, and policy mode predicts next actions from visual context—each leveraging distinct decoder heads within Cosmos 3's Mixture-of-Transformers architecture.

Cosmos 3 is NVIDIA's open-source world model that provides three distinct action-modeling surfaces for robotics and autonomous systems. According to the NVIDIA/cosmos repository, these modes differ in their conditioning inputs, decoder architectures, and API endpoints while sharing the same underlying transformer backbone, enabling everything from motion preview to real-time control.

Overview of the Three Action Modes in Cosmos 3

Cosmos 3 structures its action-modeling capabilities into three distinct surfaces, each documented in the repository's main README.md (lines 66-68). The following table summarizes their core differences:

Mode Prediction Task Primary Output API Endpoint Use Case
Forward dynamics Future video given current state + desired action Video only (MP4) POST /v1/videos/sync (synchronous) Simulating robot motion or camera trajectories
Inverse dynamics Action vector given target video or before/after frames Action chunk (JSON array) POST /v1/videos (asynchronous) Recovering control commands from recorded motion
Policy Next action given visual context + language prompt Action chunk + optional video POST /v1/videos (asynchronous) Real-time closed-loop robot control

All three modes utilize the same Mixture-of-Transformers backbone, but differ in which decoder head activates and how the model conditions on inputs.

Forward Dynamics: Simulating Visual Futures

Forward dynamics rolls the world model forward to generate a visual preview of planned motion. In cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb, this mode activates the diffusion decoder on a concatenated stream of video and action tokens, producing only visual tokens without action outputs.

Use this mode when you need a quick video preview of a planned maneuver before executing it on hardware. The endpoint returns the generated video immediately (synchronously).

from diffusers import Cosmos3OmniPipeline
import torch

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

prompt = "A mobile robot drives forward and picks up a box."

# Forward-dynamics mode: generates only video

result = pipe(
    prompt=prompt,
    action_mode="forward",
    num_inference_steps=50,
)

with open("fd_output.mp4", "wb") as f:
    f.write(result.video)

As implemented in the cookbook at cookbooks/cosmos3/generator/action/run_fd_with_cosmos_framework.ipynb, this pipeline configuration uses the diffusion decoder exclusively, making it optimal for high-fidelity visual simulation.

Inverse Dynamics: Recovering Control Commands

Inverse dynamics solves the opposite problem: given a target video showing a motion, it infers the action vector (e.g., joint angles, camera 9-D pose) that produced that transition. The model runs a causal decoder that predicts action tokens given visual context, returning a JSON array rather than pixels.

This mode operates asynchronously via POST /v1/videos, returning the inferred action data when the job completes. Use it for imitation learning or recovering control commands from human demonstrations.

from diffusers import Cosmos3OmniPipeline
import torch
import json

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

video_url = "https://example.com/robot_motion.mp4"

result = pipe(
    video_url=video_url,
    action_mode="inverse",
    num_inference_steps=50,
)

# result.action contains the inferred action array

print("Inferred action:", json.dumps(result.action, indent=2))

The reference implementation in cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb demonstrates this pattern, showing how the causal decoder extracts actionable control sequences from visual observations.

Policy Action Mode: Real-Time Decision Making

Policy mode predicts the next action a robot policy would take given the current visual context and optionally a language prompt. Like inverse dynamics, it uses the causal decoder for action prediction, but unlike inverse dynamics, it focuses on forward-looking decision making rather than retroactive analysis.

This mode supports optional video generation: if you set generate_video=True, the inferred actions feed back into the diffusion decoder to render a visual rollout. This makes it suitable for closed-loop control with visual feedback.

from diffusers import Cosmos3OmniPipeline
import torch

pipe = Cosmos3OmniPipeline.from_pretrained(
    "nvidia/Cosmos3-Nano",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

prompt = "A robot arm is holding a cup. Drop the cup onto the table."

result = pipe(
    prompt=prompt,
    action_mode="policy",
    generate_video=True,  # Optional visual rollout

    num_inference_steps=50,
)

print("Predicted action:", result.action)
with open("policy_output.mp4", "wb") as f:
    f.write(result.video)

Documentation in cookbooks/cosmos3/generator/action/run_policy_with_cosmos_framework.md shows this dual-output capability, where the same call returns both the action chunk for immediate execution and an optional video for verification.

Architectural Implementation: Diffusion vs. Causal Decoders

As documented in README.md (lines 348-350), the three modes diverge at the decoder level while sharing the same tokenizer and transformer backbone:

  • Forward dynamics activates the diffusion decoder, treating the problem as a noise-to-data generation task on the concatenated video-action token stream.
  • Inverse dynamics and policy activate the causal decoder, autoregressively predicting the next action token(s) conditioned on visual context.

This architectural distinction explains why forward dynamics cannot return actions (it lacks the causal decoder head) and why inverse/policy modes return structured JSON rather than raw pixels by default.

Summary

  • Forward dynamics uses the diffusion decoder for video-only synthesis via the synchronous POST /v1/videos/sync endpoint, ideal for motion preview.
  • Inverse dynamics employs the causal decoder to extract action chunks (joint angles, poses) from observed video via the asynchronous POST /v1/videos endpoint.
  • Policy mode combines causal action prediction with optional diffusion-based video generation, supporting real-time closed-loop control.
  • All three modes share the Mixture-of-Transformers backbone but activate different decoder heads based on the action_mode parameter.

Frequently Asked Questions

What is the difference between forward and inverse dynamics in Cosmos 3?

Forward dynamics predicts future video frames given a current state and desired action sequence, returning only visual output. Inverse dynamics takes a target video and infers the action vector that caused the observed motion, returning a JSON array of control commands. According to the Cosmos 3 source code, forward uses the diffusion decoder while inverse uses the causal decoder.

When should I use policy mode instead of inverse dynamics?

Use policy mode for real-time decision making where the model predicts the next action from the current visual context and language prompts. Use inverse dynamics when analyzing existing video recordings to recover the control commands that produced a specific motion. Policy mode is forward-looking and designed for closed-loop control, while inverse dynamics is backward-looking and designed for imitation learning.

Can I generate video with inverse dynamics or policy mode?

Yes, but only indirectly. While inverse dynamics returns only action data by default, you can feed the inferred actions back into the diffusion decoder to generate a video. Policy mode offers a convenience flag (generate_video=True) that performs this internally, returning both the action chunk and the visual rollout in a single API call, as shown in run_policy_with_cosmos_framework.md.

Which API endpoint should I use for real-time robot control?

Use the asynchronous POST /v1/videos endpoint with action_mode="policy". While forward dynamics uses the synchronous endpoint for immediate video generation, policy mode requires the asynchronous endpoint because it returns structured action data (JSON) rather than raw video streams. For true real-time applications, parse the result.action array immediately upon job completion to command the robot.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →