# Running Inverse Dynamics Predictions for Autonomous Vehicle Ego-Motion with Cosmos 3

> Learn to run inverse dynamics predictions for autonomous vehicle ego-motion with Cosmos 3. Estimate 9-DOF vehicle pose deltas from video using PyTorch or REST APIs.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Cosmos 3 treats action as a first-class modality to estimate 9-DOF vehicle pose deltas from video, supporting both native PyTorch inference and OpenAI-compatible REST APIs for autonomous vehicle development.**

NVIDIA's Cosmos 3 is an omnimodal world-model framework that unifies language, vision, audio, video, and action modalities within a single Mixture-of-Transformers architecture. For autonomous vehicle applications, inverse dynamics predictions estimate the ego-motion trajectory—the sequence of pose deltas representing the vehicle's movement—from input camera footage. This capability allows developers to extract 9-degree-of-freedom motion parameters directly from monocular or multi-view AV recordings.

## Understanding Inverse Dynamics in Cosmos 3

### Action Tokens and Ego-Motion Representation

In the Cosmos 3 architecture, **action** is treated as a first-class modality alongside text and vision. According to [`cookbooks/cosmos3/generator/action/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/README.md), each action token encodes a 9-dimensional pose delta comprising 3-D translation and 6-D continuous rotation parameters. In the autonomous vehicle embodiment, this pose delta directly represents the vehicle's change in position and orientation between consecutive frames.

### The Three Action Generation Tasks

The framework supports three distinct action-generation tasks as defined in the action generator README: forward dynamics, inverse dynamics, and policy-based generation. **Inverse dynamics** specifically refers to the task of estimating the historical trajectory that produced the observed video, making it ideal for ego-motion estimation where the goal is to reconstruct how the vehicle moved through the scene.

## Inference Backends for AV Ego-Motion

Cosmos 3 provides two inference backends that share the same underlying model weights and action representation, ensuring identical ego-motion trajectories regardless of the serving stack.

### Cosmos Framework (Native PyTorch)

The native PyTorch implementation provides low-latency inference through a Python entry point that builds an input specification pairing the AV video with an empty action trajectory. As implemented in `cosmos_framework.scripts.inference`, the command uses `torchrun` for distributed execution:

```bash
torchrun --nproc-per-node=1 \
  -m cosmos_framework.scripts.inference \
  --parallelism-preset=latency \
  -i av_input_spec.json \
  -o /tmp/cosmos3_action_id \
  --checkpoint-path Cosmos3-Nano \
  --seed 0

```

### vLLM-Omni (OpenAI-Compatible Server)

The vLLM-Omni backend exposes the model via a REST API endpoint at `/v1/videos`, enabling integration with existing AV pipelines using standard HTTP clients. Clients POST a multipart request containing the video file and a JSON payload specifying the inverse dynamics mode.

## Implementation Guide

### Preparing Input Specifications

For the Cosmos Framework backend, create a JSON input specification that references your AV video and specifies the inverse dynamics mode:

```json
{
  "input_reference": "av_video.mp4",
  "mode": "id"
}

```

### Running Inference with Cosmos Framework

Execute the inference script with the appropriate parallelism preset for latency optimization. The output directory will contain [`action.json`](https://github.com/NVIDIA/cosmos/blob/main/action.json) with the predicted trajectory:

```bash

# Create input specification

cat > av_id_input.json <<'EOF'
{
  "input_reference": "av_video.mp4",
  "mode": "id"
}
EOF

# Run inverse dynamics inference

torchrun --nproc-per-node=1 \
  -m cosmos_framework.scripts.inference \
  --parallelism-preset=latency \
  -i av_id_input.json \
  -o /tmp/cosmos3_action_id \
  --checkpoint-path Cosmos3-Nano \
  --seed 0

```

### Querying the vLLM-Omni Server

First, deploy the vLLM-Omni server using Docker Compose as described in the repository root README. Then submit the video via curl:

```bash

# Verify server availability

curl http://localhost:8001/v1/models

# Submit video for inverse dynamics prediction

curl -X POST http://localhost:8001/v1/videos \
  -F input_reference=@/path/to/av_video.mp4 \
  -F extra_params='{"mode":"id"}' \
  -o action_id.json

```

Both methods produce a JSON file containing a list of 9-D pose deltas describing the predicted ego-motion trajectory, which can be visualized or fed into downstream planning modules.

## Key Source Files

| File | Description |
|------|-------------|
| [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) | High-level overview of Cosmos 3's multimodal capabilities and action modeling |
| [`cookbooks/cosmos3/generator/action/README.md`](https://github.com/NVIDIA/cosmos/blob/main/cookbooks/cosmos3/generator/action/README.md) | Detailed documentation of inverse dynamics, forward dynamics, and policy tasks |
| `cookbooks/cosmos3/generator/action/run_id_with_cosmos_framework.ipynb` | Jupyter notebook demonstrating end-to-end inverse dynamics with native PyTorch |
| `cookbooks/cosmos3/generator/action/run_id_with_vllm.ipynb` | Notebook example using the vLLM-Omni REST API server |

## Summary

- Cosmos 3 unifies video and action modalities to perform inverse dynamics predictions for autonomous vehicle ego-motion estimation.
- Each action token encodes a 9-DOF pose delta (3-D translation + 6-D rotation) representing vehicle movement between frames.
- Two inference backends provide flexibility: the **Cosmos Framework** for native PyTorch execution and **vLLM-Omni** for OpenAI-compatible REST API access.
- Both backends produce identical JSON outputs containing sequences of pose deltas suitable for downstream AV planning and validation pipelines.
- Reference implementations are available in the `cookbooks/cosmos3/generator/action/` directory, including runnable notebooks for both deployment modes.

## Frequently Asked Questions

### What is the difference between forward and inverse dynamics in Cosmos 3?

Forward dynamics predicts future video frames given an action trajectory, while inverse dynamics reconstructs the historical action trajectory (ego-motion) that produced the observed video. For autonomous vehicle applications, inverse dynamics estimates where the vehicle has been, whereas forward dynamics simulates where it will go given a planned path.

### How does Cosmos 3 represent vehicle orientation in the pose deltas?

The 9-dimensional pose delta uses the first three dimensions for translation (x, y, z displacement) and the remaining six dimensions for continuous rotation representation. This 6-D continuous rotation encoding avoids singularities present in Euler angles and provides a smooth representation of orientation changes between frames.

### Can I run inverse dynamics inference on a single GPU?

Yes. The Cosmos Framework inference command uses `--nproc-per-node=1`, indicating single-GPU execution is supported. The vLLM-Omni server can also be configured for single-GPU deployment, though the specific GPU memory requirements depend on the model size (Cosmos3-Nano, -Base, or -Large).

### Where can I find end-to-end notebooks for testing?

The repository provides two reference notebooks in `cookbooks/cosmos3/generator/action/`: `run_id_with_cosmos_framework.ipynb` demonstrates native PyTorch inference, while `run_id_with_vllm.ipynb` shows REST API integration. Both include complete examples of preparing input data, running inference, and parsing the 9-DOF pose delta outputs.