# Integrating Newton with Reinforcement Learning for Robot Policies: A Complete Technical Guide

> Master integrating Newton with reinforcement learning for robot policies. This guide shows how to load TorchScript policies for GPU-accelerated simulation and model building with the newton-physics library.

- Repository: [Newton Physics/newton](https://github.com/newton-physics/newton)
- Tags: how-to-guide
- Published: 2026-03-19

---

**You can integrate Newton with reinforcement learning by loading a TorchScript policy into the [`example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/example_robot_policy.py) framework, which handles asset downloading, model building, observation construction, and GPU-accelerated simulation stepping.**

Newton is an open-source physics engine developed by the `newton-physics/newton` repository that combines Warp-accelerated simulation with MuJoCo-compatible solvers. Integrating Newton with reinforcement learning for robot policies enables real-time control of complex locomotion behaviors using pre-trained neural networks while leveraging GPU-accelerated physics.

## Why Use Newton for RL-Based Robot Control?

Newton provides a **high-performance simulation core** built on Warp with MuJoCo-style solvers, offering a clean Python API that bridges the gap between physics simulation and machine learning frameworks. Unlike traditional simulators that require complex middleware, Newton exposes `State` and `Control` objects that integrate directly with PyTorch tensors, enabling seamless policy inference loops.

## End-to-End Integration Architecture

The [`example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/example_robot_policy.py) script demonstrates the complete stack for integrating Newton with reinforcement learning. The architecture separates concerns into distinct layers:

- **Asset Management**: Downloads USD assets on-demand using `newton.utils.download_asset` from [`newton/_src/utils/download_assets.py`](https://github.com/newton-physics/newton/blob/main/newton/_src/utils/download_assets.py)
- **Model Construction**: Builds a `ModelBuilder` with robot USD and configures physics parameters via [`newton/_src/sim/builder.py`](https://github.com/newton-physics/newton/blob/main/newton/_src/sim/builder.py)
- **Solver Setup**: Instantiates `newton.solvers.SolverMuJoCo` from [`newton/solvers.py`](https://github.com/newton-physics/newton/blob/main/newton/solvers.py) for GPU-accelerated MuJoCo-compatible simulation
- **State & Control**: Manages simulation state through `newton.State` and actuator commands via `newton.Control` defined in [`newton/__init__.py`](https://github.com/newton-physics/newton/blob/main/newton/__init__.py)
- **Policy Interface**: Loads TorchScript policies using `torch.jit.load` and constructs observation tensors via `compute_obs` in [`newton/examples/robot/example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/newton/examples/robot/example_robot_policy.py)
- **Viewer & Interaction**: Handles keyboard input and logging through `viewer.is_key_down` and `viewer.log_state` in [`newton/viewer.py`](https://github.com/newton-physics/newton/blob/main/newton/viewer.py)
- **Simulation Loop**: Steps the solver using `Example.step()` and `Example.simulate()` with optional CUDA graph capture for maximum throughput

## Step-by-Step Implementation Guide

### Downloading Robot Assets

Newton manages robot models through the separate `newton-assets` repository. The `newton.utils.download_asset` function handles cloning and caching:

```python
import newton
import yaml

asset_dir = str(newton.utils.download_asset(robot_config.asset_dir))
yaml_path = f"{asset_dir}/{robot_config.yaml_path}"
with open(yaml_path, encoding="utf-8") as f:
    config = yaml.safe_load(f)

```

This YAML file stores joint-position defaults, stiffness, damping, and metadata required for policy initialization.

### Building the Simulation Model

After downloading assets, construct the simulation using `ModelBuilder` from [`newton/_src/sim/builder.py`](https://github.com/newton-physics/newton/blob/main/newton/_src/sim/builder.py):

1. Instantiate `newton.ModelBuilder`
2. Load the robot USD and configure physics parameters
3. Set initial joint configuration via `builder.joint_q`
4. Apply per-joint PD gains using `builder.joint_target_ke` and `builder.joint_target_kd`
5. Add ground plane and finalize with `builder.finalize()`

### Configuring the Physics Solver

Newton provides GPU-accelerated solvers compatible with MuJoCo conventions. Instantiate `SolverMuJoCo` from [`newton/solvers.py`](https://github.com/newton-physics/newton/blob/main/newton/solvers.py):

```python
solver = newton.solvers.SolverMuJoCo(
    model=builder.model,
    device="cuda",
    # Additional solver parameters

)

```

This solver runs on the GPU via Warp, enabling real-time simulation of complex robots.

### Loading and Running the RL Policy

The policy integration happens in [`newton/examples/robot/example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/newton/examples/robot/example_robot_policy.py). First, load the TorchScript policy:

```python
import torch

policy = torch.jit.load(policy_path, map_location=example.torch_device)

# Prepare static tensors (run once)

joint_q = example.state_0.joint_q
example.joint_pos_initial = torch.tensor(
    joint_q[7:], device=example.torch_device, dtype=torch.float32
).unsqueeze(0)
example.act = torch.zeros(1, num_dofs, device=example.torch_device)

```

Then implement the main simulation loop:

```python
while running:
    # ① Read keyboard command (forward / lateral / yaw)

    example.step()            # handles command → obs → policy → act

    example.render()          # draw frame

```

## Key Code Implementation Details

### Observation Construction for Locomotion

The `compute_obs` function in [`example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/example_robot_policy.py) constructs the observation vector required by locomotion policies:

```python
def compute_obs(
    actions, state, joint_pos_initial, device,
    indices, gravity_vec, command,
):
    # Extract root pose/velocity

    root_quat = torch.tensor(state.joint_q[3:7], device=device).unsqueeze(0)
    root_vel  = torch.tensor(state.joint_qd[:3], device=device).unsqueeze(0)
    root_ang  = torch.tensor(state.joint_qd[3:6], device=device).unsqueeze(0)

    # Transform to robot base frame

    vel_b   = quat_rotate_inverse(root_quat, root_vel)
    ang_b   = quat_rotate_inverse(root_quat, root_ang)
    grav_b  = quat_rotate_inverse(root_quat, gravity_vec)

    # Joint-space relative positions / velocities

    joint_pos_rel = torch.tensor(state.joint_q[7:], device=device).unsqueeze(0) - joint_pos_initial
    joint_vel_rel = torch.tensor(state.joint_qd[6:], device=device).unsqueeze(0)

    # Re-order joints to match policy ordering

    joint_pos = torch.index_select(joint_pos_rel, 1, indices)
    joint_vel = torch.index_select(joint_vel_rel, 1, indices)

    # Concatenate all terms

    return torch.cat([vel_b, ang_b, grav_b, command, joint_pos, joint_vel, actions], dim=1)

```

This function handles coordinate frame transformations, joint reordering between MuJoCo and PhysX conventions, and tensor preparation for neural network input.

### Policy Inference Loop

The `Example.step()` method orchestrates the policy inference:

1. **Command Generation**: Reads keyboard input via `viewer.is_key_down` to generate velocity commands
2. **Observation Computation**: Calls `compute_obs` to build the observation tensor
3. **Policy Execution**: Runs `example.policy(obs)` to obtain actions
4. **Action Processing**: Re-orders actions to PhysX joint order and applies scaling
5. **Control Update**: Writes targets to `example.control.joint_target_pos`
6. **Physics Stepping**: Calls the solver for `decimation` sub-steps with optional CUDA graph capture

## Performance Optimization with CUDA Graphs

Newton supports **CUDA graph capture** for eliminating CPU launch overhead during the simulation loop. When `example.use_cuda_graph` is enabled, the solver execution is captured as a CUDA graph and replayed for subsequent steps, significantly improving throughput for RL training and deployment scenarios where the simulation loop runs thousands of iterations with identical kernel arguments.

## Summary

- Newton provides a **high-performance simulation core** using Warp and MuJoCo-compatible solvers for GPU-accelerated robot simulation.
- The [`example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/example_robot_policy.py) script in `newton/examples/robot/` demonstrates the complete pipeline for **integrating Newton with reinforcement learning**.
- Key components include `newton.ModelBuilder` for constructing robot models, `newton.solvers.SolverMuJoCo` for physics stepping, and `newton.State`/`newton.Control` for state management.
- The observation pipeline handles coordinate transformations, joint reordering between MuJoCo and PhysX conventions, and tensor preparation for PyTorch policies.
- CUDA graph capture support enables low-latency simulation loops suitable for real-time RL policy deployment.

## Frequently Asked Questions

### What physics backends does Newton support for reinforcement learning?

Newton supports both **MuJoCo-style** and **PhysX** physics backends through its solver abstraction. The `SolverMuJoCo` class provides GPU-accelerated simulation compatible with MuJoCo conventions, while the underlying Warp framework enables cross-platform execution. You can switch between backends by selecting the appropriate solver class when instantiating your simulation.

### How do I map joint indices between MuJoCo and PhysX in Newton?

Newton provides the `find_physx_mjwarp_mapping` function to handle joint ordering differences between MuJoCo-style (`mjw`) and PhysX (`physx`) conventions. This mapping is essential when loading policies trained in one simulator but executed in another. The `compute_obs` function in [`example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/example_robot_policy.py) demonstrates using `torch.index_select` with these indices to reorder joint positions and velocities before feeding them to the policy network.

### Can I use Newton with libraries other than PyTorch?

While the reference implementation in [`example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/example_robot_policy.py) uses **PyTorch** and **TorchScript** for policy inference, Newton's Python API is framework-agnostic regarding the physics simulation itself. You can integrate TensorFlow, JAX, or NumPy-based policies by converting the `newton.State` tensors (accessible via `state.joint_q` and `state.joint_qd`) to your framework's array format, running inference, and writing back to `newton.Control` objects.

### Where can I find working examples of RL integration?

The primary reference is [`newton/examples/robot/example_robot_policy.py`](https://github.com/newton-physics/newton/blob/main/newton/examples/robot/example_robot_policy.py), which demonstrates loading a TorchScript policy, constructing observations, and running a real-time simulation loop with keyboard commands. Additional validation exists in [`newton/tests/test_examples.py`](https://github.com/newton-physics/newton/blob/main/newton/tests/test_examples.py), which runs the RL policy example as part of the continuous integration suite to prevent regressions. The `newton-assets` repository contains the USD models and YAML configurations required by these examples.