Integrating Newton with Reinforcement Learning for Robot Policies: A Complete Technical Guide

You can integrate Newton with reinforcement learning by loading a TorchScript policy into the example_robot_policy.py framework, which handles asset downloading, model building, observation construction, and GPU-accelerated simulation stepping.

Newton is an open-source physics engine developed by the newton-physics/newton repository that combines Warp-accelerated simulation with MuJoCo-compatible solvers. Integrating Newton with reinforcement learning for robot policies enables real-time control of complex locomotion behaviors using pre-trained neural networks while leveraging GPU-accelerated physics.

Why Use Newton for RL-Based Robot Control?

Newton provides a high-performance simulation core built on Warp with MuJoCo-style solvers, offering a clean Python API that bridges the gap between physics simulation and machine learning frameworks. Unlike traditional simulators that require complex middleware, Newton exposes State and Control objects that integrate directly with PyTorch tensors, enabling seamless policy inference loops.

End-to-End Integration Architecture

The example_robot_policy.py script demonstrates the complete stack for integrating Newton with reinforcement learning. The architecture separates concerns into distinct layers:

  • Asset Management: Downloads USD assets on-demand using newton.utils.download_asset from newton/_src/utils/download_assets.py
  • Model Construction: Builds a ModelBuilder with robot USD and configures physics parameters via newton/_src/sim/builder.py
  • Solver Setup: Instantiates newton.solvers.SolverMuJoCo from newton/solvers.py for GPU-accelerated MuJoCo-compatible simulation
  • State & Control: Manages simulation state through newton.State and actuator commands via newton.Control defined in newton/__init__.py
  • Policy Interface: Loads TorchScript policies using torch.jit.load and constructs observation tensors via compute_obs in newton/examples/robot/example_robot_policy.py
  • Viewer & Interaction: Handles keyboard input and logging through viewer.is_key_down and viewer.log_state in newton/viewer.py
  • Simulation Loop: Steps the solver using Example.step() and Example.simulate() with optional CUDA graph capture for maximum throughput

Step-by-Step Implementation Guide

Downloading Robot Assets

Newton manages robot models through the separate newton-assets repository. The newton.utils.download_asset function handles cloning and caching:

import newton
import yaml

asset_dir = str(newton.utils.download_asset(robot_config.asset_dir))
yaml_path = f"{asset_dir}/{robot_config.yaml_path}"
with open(yaml_path, encoding="utf-8") as f:
    config = yaml.safe_load(f)

This YAML file stores joint-position defaults, stiffness, damping, and metadata required for policy initialization.

Building the Simulation Model

After downloading assets, construct the simulation using ModelBuilder from newton/_src/sim/builder.py:

  1. Instantiate newton.ModelBuilder
  2. Load the robot USD and configure physics parameters
  3. Set initial joint configuration via builder.joint_q
  4. Apply per-joint PD gains using builder.joint_target_ke and builder.joint_target_kd
  5. Add ground plane and finalize with builder.finalize()

Configuring the Physics Solver

Newton provides GPU-accelerated solvers compatible with MuJoCo conventions. Instantiate SolverMuJoCo from newton/solvers.py:

solver = newton.solvers.SolverMuJoCo(
    model=builder.model,
    device="cuda",
    # Additional solver parameters

)

This solver runs on the GPU via Warp, enabling real-time simulation of complex robots.

Loading and Running the RL Policy

The policy integration happens in newton/examples/robot/example_robot_policy.py. First, load the TorchScript policy:

import torch

policy = torch.jit.load(policy_path, map_location=example.torch_device)

# Prepare static tensors (run once)

joint_q = example.state_0.joint_q
example.joint_pos_initial = torch.tensor(
    joint_q[7:], device=example.torch_device, dtype=torch.float32
).unsqueeze(0)
example.act = torch.zeros(1, num_dofs, device=example.torch_device)

Then implement the main simulation loop:

while running:
    # ① Read keyboard command (forward / lateral / yaw)

    example.step()            # handles command → obs → policy → act

    example.render()          # draw frame

Key Code Implementation Details

Observation Construction for Locomotion

The compute_obs function in example_robot_policy.py constructs the observation vector required by locomotion policies:

def compute_obs(
    actions, state, joint_pos_initial, device,
    indices, gravity_vec, command,
):
    # Extract root pose/velocity

    root_quat = torch.tensor(state.joint_q[3:7], device=device).unsqueeze(0)
    root_vel  = torch.tensor(state.joint_qd[:3], device=device).unsqueeze(0)
    root_ang  = torch.tensor(state.joint_qd[3:6], device=device).unsqueeze(0)

    # Transform to robot base frame

    vel_b   = quat_rotate_inverse(root_quat, root_vel)
    ang_b   = quat_rotate_inverse(root_quat, root_ang)
    grav_b  = quat_rotate_inverse(root_quat, gravity_vec)

    # Joint-space relative positions / velocities

    joint_pos_rel = torch.tensor(state.joint_q[7:], device=device).unsqueeze(0) - joint_pos_initial
    joint_vel_rel = torch.tensor(state.joint_qd[6:], device=device).unsqueeze(0)

    # Re-order joints to match policy ordering

    joint_pos = torch.index_select(joint_pos_rel, 1, indices)
    joint_vel = torch.index_select(joint_vel_rel, 1, indices)

    # Concatenate all terms

    return torch.cat([vel_b, ang_b, grav_b, command, joint_pos, joint_vel, actions], dim=1)

This function handles coordinate frame transformations, joint reordering between MuJoCo and PhysX conventions, and tensor preparation for neural network input.

Policy Inference Loop

The Example.step() method orchestrates the policy inference:

  1. Command Generation: Reads keyboard input via viewer.is_key_down to generate velocity commands
  2. Observation Computation: Calls compute_obs to build the observation tensor
  3. Policy Execution: Runs example.policy(obs) to obtain actions
  4. Action Processing: Re-orders actions to PhysX joint order and applies scaling
  5. Control Update: Writes targets to example.control.joint_target_pos
  6. Physics Stepping: Calls the solver for decimation sub-steps with optional CUDA graph capture

Performance Optimization with CUDA Graphs

Newton supports CUDA graph capture for eliminating CPU launch overhead during the simulation loop. When example.use_cuda_graph is enabled, the solver execution is captured as a CUDA graph and replayed for subsequent steps, significantly improving throughput for RL training and deployment scenarios where the simulation loop runs thousands of iterations with identical kernel arguments.

Summary

  • Newton provides a high-performance simulation core using Warp and MuJoCo-compatible solvers for GPU-accelerated robot simulation.
  • The example_robot_policy.py script in newton/examples/robot/ demonstrates the complete pipeline for integrating Newton with reinforcement learning.
  • Key components include newton.ModelBuilder for constructing robot models, newton.solvers.SolverMuJoCo for physics stepping, and newton.State/newton.Control for state management.
  • The observation pipeline handles coordinate transformations, joint reordering between MuJoCo and PhysX conventions, and tensor preparation for PyTorch policies.
  • CUDA graph capture support enables low-latency simulation loops suitable for real-time RL policy deployment.

Frequently Asked Questions

What physics backends does Newton support for reinforcement learning?

Newton supports both MuJoCo-style and PhysX physics backends through its solver abstraction. The SolverMuJoCo class provides GPU-accelerated simulation compatible with MuJoCo conventions, while the underlying Warp framework enables cross-platform execution. You can switch between backends by selecting the appropriate solver class when instantiating your simulation.

How do I map joint indices between MuJoCo and PhysX in Newton?

Newton provides the find_physx_mjwarp_mapping function to handle joint ordering differences between MuJoCo-style (mjw) and PhysX (physx) conventions. This mapping is essential when loading policies trained in one simulator but executed in another. The compute_obs function in example_robot_policy.py demonstrates using torch.index_select with these indices to reorder joint positions and velocities before feeding them to the policy network.

Can I use Newton with libraries other than PyTorch?

While the reference implementation in example_robot_policy.py uses PyTorch and TorchScript for policy inference, Newton's Python API is framework-agnostic regarding the physics simulation itself. You can integrate TensorFlow, JAX, or NumPy-based policies by converting the newton.State tensors (accessible via state.joint_q and state.joint_qd) to your framework's array format, running inference, and writing back to newton.Control objects.

Where can I find working examples of RL integration?

The primary reference is newton/examples/robot/example_robot_policy.py, which demonstrates loading a TorchScript policy, constructing observations, and running a real-time simulation loop with keyboard commands. Additional validation exists in newton/tests/test_examples.py, which runs the RL policy example as part of the continuous integration suite to prevent regressions. The newton-assets repository contains the USD models and YAML configurations required by these examples.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →