Integrating Newton with Reinforcement Learning for Robot Policies: A Complete Technical Guide
You can integrate Newton with reinforcement learning by loading a TorchScript policy into the example_robot_policy.py framework, which handles asset downloading, model building, observation construction, and GPU-accelerated simulation stepping.
Newton is an open-source physics engine developed by the newton-physics/newton repository that combines Warp-accelerated simulation with MuJoCo-compatible solvers. Integrating Newton with reinforcement learning for robot policies enables real-time control of complex locomotion behaviors using pre-trained neural networks while leveraging GPU-accelerated physics.
Why Use Newton for RL-Based Robot Control?
Newton provides a high-performance simulation core built on Warp with MuJoCo-style solvers, offering a clean Python API that bridges the gap between physics simulation and machine learning frameworks. Unlike traditional simulators that require complex middleware, Newton exposes State and Control objects that integrate directly with PyTorch tensors, enabling seamless policy inference loops.
End-to-End Integration Architecture
The example_robot_policy.py script demonstrates the complete stack for integrating Newton with reinforcement learning. The architecture separates concerns into distinct layers:
- Asset Management: Downloads USD assets on-demand using
newton.utils.download_assetfromnewton/_src/utils/download_assets.py - Model Construction: Builds a
ModelBuilderwith robot USD and configures physics parameters vianewton/_src/sim/builder.py - Solver Setup: Instantiates
newton.solvers.SolverMuJoCofromnewton/solvers.pyfor GPU-accelerated MuJoCo-compatible simulation - State & Control: Manages simulation state through
newton.Stateand actuator commands vianewton.Controldefined innewton/__init__.py - Policy Interface: Loads TorchScript policies using
torch.jit.loadand constructs observation tensors viacompute_obsinnewton/examples/robot/example_robot_policy.py - Viewer & Interaction: Handles keyboard input and logging through
viewer.is_key_downandviewer.log_stateinnewton/viewer.py - Simulation Loop: Steps the solver using
Example.step()andExample.simulate()with optional CUDA graph capture for maximum throughput
Step-by-Step Implementation Guide
Downloading Robot Assets
Newton manages robot models through the separate newton-assets repository. The newton.utils.download_asset function handles cloning and caching:
import newton
import yaml
asset_dir = str(newton.utils.download_asset(robot_config.asset_dir))
yaml_path = f"{asset_dir}/{robot_config.yaml_path}"
with open(yaml_path, encoding="utf-8") as f:
config = yaml.safe_load(f)
This YAML file stores joint-position defaults, stiffness, damping, and metadata required for policy initialization.
Building the Simulation Model
After downloading assets, construct the simulation using ModelBuilder from newton/_src/sim/builder.py:
- Instantiate
newton.ModelBuilder - Load the robot USD and configure physics parameters
- Set initial joint configuration via
builder.joint_q - Apply per-joint PD gains using
builder.joint_target_keandbuilder.joint_target_kd - Add ground plane and finalize with
builder.finalize()
Configuring the Physics Solver
Newton provides GPU-accelerated solvers compatible with MuJoCo conventions. Instantiate SolverMuJoCo from newton/solvers.py:
solver = newton.solvers.SolverMuJoCo(
model=builder.model,
device="cuda",
# Additional solver parameters
)
This solver runs on the GPU via Warp, enabling real-time simulation of complex robots.
Loading and Running the RL Policy
The policy integration happens in newton/examples/robot/example_robot_policy.py. First, load the TorchScript policy:
import torch
policy = torch.jit.load(policy_path, map_location=example.torch_device)
# Prepare static tensors (run once)
joint_q = example.state_0.joint_q
example.joint_pos_initial = torch.tensor(
joint_q[7:], device=example.torch_device, dtype=torch.float32
).unsqueeze(0)
example.act = torch.zeros(1, num_dofs, device=example.torch_device)
Then implement the main simulation loop:
while running:
# ① Read keyboard command (forward / lateral / yaw)
example.step() # handles command → obs → policy → act
example.render() # draw frame
Key Code Implementation Details
Observation Construction for Locomotion
The compute_obs function in example_robot_policy.py constructs the observation vector required by locomotion policies:
def compute_obs(
actions, state, joint_pos_initial, device,
indices, gravity_vec, command,
):
# Extract root pose/velocity
root_quat = torch.tensor(state.joint_q[3:7], device=device).unsqueeze(0)
root_vel = torch.tensor(state.joint_qd[:3], device=device).unsqueeze(0)
root_ang = torch.tensor(state.joint_qd[3:6], device=device).unsqueeze(0)
# Transform to robot base frame
vel_b = quat_rotate_inverse(root_quat, root_vel)
ang_b = quat_rotate_inverse(root_quat, root_ang)
grav_b = quat_rotate_inverse(root_quat, gravity_vec)
# Joint-space relative positions / velocities
joint_pos_rel = torch.tensor(state.joint_q[7:], device=device).unsqueeze(0) - joint_pos_initial
joint_vel_rel = torch.tensor(state.joint_qd[6:], device=device).unsqueeze(0)
# Re-order joints to match policy ordering
joint_pos = torch.index_select(joint_pos_rel, 1, indices)
joint_vel = torch.index_select(joint_vel_rel, 1, indices)
# Concatenate all terms
return torch.cat([vel_b, ang_b, grav_b, command, joint_pos, joint_vel, actions], dim=1)
This function handles coordinate frame transformations, joint reordering between MuJoCo and PhysX conventions, and tensor preparation for neural network input.
Policy Inference Loop
The Example.step() method orchestrates the policy inference:
- Command Generation: Reads keyboard input via
viewer.is_key_downto generate velocity commands - Observation Computation: Calls
compute_obsto build the observation tensor - Policy Execution: Runs
example.policy(obs)to obtain actions - Action Processing: Re-orders actions to PhysX joint order and applies scaling
- Control Update: Writes targets to
example.control.joint_target_pos - Physics Stepping: Calls the solver for
decimationsub-steps with optional CUDA graph capture
Performance Optimization with CUDA Graphs
Newton supports CUDA graph capture for eliminating CPU launch overhead during the simulation loop. When example.use_cuda_graph is enabled, the solver execution is captured as a CUDA graph and replayed for subsequent steps, significantly improving throughput for RL training and deployment scenarios where the simulation loop runs thousands of iterations with identical kernel arguments.
Summary
- Newton provides a high-performance simulation core using Warp and MuJoCo-compatible solvers for GPU-accelerated robot simulation.
- The
example_robot_policy.pyscript innewton/examples/robot/demonstrates the complete pipeline for integrating Newton with reinforcement learning. - Key components include
newton.ModelBuilderfor constructing robot models,newton.solvers.SolverMuJoCofor physics stepping, andnewton.State/newton.Controlfor state management. - The observation pipeline handles coordinate transformations, joint reordering between MuJoCo and PhysX conventions, and tensor preparation for PyTorch policies.
- CUDA graph capture support enables low-latency simulation loops suitable for real-time RL policy deployment.
Frequently Asked Questions
What physics backends does Newton support for reinforcement learning?
Newton supports both MuJoCo-style and PhysX physics backends through its solver abstraction. The SolverMuJoCo class provides GPU-accelerated simulation compatible with MuJoCo conventions, while the underlying Warp framework enables cross-platform execution. You can switch between backends by selecting the appropriate solver class when instantiating your simulation.
How do I map joint indices between MuJoCo and PhysX in Newton?
Newton provides the find_physx_mjwarp_mapping function to handle joint ordering differences between MuJoCo-style (mjw) and PhysX (physx) conventions. This mapping is essential when loading policies trained in one simulator but executed in another. The compute_obs function in example_robot_policy.py demonstrates using torch.index_select with these indices to reorder joint positions and velocities before feeding them to the policy network.
Can I use Newton with libraries other than PyTorch?
While the reference implementation in example_robot_policy.py uses PyTorch and TorchScript for policy inference, Newton's Python API is framework-agnostic regarding the physics simulation itself. You can integrate TensorFlow, JAX, or NumPy-based policies by converting the newton.State tensors (accessible via state.joint_q and state.joint_qd) to your framework's array format, running inference, and writing back to newton.Control objects.
Where can I find working examples of RL integration?
The primary reference is newton/examples/robot/example_robot_policy.py, which demonstrates loading a TorchScript policy, constructing observations, and running a real-time simulation loop with keyboard commands. Additional validation exists in newton/tests/test_examples.py, which runs the RL policy example as part of the continuous integration suite to prevent regressions. The newton-assets repository contains the USD models and YAML configurations required by these examples.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →