Newton GPU vs CPU Simulation: A Complete Performance Optimization Guide

Newton's unified API delivers 2–5× faster physics simulations on CUDA GPUs compared to CPU execution by leveraging 3-D texture-based SDF collision detection and CSR adjacency structures, while maintaining full code compatibility across both backends.

Newton is a Python-centric physics engine built on NVIDIA’s Warp (wp) that enables seamless Newton GPU vs CPU simulation without modifying user code. The engine automatically detects the target device and dispatches optimized kernels for either CUDA GPUs or multithreaded CPU backends. This architecture allows developers to prototype on CPU hardware and deploy production simulations on GPU clusters using identical Python scripts.

Architecture Overview and Device Abstraction

Newton centralizes device selection through Warp's device abstraction layer. The engine queries wp.get_device() to determine whether to target CUDA hardware or fall back to CPU execution, with device.is_cuda and device.is_cpu boolean flags guarding GPU-specific code paths.

In newton/_src/sim/builder.py (lines 9583-9599), the simulation builder validates device capabilities during model construction. When a user requests a CPU device, the system checks for incompatible GPU-only features such as mesh-mesh SDF collisions. If detected, the builder raises a ValueError at lines 9597-9599 to prevent runtime crashes, ensuring predictable performance characteristics across hardware configurations.

Execution Flow: How Newton Dispatches GPU vs CPU Kernels

Newton follows a four-stage execution pipeline that remains transparent to end users:

  1. Device Selection – Users explicitly request hardware via wp.get_device("cuda:0") or accept the default. The system stores this choice in the model's device property.

  2. Data Placement – Data structures including adjacency lists and meshes call .to(device) to migrate memory to GPU VRAM when necessary. The utility helper in newton/_src/utils/mesh.py (lines 105-110) handles NumPy-to-Warp array conversions with optional device transfer.

  3. Kernel Launch – wp.launch() dispatches computations to Warp's backend. For GPUs, this invokes the CUDA driver; for CPUs, it uses a multithreaded host backend executing identical @wp.func kernels.

  4. Fallback Validation – Before simulation begins, newton/_src/geometry/narrow_phase.py (lines 1415-1424) emits warnings when CPU execution encounters mesh-mesh SDF contacts, skipping unsupported collision paths rather than failing silently.

Performance-Critical Optimizations

3-D Texture-Based SDF Pipeline

Newton replaced NanoVDB volume representations with a compact 3-D texture-based signed-distance-field implementation in newton/_src/geometry/sdf_texture.py (lines 991-1005). This change drastically reduces memory bandwidth consumption during texture fetches and enables CPU compatibility for the SDF pipeline while maintaining GPU acceleration. The texture representation cuts GPU memory pressure for complex mesh-mesh collisions compared to traditional voxel grids.

CSR Adjacency Structures

Joint and particle adjacency data utilize Compressed-Sparse-Row (CSR) format storage across both rigid and particle VBD solvers. In newton/_src/solvers/vbd/rigid_vbd_kernels.py (lines 511-518), the RigidForceElementAdjacencyInfo.to() method efficiently transfers CSR structures to the target device with minimal overhead. This format enables fast GPU traversal while reducing per-step memory transfer costs for large particle systems.

Early Warning System for GPU-Only Features

The engine implements defensive checks that prevent costly runtime failures. When running on CPU hardware, narrow_phase.py (lines 1415-1419) detects mesh-mesh SDF contact requests and emits a user warning before skipping those contact pairs. Similarly, builder.py validates SDF requirements during model construction, aborting early if the simulation requires CUDA-specific texture features unavailable on CPU.

Practical Implementation Examples

Switching Between CPU and GPU Execution

Changing hardware targets requires only modifying the device initialization string. The following example demonstrates identical pendulum simulations on both backends:

import newton
import warp as wp

# Toggle between backends by changing one line

device = wp.get_device("cpu")          # Multithreaded CPU execution

# device = wp.get_device("cuda:0")      # CUDA GPU acceleration

model = newton.examples.basic.example_basic_pendulum.load_model()
model.device = device                    # Force data placement

solver = newton.solvers.SolverMuJoCo(
    model,
    use_mujoco_cpu=device.is_cpu,
    iterations=5,
    dt=0.01,
)

# Benchmark 100 simulation steps

for _ in range(100):
    solver.step()

angle = model.joints[0].q[0].numpy()
print(f"Final angle: {angle:.3f} rad")

GPU execution typically achieves 2–5× speed-up for equivalent step counts, with larger gains for particle-rich scenes or complex contact geometries.

Handling GPU-Only SDF Constraints

Certain collision features explicitly require CUDA hardware. Attempting to run mesh-mesh SDF collisions on CPU triggers an informative error:

model = newton.examples.contacts.example_nut_bolt_hydro.load_model()

try:
    model.device = wp.get_device("cpu")
    solver = newton.solvers.SolverMuJoCo(model, use_mujoco_cpu=True)
    solver.step()
except ValueError as e:
    print(e)   # "SDF collision paths require a CUDA-capable GPU device..."

When executed on a CUDA device, the texture-based SDF generates on-the-fly and the simulation proceeds with full contact fidelity.

Monitoring Device-Specific Resource Usage

Inspect GPU memory allocation directly through Warp's profiling utilities:

if device.is_cuda:
    mem_bytes = wp.memory_usage(device)
    print(f"GPU memory allocated: {mem_bytes/1e6:.1f} MiB")

Key Source Files and Implementation Details

File Purpose Critical Lines
newton/_src/sim/builder.py Device validation and model construction 9583-9599 (device detection), 9597-9599 (SDF requirements)
newton/_src/geometry/narrow_phase.py Contact pair generation with CPU fallbacks 1415-1424 (warning system)
newton/_src/geometry/sdf_texture.py 3-D texture SDF implementation 991-1005 (texture creation)
newton/_src/solvers/vbd/rigid_vbd_kernels.py Rigid-body contact kernels 511-518 (device transfer methods)
newton/_src/solvers/vbd/particle_vbd_kernels.py Particle adjacency computations Entire module (CSR traversal)
newton/_src/utils/mesh.py Array conversion and device migration 105-110 (device handling)

Summary

  • Unified API: Newton GPU vs CPU simulation requires zero code changes—simply modify the wp.get_device() argument to switch backends.
  • Performance Gains: CUDA execution delivers 2–5× faster simulation steps compared to CPU, particularly for scenes with dense particle systems or complex mesh contacts.
  • Memory Efficiency: The 3-D texture-based SDF implementation reduces GPU memory bandwidth while maintaining CPU compatibility for the pipeline infrastructure.
  • Safety Mechanisms: Early validation in builder.py and narrow_phase.py prevents silent failures by explicitly warning users when CPU execution encounters GPU-only collision features.
  • Data Structure Optimization: CSR adjacency formats minimize device-to-host transfer overhead, critical for maintaining GPU utilization during iterative solver steps.

Frequently Asked Questions

Can I run Newton simulations without a CUDA-capable GPU?

Yes. Newton fully supports CPU execution through Warp's multithreaded host backend. While GPU-only features such as mesh-mesh SDF texture collisions trigger warnings or errors, the core physics kernels—including rigid-body dynamics and particle systems—execute correctly on CPU hardware using identical Python APIs.

How do I switch between CPU and GPU execution in Newton?

Change the device string passed to wp.get_device(). Use wp.get_device("cpu") for host execution or wp.get_device("cuda:0") for the first GPU. Assign this device to model.device before solver initialization; all subsequent data placement and kernel launches automatically target the selected hardware.

Why does my simulation raise an error about SDF collision on CPU?

The 3-D texture-based signed-distance-field implementation requires CUDA for mesh-mesh collision detection. When newton/_src/sim/builder.py detects an SDF collision path with a CPU device (lines 9597-9599), it raises a ValueError to prevent unsupported execution. Either switch to a CUDA device or disable SDF-based collision geometries in your model definition.

What performance difference should I expect between CPU and GPU?

Benchmarks indicate 2–5× speed improvements on GPU for standard benchmarks, with higher multiples for scenes containing hundreds of particles or complex contact manifolds. The CSR adjacency structures and texture-based SDF specifically optimize GPU memory access patterns, while CPU execution remains valuable for debugging and development on laptops without discrete graphics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →