NVIDIA Engineers Detail GPU Scaling for MuJoCo Robotics Simulation Using Warp
A new guide from NVIDIA demonstrates how MuJoCo Warp leverages GPU kernel compilation to scale robotics environments like the SO-101 arm up to 2,048 parallel worlds for reinforcement learning.
Scaling Robot Simulation from Single CPU Worlds to Massive GPU Batches
On September 23, 2026, NVIDIA researchers Johnny Nuñez Cano, Asier Arranz, Rishabh Chadha, and Ben Oliveri published an implementation guide showing how robotics developers can transition traditional CPU-based MuJoCo simulations into high-throughput GPU workloads. As physical AI and reinforcement learning pipelines expand, training speed increasingly depends on running hundreds or thousands of environments at once rather than simply optimizing single-world latency.
To bridge this gap, NVIDIA introduced a workflow using MuJoCo Warp (MJWarp), a GPU-accelerated implementation of the MuJoCo physics engine built on NVIDIA Warp. As the second entry in NVIDIA's "State of Simulation for Physical AI" series, the walkthrough demonstrates moving an SO-101 follower robotic arm running a block-stacking task from a standard Python-driven CPU environment directly into 2,048 parallelized simulation states on an NVIDIA GPU.
Porting the SO-101 Arm: Buffer Sizing, CUDA Graphs, and Parity Validation
At the foundation of this setup is NVIDIA Warp, a Python-based framework that compiles statically typed Python kernels into native CUDA code with support for just-in-time (JIT) compilation, kernel fusion, automatic differentiation via wp.Tape, and deterministic execution modes introduced in version 1.15. MJWarp applies this framework to MuJoCo’s physics engine, processing standard MJCF XML scene descriptions while shifting arrays to the GPU with an added leading batch dimension.
The migration begins with an SO-101 pick-and-place task configured at 50 Hz control frequency with 10 physics substeps per frame, targeting a 0.002-second physics timestep. The arm's task is to grasp a 44 mm red cube and place it atop an identically sized blue cube. Moving this workflow to MJWarp involves an API transition: developers upload the model using mjw.put_model(), allocate GPU-resident state using mjw.make_data() or mjw.put_data(), and advance all parallel worlds simultaneously via mjw.step().
A crucial operational requirement is managing device-side contact and constraint buffers. Developers must define capacities such as nconmax (maximum contacts per environment), naconmax (global contact ceiling), and njmax (maximum constraints per environment). Because MJWarp issues warnings rather than terminating immediately when narrowphase collisions exceed capacity, runs risk silent invalidation unless verified. NVIDIA recommends sizing these limits during the highest-contact moment of a task—such as when both gripper jaws and the table touch a cube—and diagnosing allocations using tools like mjwarp-testspeed --measure_alloc.
Once single-world physics parity is verified against the CPU baseline, environments can scale up to 2,048 worlds or more. To eliminate dispatch latency across thousands of batched steps, developers wrap mjw.step() within wp.ScopedCapture() to create replayable CUDA Graphs. Profiling must account for GPU asynchronous execution by inserting wp.synchronize() before and after timed loops, and by keeping simulation data on-device rather than pulling arrays back to host memory via .numpy() during rollout steps.
Expanding the Physical AI Simulation Stack Toward Newton and Isaac Lab
The shift to MJWarp highlights a growing industry division between single-robot model-predictive control (MPC) and batched reinforcement learning. While classic CPU MuJoCo remains standard for teleoperation, single-world debugging, and low-latency MPC, batched GPU environments target massive data collection.
MJWarp is designed to interface across several modern physical AI frameworks. It serves as an execution backend for mjlab (which pairs MJWarp directly with PyTorch), MuJoCo Playground via MJX's Warp implementation, and broader learning pipelines. NVIDIA stated that the next installment of the simulation series will demonstrate how to import this same SO-101 environment into its multi-solver framework, Newton, where MJWarp will operate as the underlying rigid-body solver (newton.solvers.SolverMuJoCo) to connect scenes directly with Isaac Lab training workflows.



