How to Optimize Box3D Performance for Large Piles of Rigid Bodies

To optimize Box3D performance for large piles of rigid bodies, enable multithreading via workerCount, tune contact recycling distances, use collision filtering, and leverage SIMD acceleration while pre-allocating capacity to avoid runtime allocations.

Simulating hundreds of interacting rigid bodies in real-time requires careful tuning of the physics engine parameters. The erincatto/box3d repository provides specific architectural optimizations—ranging from broad-phase dynamic trees to contact recycling—that allow you to optimize Box3D performance for dense piles without sacrificing stability. By configuring the world definition and leveraging parallel task scheduling, you can maintain high frame rates even with complex stacking scenarios.

Parallelize with Multithreading and Task Scheduling

Box3D splits simulation work across CPU cores using a task-based parallelism model. The engine automatically divides broad-phase pair generation and narrow-phase collision detection into independent chunks processed via b3ParallelFor in src/parallel_for.c.

To enable multithreading, set the workerCount field in b3WorldDef before creating the world:

b3WorldDef def = b3DefaultWorldDef();
def.workerCount = 8;  // Use 8 worker threads
b3WorldId world = b3CreateWorld(&def);

The scheduler distributes moved proxies across threads in b3UpdateBroadPhasePairs (src/broad_phase.c line 998) and parallelizes the contact generation in b3Collide (src/physics_world.c line 815). Scaling is roughly linear with core count until memory bandwidth limits are reached.

Reduce SAT Overhead with Contact Recycling

For stable piles, Box3D reuses existing contact manifolds instead of recomputing them every step via the Separating Axis Theorem (SAT). This recycling logic in b3CollideTask (src/physics_world.c line 558) checks if bodies moved less than recycleDistance since the last frame.

Tune the recycling tolerance for tight stacks by adjusting contactRecycleDistance:

def.contactRecycleDistance = 0.02f;  // Default is ~0.01m; increase for stable stacks

Combine this with warm-starting to reduce solver iteration costs:

def.enableWarmStarting = true;
def.contactHertz = 20.0f;  // Lower frequency reduces iteration count (~6 instead of 10)

Filter Collision Pairs Early

Prevent unnecessary narrow-phase tests using bit-mask collision filtering. The filter check occurs in b3ShouldShapesCollide (src/contact.c line 70), which runs during broad-phase pair generation.

Assign category and mask bits to skip interactions between non-colliding groups:

b3ShapeDef terrainDef = b3DefaultShapeDef();
terrainDef.filter.categoryBits = 0x0001;  // Ground category
terrainDef.filter.maskBits = 0xFFFE;      // Collides with everything except debris (0x0002)
b3ShapeId terrain = b3CreateHullShape(bodyId, &terrainDef, &terrainHull);

This eliminates SAT calls for pairs that never interact, significantly reducing CPU load in scenes with static terrain and dynamic debris.

Accelerate with SIMD and Memory Pools

Box3D includes vectorized math kernels for AABB tests and dot products in src/simd.c. SIMD is enabled by default via BOX3D_ENABLE_SIMD in the CMake configuration, but verify it is active for your target platform:

cmake -DBOX3D_DISABLE_SIMD=OFF ..

Pre-allocate memory pools to avoid runtime reallocations when bodies enter the simulation. Set the b3Capacity struct in b3WorldDef:

def.capacity.dynamicBodyCount = 2000;   // Reserve for 2k dynamic bodies
def.capacity.dynamicShapeCount = 4000;  // Reserve for shape count
def.capacity.contactCount = 8000;       // Enough contacts for dense stacks

This initialization in b3CreateWorld (src/physics_world.c line 56) ensures bit-sets and arrays are sized once, improving cache locality.

Configure Large Worlds and Solver Precision

For worlds spanning distances greater than ±10⁶ meters, enable double-precision positioning to keep broad-phase quantization error bounded. This mode stores world positions as doubles while keeping the dynamic tree in floats for speed:

set(BOX3D_DOUBLE_PRECISION ON)  # In CMakeLists.txt

Documentation in docs/large_worlds.md explains the trade-offs. For standard piles, keep solver iterations modest (6-8 instead of the default 10) by lowering contactHertz to reduce per-step CPU usage.

Profile Bottlenecks with Built-in Timers

Identify performance constraints using the internal profiling structure updated in b3World_Step (src/physics_world.c line 332). Access timing data after each step:

b3World_Step(worldId, timeStep);
printf("Pairs: %.2f ms, Collide: %.2f ms, Solve: %.2f ms\n",
       world->profile.pairs, world->profile.collide, world->profile.solve);

This breaks down time spent in broad-phase pair generation (b3UpdateBroadPhasePairs), narrow-phase collision (b3CollideTask), and constraint solving, allowing targeted optimization.

Summary

  • Enable multithreading by setting workerCount in b3WorldDef to utilize b3ParallelFor for broad-phase and narrow-phase tasks.
  • Tune contact recycling via contactRecycleDistance and warm-starting to avoid redundant SAT calculations in b3CollideTask.
  • Use collision filtering with categoryBits and maskBits to skip unnecessary tests in b3ShouldShapesCollide.
  • Leverage SIMD instructions in src/simd.c for vectorized AABB and math operations.
  • Pre-allocate capacity for bodies, shapes, and contacts to eliminate runtime memory churn.
  • Profile with world->profile to identify whether broad-phase, collision, or solving dominates frame time.

Frequently Asked Questions

How many worker threads should I configure for Box3D?

Start with the number of physical cores on your CPU, typically between 4 and 16. The b3ParallelFor implementation in src/parallel_for.c efficiently distributes work for broad-phase updates and contact generation, but excessive threads may cause contention. Profile with world->profile to ensure the overhead of task splitting does not exceed the gains from parallel execution.

What is contact recycling and why does it improve stack performance?

Contact recycling reuses existing contact manifolds from the previous frame when bodies remain within recycleDistance, avoiding expensive SAT recomputation. In src/physics_world.c, the b3CollideTask function checks angularDistance and distSquared against the threshold before regenerating contacts. For stable piles where bodies barely move, this skips thousands of redundant collision calculations per step.

When should I enable double-precision mode for large piles?

Enable BOX3D_DOUBLE_PRECISION in CMake when your simulation world extends beyond approximately 10⁶ meters in any dimension. According to docs/large_worlds.md, this keeps the broad-phase tree (src/broad_phase.c) fast while preventing numerical drift in body positions. For smaller scenes, single-precision maintains better cache performance and SIMD utilization.

How can I reduce solver iteration costs without destabilizing stacks?

Lower the contactHertz value in b3WorldDef to reduce the number of constraint solver iterations per step (e.g., from 10 to 6). Combine this with enableWarmStarting = true to reuse previous frame impulses as initial guesses. This configuration minimizes CPU usage while maintaining stability for resting contacts, as the warm-starting reduces the iterations needed for convergence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →