How to Optimize Box3D Performance for Large Piles of Rigid Bodies
To optimize Box3D performance for large piles of rigid bodies, enable multithreading via workerCount, tune contact recycling distances, use collision filtering, and leverage SIMD acceleration while pre-allocating capacity to avoid runtime allocations.
Simulating hundreds of interacting rigid bodies in real-time requires careful tuning of the physics engine parameters. The erincatto/box3d repository provides specific architectural optimizations—ranging from broad-phase dynamic trees to contact recycling—that allow you to optimize Box3D performance for dense piles without sacrificing stability. By configuring the world definition and leveraging parallel task scheduling, you can maintain high frame rates even with complex stacking scenarios.
Parallelize with Multithreading and Task Scheduling
Box3D splits simulation work across CPU cores using a task-based parallelism model. The engine automatically divides broad-phase pair generation and narrow-phase collision detection into independent chunks processed via b3ParallelFor in src/parallel_for.c.
To enable multithreading, set the workerCount field in b3WorldDef before creating the world:
b3WorldDef def = b3DefaultWorldDef();
def.workerCount = 8; // Use 8 worker threads
b3WorldId world = b3CreateWorld(&def);
The scheduler distributes moved proxies across threads in b3UpdateBroadPhasePairs (src/broad_phase.c line 998) and parallelizes the contact generation in b3Collide (src/physics_world.c line 815). Scaling is roughly linear with core count until memory bandwidth limits are reached.
Reduce SAT Overhead with Contact Recycling
For stable piles, Box3D reuses existing contact manifolds instead of recomputing them every step via the Separating Axis Theorem (SAT). This recycling logic in b3CollideTask (src/physics_world.c line 558) checks if bodies moved less than recycleDistance since the last frame.
Tune the recycling tolerance for tight stacks by adjusting contactRecycleDistance:
def.contactRecycleDistance = 0.02f; // Default is ~0.01m; increase for stable stacks
Combine this with warm-starting to reduce solver iteration costs:
def.enableWarmStarting = true;
def.contactHertz = 20.0f; // Lower frequency reduces iteration count (~6 instead of 10)
Filter Collision Pairs Early
Prevent unnecessary narrow-phase tests using bit-mask collision filtering. The filter check occurs in b3ShouldShapesCollide (src/contact.c line 70), which runs during broad-phase pair generation.
Assign category and mask bits to skip interactions between non-colliding groups:
b3ShapeDef terrainDef = b3DefaultShapeDef();
terrainDef.filter.categoryBits = 0x0001; // Ground category
terrainDef.filter.maskBits = 0xFFFE; // Collides with everything except debris (0x0002)
b3ShapeId terrain = b3CreateHullShape(bodyId, &terrainDef, &terrainHull);
This eliminates SAT calls for pairs that never interact, significantly reducing CPU load in scenes with static terrain and dynamic debris.
Accelerate with SIMD and Memory Pools
Box3D includes vectorized math kernels for AABB tests and dot products in src/simd.c. SIMD is enabled by default via BOX3D_ENABLE_SIMD in the CMake configuration, but verify it is active for your target platform:
cmake -DBOX3D_DISABLE_SIMD=OFF ..
Pre-allocate memory pools to avoid runtime reallocations when bodies enter the simulation. Set the b3Capacity struct in b3WorldDef:
def.capacity.dynamicBodyCount = 2000; // Reserve for 2k dynamic bodies
def.capacity.dynamicShapeCount = 4000; // Reserve for shape count
def.capacity.contactCount = 8000; // Enough contacts for dense stacks
This initialization in b3CreateWorld (src/physics_world.c line 56) ensures bit-sets and arrays are sized once, improving cache locality.
Configure Large Worlds and Solver Precision
For worlds spanning distances greater than ±10⁶ meters, enable double-precision positioning to keep broad-phase quantization error bounded. This mode stores world positions as doubles while keeping the dynamic tree in floats for speed:
set(BOX3D_DOUBLE_PRECISION ON) # In CMakeLists.txt
Documentation in docs/large_worlds.md explains the trade-offs. For standard piles, keep solver iterations modest (6-8 instead of the default 10) by lowering contactHertz to reduce per-step CPU usage.
Profile Bottlenecks with Built-in Timers
Identify performance constraints using the internal profiling structure updated in b3World_Step (src/physics_world.c line 332). Access timing data after each step:
b3World_Step(worldId, timeStep);
printf("Pairs: %.2f ms, Collide: %.2f ms, Solve: %.2f ms\n",
world->profile.pairs, world->profile.collide, world->profile.solve);
This breaks down time spent in broad-phase pair generation (b3UpdateBroadPhasePairs), narrow-phase collision (b3CollideTask), and constraint solving, allowing targeted optimization.
Summary
- Enable multithreading by setting
workerCountinb3WorldDefto utilizeb3ParallelForfor broad-phase and narrow-phase tasks. - Tune contact recycling via
contactRecycleDistanceand warm-starting to avoid redundant SAT calculations inb3CollideTask. - Use collision filtering with
categoryBitsandmaskBitsto skip unnecessary tests inb3ShouldShapesCollide. - Leverage SIMD instructions in
src/simd.cfor vectorized AABB and math operations. - Pre-allocate capacity for bodies, shapes, and contacts to eliminate runtime memory churn.
- Profile with
world->profileto identify whether broad-phase, collision, or solving dominates frame time.
Frequently Asked Questions
How many worker threads should I configure for Box3D?
Start with the number of physical cores on your CPU, typically between 4 and 16. The b3ParallelFor implementation in src/parallel_for.c efficiently distributes work for broad-phase updates and contact generation, but excessive threads may cause contention. Profile with world->profile to ensure the overhead of task splitting does not exceed the gains from parallel execution.
What is contact recycling and why does it improve stack performance?
Contact recycling reuses existing contact manifolds from the previous frame when bodies remain within recycleDistance, avoiding expensive SAT recomputation. In src/physics_world.c, the b3CollideTask function checks angularDistance and distSquared against the threshold before regenerating contacts. For stable piles where bodies barely move, this skips thousands of redundant collision calculations per step.
When should I enable double-precision mode for large piles?
Enable BOX3D_DOUBLE_PRECISION in CMake when your simulation world extends beyond approximately 10⁶ meters in any dimension. According to docs/large_worlds.md, this keeps the broad-phase tree (src/broad_phase.c) fast while preventing numerical drift in body positions. For smaller scenes, single-precision maintains better cache performance and SIMD utilization.
How can I reduce solver iteration costs without destabilizing stacks?
Lower the contactHertz value in b3WorldDef to reduce the number of constraint solver iterations per step (e.g., from 10 to 6). Combine this with enableWarmStarting = true to reuse previous frame impulses as initial guesses. This configuration minimizes CPU usage while maintaining stability for resting contacts, as the warm-starting reduces the iterations needed for convergence.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →