# How to Optimize Box3D Performance for Large Piles of Rigid Bodies

> Boost Box3D performance for hundreds of rigid bodies. Learn to optimize multithreading, contact recycling, collision filtering, SIMD, and pre-allocation for faster simulations.

- Repository: [Erin Catto/box3d](https://github.com/erincatto/box3d)
- Tags: performance
- Published: 2026-08-01

---

**To optimize Box3D performance for large piles of rigid bodies, enable multithreading via `workerCount`, tune contact recycling distances, use collision filtering, and leverage SIMD acceleration while pre-allocating capacity to avoid runtime allocations.**

Simulating hundreds of interacting rigid bodies in real-time requires careful tuning of the physics engine parameters. The erincatto/box3d repository provides specific architectural optimizations—ranging from broad-phase dynamic trees to contact recycling—that allow you to optimize Box3D performance for dense piles without sacrificing stability. By configuring the world definition and leveraging parallel task scheduling, you can maintain high frame rates even with complex stacking scenarios.

## Parallelize with Multithreading and Task Scheduling

Box3D splits simulation work across CPU cores using a task-based parallelism model. The engine automatically divides broad-phase pair generation and narrow-phase collision detection into independent chunks processed via `b3ParallelFor` in [`src/parallel_for.c`](https://github.com/erincatto/box3d/blob/main/src/parallel_for.c).

To enable multithreading, set the `workerCount` field in `b3WorldDef` before creating the world:

```c
b3WorldDef def = b3DefaultWorldDef();
def.workerCount = 8;  // Use 8 worker threads
b3WorldId world = b3CreateWorld(&def);

```

The scheduler distributes moved proxies across threads in `b3UpdateBroadPhasePairs` ([`src/broad_phase.c`](https://github.com/erincatto/box3d/blob/main/src/broad_phase.c) line 998) and parallelizes the contact generation in `b3Collide` ([`src/physics_world.c`](https://github.com/erincatto/box3d/blob/main/src/physics_world.c) line 815). Scaling is roughly linear with core count until memory bandwidth limits are reached.

## Reduce SAT Overhead with Contact Recycling

For stable piles, Box3D reuses existing contact manifolds instead of recomputing them every step via the Separating Axis Theorem (SAT). This recycling logic in `b3CollideTask` ([`src/physics_world.c`](https://github.com/erincatto/box3d/blob/main/src/physics_world.c) line 558) checks if bodies moved less than `recycleDistance` since the last frame.

Tune the recycling tolerance for tight stacks by adjusting `contactRecycleDistance`:

```c
def.contactRecycleDistance = 0.02f;  // Default is ~0.01m; increase for stable stacks

```

Combine this with warm-starting to reduce solver iteration costs:

```c
def.enableWarmStarting = true;
def.contactHertz = 20.0f;  // Lower frequency reduces iteration count (~6 instead of 10)

```

## Filter Collision Pairs Early

Prevent unnecessary narrow-phase tests using bit-mask collision filtering. The filter check occurs in `b3ShouldShapesCollide` ([`src/contact.c`](https://github.com/erincatto/box3d/blob/main/src/contact.c) line 70), which runs during broad-phase pair generation.

Assign category and mask bits to skip interactions between non-colliding groups:

```c
b3ShapeDef terrainDef = b3DefaultShapeDef();
terrainDef.filter.categoryBits = 0x0001;  // Ground category
terrainDef.filter.maskBits = 0xFFFE;      // Collides with everything except debris (0x0002)
b3ShapeId terrain = b3CreateHullShape(bodyId, &terrainDef, &terrainHull);

```

This eliminates SAT calls for pairs that never interact, significantly reducing CPU load in scenes with static terrain and dynamic debris.

## Accelerate with SIMD and Memory Pools

Box3D includes vectorized math kernels for AABB tests and dot products in [`src/simd.c`](https://github.com/erincatto/box3d/blob/main/src/simd.c). SIMD is enabled by default via `BOX3D_ENABLE_SIMD` in the CMake configuration, but verify it is active for your target platform:

```bash
cmake -DBOX3D_DISABLE_SIMD=OFF ..

```

Pre-allocate memory pools to avoid runtime reallocations when bodies enter the simulation. Set the `b3Capacity` struct in `b3WorldDef`:

```c
def.capacity.dynamicBodyCount = 2000;   // Reserve for 2k dynamic bodies
def.capacity.dynamicShapeCount = 4000;  // Reserve for shape count
def.capacity.contactCount = 8000;       // Enough contacts for dense stacks

```

This initialization in `b3CreateWorld` ([`src/physics_world.c`](https://github.com/erincatto/box3d/blob/main/src/physics_world.c) line 56) ensures bit-sets and arrays are sized once, improving cache locality.

## Configure Large Worlds and Solver Precision

For worlds spanning distances greater than ±10⁶ meters, enable double-precision positioning to keep broad-phase quantization error bounded. This mode stores world positions as doubles while keeping the dynamic tree in floats for speed:

```cmake
set(BOX3D_DOUBLE_PRECISION ON)  # In CMakeLists.txt

```

Documentation in [`docs/large_worlds.md`](https://github.com/erincatto/box3d/blob/main/docs/large_worlds.md) explains the trade-offs. For standard piles, keep solver iterations modest (6-8 instead of the default 10) by lowering `contactHertz` to reduce per-step CPU usage.

## Profile Bottlenecks with Built-in Timers

Identify performance constraints using the internal profiling structure updated in `b3World_Step` ([`src/physics_world.c`](https://github.com/erincatto/box3d/blob/main/src/physics_world.c) line 332). Access timing data after each step:

```c
b3World_Step(worldId, timeStep);
printf("Pairs: %.2f ms, Collide: %.2f ms, Solve: %.2f ms\n",
       world->profile.pairs, world->profile.collide, world->profile.solve);

```

This breaks down time spent in broad-phase pair generation (`b3UpdateBroadPhasePairs`), narrow-phase collision (`b3CollideTask`), and constraint solving, allowing targeted optimization.

## Summary

- **Enable multithreading** by setting `workerCount` in `b3WorldDef` to utilize `b3ParallelFor` for broad-phase and narrow-phase tasks.
- **Tune contact recycling** via `contactRecycleDistance` and warm-starting to avoid redundant SAT calculations in `b3CollideTask`.
- **Use collision filtering** with `categoryBits` and `maskBits` to skip unnecessary tests in `b3ShouldShapesCollide`.
- **Leverage SIMD** instructions in [`src/simd.c`](https://github.com/erincatto/box3d/blob/main/src/simd.c) for vectorized AABB and math operations.
- **Pre-allocate capacity** for bodies, shapes, and contacts to eliminate runtime memory churn.
- **Profile with** `world->profile` to identify whether broad-phase, collision, or solving dominates frame time.

## Frequently Asked Questions

### How many worker threads should I configure for Box3D?

Start with the number of physical cores on your CPU, typically between 4 and 16. The `b3ParallelFor` implementation in [`src/parallel_for.c`](https://github.com/erincatto/box3d/blob/main/src/parallel_for.c) efficiently distributes work for broad-phase updates and contact generation, but excessive threads may cause contention. Profile with `world->profile` to ensure the overhead of task splitting does not exceed the gains from parallel execution.

### What is contact recycling and why does it improve stack performance?

Contact recycling reuses existing contact manifolds from the previous frame when bodies remain within `recycleDistance`, avoiding expensive SAT recomputation. In [`src/physics_world.c`](https://github.com/erincatto/box3d/blob/main/src/physics_world.c), the `b3CollideTask` function checks `angularDistance` and `distSquared` against the threshold before regenerating contacts. For stable piles where bodies barely move, this skips thousands of redundant collision calculations per step.

### When should I enable double-precision mode for large piles?

Enable `BOX3D_DOUBLE_PRECISION` in CMake when your simulation world extends beyond approximately 10⁶ meters in any dimension. According to [`docs/large_worlds.md`](https://github.com/erincatto/box3d/blob/main/docs/large_worlds.md), this keeps the broad-phase tree ([`src/broad_phase.c`](https://github.com/erincatto/box3d/blob/main/src/broad_phase.c)) fast while preventing numerical drift in body positions. For smaller scenes, single-precision maintains better cache performance and SIMD utilization.

### How can I reduce solver iteration costs without destabilizing stacks?

Lower the `contactHertz` value in `b3WorldDef` to reduce the number of constraint solver iterations per step (e.g., from 10 to 6). Combine this with `enableWarmStarting = true` to reuse previous frame impulses as initial guesses. This configuration minimizes CPU usage while maintaining stability for resting contacts, as the warm-starting reduces the iterations needed for convergence.