# How Peng's Multi-Threaded Depth Rendering Works

> Discover how Peng's multi-threaded depth rendering parallelizes ray casting across CPU cores for faster depth map generation. Learn about chunking and performance gains.

- Repository: [Yang Zhou/peng](https://github.com/makeecat/peng)
- Tags: internals
- Published: 2026-03-06

---

**Peng renders depth maps by parallelizing per-pixel ray casting across CPU cores using Rayon's `par_chunks_mut`, automatically splitting the workload into 64-element chunks when the `use_multi_threading` flag is enabled.**

Peng is a quadrotor simulation environment that generates depth maps by casting rays from a virtual camera through every pixel to measure distances to obstacles. When configured for high-performance rendering, the simulation utilizes a **Rayon-based parallel implementation** to accelerate per-pixel calculations without manual thread pool management. The core implementation resides in `Camera::render_depth` within [[`src/lib.rs`](https://github.com/makeecat/peng/blob/main/src/lib.rs)](https://github.com/makeecat/peng/blob/main/src/lib.rs#L2918-L2970), which dynamically partitions the depth buffer across available CPU cores based on configuration flags defined in [`src/config.rs`](https://github.com/makeecat/peng/blob/main/src/config.rs).

## The Depth Rendering Pipeline

At its foundation, Peng's depth rendering operates by shooting a ray from the virtual camera through every pixel in the image, calculating the distance to the nearest obstacle or wall, and storing these distances in a flat buffer. The method prepares a rotation matrix (`rotation_camera_to_world`) that transforms ray directions from camera space to world space, along with its transposed inverse for orientation calculations.

When the resolution warrants parallel processing, the system splits the depth buffer—represented as a `Vec<f32>` of size `width × height`—into fixed-size chunks for concurrent processing.

## Multi-Threaded Implementation Details

The parallel execution path leverages Rayon to distribute work across CPU cores without explicit synchronization code.

### Chunking Strategy with par_chunks_mut

Inside `Camera::render_depth`, the implementation divides the flat depth buffer into chunks of **64 elements** (`CHUNK_SIZE = 64`). When the `use_multi_threading` parameter is true, the code invokes `rayon::prelude::ParallelIterator::par_chunks_mut` to process each chunk concurrently:

```rust
// Simplified logic from src/lib.rs Camera::render_depth
if use_multi_threading {
    buffer.par_chunks_mut(CHUNK_SIZE).enumerate().try_for_each(|(chunk_idx, chunk)| {
        let start_idx = chunk_idx * CHUNK_SIZE;
        for (i, depth) in chunk.iter_mut().enumerate() {
            let ray_idx = start_idx + i;
            // compute ray direction, cast ray, store result
            *depth = ray_cast(ray_origin, ray_direction, maze);
        }
        Ok(())
    })?;
}

```

Each thread executes identical per-pixel logic (`ray_cast`) on its assigned chunk. The `CHUNK_SIZE` of 64 balances workload granularity against parallel overhead, ensuring efficient utilization across varying CPU core counts.

### Configuration and Entry Points

Multi-threading is controlled via the `Config::use_multithreading_depth_rendering` boolean field defined in [[`src/config.rs`](https://github.com/makeecat/peng/blob/main/src/config.rs)](https://github.com/makeecat/peng/blob/main/src/config.rs#L27-L36). Users can optionally limit the global Rayon thread pool size using `max_render_threads` to prevent resource contention during simulation.

The rendering call in [[`src/main.rs`](https://github.com/makeecat/peng/blob/main/src/main.rs)](https://github.com/makeecat/peng/blob/main/src/main.rs#L155-L162) passes this configuration flag directly to the camera:

```rust
camera.render_depth(
    &quad.position,
    &quad.orientation,
    &maze,
    config.use_multithreading_depth_rendering,
)?;

```

### Single-Threaded Fallback

When multi-threading is disabled, the implementation falls back to a simple sequential loop over the pixel indices. This path uses identical `ray_cast` logic but processes pixels `0..total_pixels` in a single thread, avoiding any parallelization overhead for small renders or debugging scenarios.

## Configuring Multi-Threaded Rendering

Enable parallel depth rendering via the YAML configuration file or direct API calls.

### YAML Configuration

Set the flag in your [`config/quad.yaml`](https://github.com/makeecat/peng/blob/main/config/quad.yaml):

```yaml
use_multithreading_depth_rendering: true
max_render_threads: 8  # Optional: caps Rayon thread pool

```

### Programmatic Control

Force multi-threading for specific camera instances in Rust:

```rust
let mut camera = Camera::new((800, 600), 60.0, 0.1, 100.0);
camera.render_depth(&quad_pos, &quad_ori, &maze, true)?;

```

### Measuring Performance Gains

Benchmark the implementation to verify speedup on your hardware:

```rust
use std::time::Instant;

let start = Instant::now();
camera.render_depth(&pos, &ori, &maze, true)?;
println!("Multi-threaded: {:.2?}", start.elapsed());

let start = Instant::now();
camera.render_depth(&pos, &ori, &maze, false)?;
println!("Single-threaded: {:.2?}", start.elapsed());

```

## Performance Characteristics

Rayon provides **dynamic work stealing** that automatically balances chunks across threads as they complete, eliminating the need for manual load balancing. The fixed `CHUNK_SIZE` of 64 ensures that overhead from thread synchronization remains minimal while providing enough granularity to saturate multiple cores. Because the depth buffer is split into independent chunks, the implementation scales linearly with CPU core count until memory bandwidth becomes the limiting factor.

## Summary

- **Peng** accelerates depth map generation using **Rayon's `par_chunks_mut`** to parallelize per-pixel ray casting across CPU cores.
- The **64-element chunk size** balances parallel granularity with synchronization overhead in `Camera::render_depth` ([`src/lib.rs`](https://github.com/makeecat/peng/blob/main/src/lib.rs)).
- Multi-threading is toggled via **`Config::use_multithreading_depth_rendering`** in [`src/config.rs`](https://github.com/makeecat/peng/blob/main/src/config.rs) and passed through the simulation loop in [`src/main.rs`](https://github.com/makeecat/peng/blob/main/src/main.rs).
- The implementation automatically falls back to single-threaded execution when the flag is disabled, using identical ray casting logic.
- Thread pool size can be constrained using **`max_render_threads`** to manage system resources during simulation.

## Frequently Asked Questions

### What triggers multi-threaded rendering in Peng?

Multi-threaded rendering activates when the `use_multi_threading` parameter passed to `Camera::render_depth` is set to `true`, typically controlled by the `use_multithreading_depth_rendering` configuration flag. Unlike automatic resolution-based switching, this explicit boolean flag allows developers to force single-threaded execution for debugging or low-power scenarios regardless of image size.

### How does Peng divide work among CPU cores?

Peng splits the flat depth buffer (`Vec<f32>`) into contiguous chunks of 64 elements using `par_chunks_mut` from the Rayon library. Each chunk processes independently on separate threads, with Rayon dynamically scheduling work to minimize idle time through work-stealing algorithms.

### Can I limit the number of threads used for depth rendering?

Yes. Set the `max_render_threads` field in your configuration YAML or initialize the Rayon global thread pool with a specific size before running the simulation. This prevents Peng from consuming all available CPU cores, leaving resources for other simulation components or system processes.

### Is the ray casting logic different between single and multi-threaded modes?

No. Both execution paths use identical `ray_cast` implementations and mathematical calculations for ray direction and obstacle intersection. The only difference is the iteration strategy: `par_chunks_mut` for parallel execution versus a standard `for` loop for sequential processing, ensuring consistent depth results regardless of thread count.