# RenodX Performance Considerations: Optimizing Frame Rate, Swap-Chain Resizing, and HDR Handling

> Optimize RenodX performance by addressing CPU overhead from FPS limiting, GPU stalls from swap-chain resizing, and driver latency with HDR handling.

- Repository: [Carlos Lopez/renodx](https://github.com/clshortfuse/renodx)
- Tags: performance
- Published: 2026-09-06

---

**RenodX performance considerations center on CPU overhead from the optional FPS limiter's busy-spin loop, GPU stalls caused by swap-chain buffer reallocations, and driver latency from HDR capability queries.**

RenodX is a ReShade add-on that intercepts DirectX 12 and Vulkan swap-chains to modify formats, color spaces, and frame-rate limits. While the core functionality adds minimal overhead, several specific code paths in [`src/utils/swapchain.hpp`](https://github.com/clshortfuse/renodx/blob/main/src/utils/swapchain.hpp) can measurably impact runtime performance. This guide breaks down the performance-critical sections and provides optimization strategies based on the actual source implementation.

---

## FPS Limiter: CPU Spin and Latency Tracking

The **FPS limiter** is the most CPU-intensive optional feature in RenodX. When enabled via the static `fps_limit` variable, it runs on every `present` call between lines 717–805 in [`src/utils/swapchain.hpp`](https://github.com/clshortfuse/renodx/blob/main/src/utils/swapchain.hpp).

### How the Limiter Works

The implementation uses a **hybrid sleep-and-spin approach**:

1. Calculates time remaining until the next frame boundary
2. Attempts `std::this_thread::sleep_for` for coarse timing
3. Enters a **busy-spin loop** (`YieldProcessor()`) for fine-grained precision

```cpp
// Set a 60 fps cap
renodx::utils::swapchain::fps_limit = 60.f;

```

The spin duration is dynamically adjusted based on the **worst 1% observed latency** (P99), tracked in a `wait_latency_history` deque.

### Latency History Overhead

The history mechanism adds measurable CPU cost:

- Stores up to **1,000 latency samples** (`MAX_LATENCY_HISTORY_SIZE`)
- **Sorts the deque each frame** to compute the P99 percentile
- May trigger reallocations during resize operations

This creates O(N log N) work per present call when the limiter is active.

### Optimization

Disable the limiter completely when not needed:

```cpp
renodx::utils::swapchain::fps_limit = 0.f;  // Zero CPU overhead

```

The limiter uses a **lock-free read path** (no mutex acquisition), but the busy-spin alone can increase CPU usage significantly on high-refresh displays.

---

## Swap-Chain Resizing: GPU Stalls and Stutter

Changing the back-buffer **format** or **color-space** triggers `ResizeBuffer` (lines 492–525), which invokes the heavyweight `IDXGISwapChain4::ResizeBuffers` DXGI operation.

### Performance Impact

| Operation | Cost | Source Location |
|-----------|------|-----------------|
| `ResizeBuffers` call | Buffer reallocation, GPU stall | `swapchain.hpp:L492-L525` |
| `DXGI_ERROR_INVALID_CALL` handling | Early exit with logging | `swapchain.hpp:L537-L548` |
| `ChangeColorSpace` follow-up | Additional `SetColorSpace1` call | `swapchain.hpp:L557-L566` |

### Best Practices

- **Batch format/color-space changes** together to minimize resize calls
- **Avoid toggling HDR on-the-fly** — each toggle triggers a full resize
- Check current format before calling; the code has an early-exit optimization

```cpp
// Example: Force 16-bit float HDR back-buffer (call once, not per-frame)
renodx::utils::swapchain::ResizeBuffer(
    swapchain,
    renodx::utils::swapchain::format::r16g16b16a16_float,
    renodx::utils::swapchain::color_space::hdr10_hlg);

```

The resize path also acquires the `DeviceData` mutex, adding potential contention in multi-threaded command-list scenarios.

---

## HDR and Color-Space Queries: Driver Round-Trips

RenodX queries display HDR capabilities through Windows' **DisplayConfig APIs**, which require driver round-trips and can pause the frame pipeline.

### Slow Operations

| Function | API Used | Lines |
|----------|----------|-------|
| `GetHDRSupported` | `DisplayConfigGetDeviceInfo` | `L420-L604` |
| `GetHDREnabled` | `DisplayConfigGetDeviceInfo` | `L420-L604` |
| `SetHDREnabled` | `DisplayConfigSetDeviceInfo` | `L420-L604` |
| `GetDirectXOutputDesc1` | `IDXGIOutput6::GetDesc1` | `L136-L182` |

### Optimization Strategy

**Cache display information** after the first query:

```cpp
// Query once, reuse throughout session
auto displayInfo = renodx::utils::swapchain::GetDisplayInfo(swapchain, /*force_hdr=*/true);

if (displayInfo.hdr_supported && !displayInfo.hdr_enabled) {
    renodx::utils::swapchain::SetHDREnabled(swapchain, true);
}
// Store displayInfo — no further driver calls needed

```

HDR functions use a **`std::shared_mutex`** allowing concurrent reads but exclusive writes during `OnInitEffectRuntime` and `ChangeColorSpace` (lines 740–782).

---

## Synchronization: DeviceData and CommandListData

RenodX maintains shared state in two primary structures:

- **`DeviceData`** (lines 34–40): `effect_runtime*` set, back-buffer description, current color-space
- **`CommandListData`** (lines 41–48): Render-target tracking, dirty flags

### Locking Behavior

| Scenario | Lock Type | Performance Characteristic |
|----------|-----------|---------------------------|
| Reading back-buffer description | **Lock-free** | No contention |
| Mutating shared state | `std::unique_lock` | Blocks readers |
| HDR color-space updates | `std::unique_lock` on `DeviceData::mutex` | Present-thread only |

Contention is generally low because mutations occur primarily on the **present thread**. However, games with heavy multi-threaded command-list usage may experience bottlenecking at the `unique_lock` points.

---

## Summary of Key Performance Knobs

| Feature | Performance Impact | Recommended Setting |
|---------|------------------|---------------------|
| **`fps_limit`** | CPU spin + latency history overhead | `0.f` to disable; enable only when throttling needed |
| **`ResizeBuffer`** | GPU stalls from buffer reallocation | Call once per format change; batch with color-space updates |
| **HDR queries** | Driver round-trip latency | Cache `GetDisplayInfo` results; avoid per-frame polling |
| **`wait_latency_history`** | O(N log N) sort per frame (max 1,000) | Reduce `MAX_LATENCY_HISTORY_SIZE` if CPU-constrained |
| **`DeviceData` mutex** | Minimal contention under normal use | Profile if lock contention suspected |

---

## Frequently Asked Questions

### Does RenodX affect performance when all features are disabled?

When `fps_limit` is `0.f` and no resize or HDR operations are performed, RenodX adds negligible overhead. The present hook reads only static variables without acquiring locks, and no driver calls occur unless explicitly triggered.

### Why does the FPS limiter increase CPU usage instead of reducing it?

The limiter achieves precise frame pacing through **busy-spinning** (`YieldProcessor()`) after coarse sleep. This burns CPU cycles to hit exact frame boundaries with minimal jitter. The trade-off favors timing accuracy over power efficiency, which matters for high-refresh-rate displays.

### How can I minimize stutter when enabling HDR?

Call `SetHDREnabled` **once during initialization**, not during gameplay. HDR enablement triggers a full mode change and potential `ResizeBuffer` call. Pre-caching display capabilities with `GetDisplayInfo` and reusing the result avoids repeated driver queries.

### What causes occasional frame-time spikes with the limiter enabled?

The **latency history sorting** (up to 1,000 samples) runs every frame to compute the P99 adjustment. While typically fast, memory reallocation or cache effects can occasionally cause **GC-like spikes**. Reduce `MAX_LATENCY_HISTORY_SIZE` in [`src/utils/swapchain.hpp`](https://github.com/clshortfuse/renodx/blob/main/src/utils/swapchain.hpp) if this occurs.