# How Modly Handles GPU Memory Management During Generation: 3 Core Mechanisms Explained

> Discover how Modly manages GPU memory during generation with its 3 core mechanisms: hardware detection, explicit cache clearing, and process restarts. Optimize your generation workflow today.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: internals
- Published: 2026-08-20

---

**Modly handles GPU memory during generation through a three-layer system: hardware capability detection before model loading, explicit cache clearing after each generation, and process-level bridge restarts when memory cannot be safely reclaimed.**

Stable diffusion and 3-D generation models are notorious for consuming large amounts of **GPU memory**, often causing out-of-memory crashes or degraded performance. According to the **lightningpixel/modly** source code, Modly addresses this with a coordinated architecture spanning the Electron main process and Python backend. This article breaks down exactly how GPU memory management works in Modly, with references to specific implementation files and methods.

---

## GPU Capability Detection Sets the Foundation

Before any model touches GPU memory, Modly determines what hardware is available and which **accelerator** to use.

In [`electron/main/ipc-handlers.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts), the `detectGpuInfo()` function (lines 42-82) runs system queries to establish the execution environment:

```typescript
// Detector implementation in electron/main/ipc-handlers.ts
const { sm, cudaVersion, accelerator } = await detectGpuInfo();

```

This function executes `nvidia-smi` on NVIDIA systems or falls back to **MPS** (Metal Performance Shaders) on Apple Silicon. It returns a structured object containing:

- `sm` — the GPU's compute capability
- `cudaVersion` — mapped driver version
- `accelerator` — one of `'cuda' | 'mps' | 'cpu'`

The detected `accelerator` value propagates through `runExtensionSetup()` (lines 110-117 in the same file) and becomes a launch argument for the Python bridge. This ensures the model loads onto the correct device from the start, preventing unnecessary memory copying or failed CUDA allocations.

---

## Explicit GPU Cache Clearing After Generation

Once generation completes, Modly aggressively releases GPU memory through the **generator base class**.

### The unload() Method

Every generator in Modly inherits from `BaseGenerator` in [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py). Its `unload()` method (lines 77-86) implements a three-step cleanup:

```python
def unload(self):
    self._model = None          # 1. Drop the model reference

    import gc, torch
    gc.collect()                # 2. Force Python garbage collection

    if torch.cuda.is_available():
        torch.cuda.empty_cache()  # 3. Release cached CUDA memory

```

The `torch.cuda.empty_cache()` call is critical—PyTorch retains allocated GPU memory in a pool to avoid expensive reallocation overhead. Without this explicit call, subsequent generations see inflated memory usage even after the model appears "unloaded."

### Windows-Specific Memory Reclamation

On Windows, Modly goes further. Lines 90-97 of [`base.py`](https://github.com/lightningpixel/modly/blob/main/base.py) invoke a Windows API call to trim the process working set:

```python
import sys, ctypes

if sys.platform == "win32":
    kernel32 = ctypes.windll.kernel32
    kernel32.SetProcessWorkingSetSizeEx(
        kernel32.GetCurrentProcess(),
        -1, -1, 0  # -1 signals "reduce as much as possible"

    )

```

This `SetProcessWorkingSetSizeEx` call prompts the Windows memory manager to reclaim private pages that the Python process no longer actively uses, reducing the reported memory footprint without terminating the process.

---

## Process-Level Bridge Restart for Stubborn Allocations

Some GPU memory cannot be safely freed—fragmented heaps, driver-level allocations, or third-party library caches may persist despite `empty_cache()`. Modly handles these cases by **restarting the Python bridge entirely**.

In [`electron/main/python-bridge.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/python-bridge.ts), lines 124-132 contain the restart logic:

```typescript
logger.info('[PythonBridge] Restarting to free memory…');
this.stop();   // Terminate the Python subprocess
await this.start();  // Respawn fresh process

```

This restart eliminates any dangling allocations from the previous generation's Python process. The Electron main process orchestrates this without disrupting the user experience, queueing the restart between generation requests.

---

## Hardware-to-Code Flow: How the Pieces Connect

| Stage | Component | Key File | Memory Impact |
|-------|-----------|----------|---------------|
| Detection | GPU capability probe | [`electron/main/ipc-handlers.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts) (42-82) | Prevents loading wrong accelerator |
| Loading | Model instantiation | Python generators (device-aware) | Allocates GPU tensors |
| Cleanup | Explicit cache clear | [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) (77-86) | Frees PyTorch CUDA cache |
| Reclamation | OS-level trim | [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) (90-97) | Reduces process working set |
| Safety net | Bridge restart | [`electron/main/python-bridge.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/python-bridge.ts) (124-132) | Eliminates unrecoverable leaks |

---

## Summary

Modly's **GPU memory management** combines detection, cleanup, and failsafe mechanisms:

- **Hardware detection** via `detectGpuInfo()` ensures models load on compatible accelerators
- **Explicit cache clearing** through `unload()` releases `torch.cuda` memory and trims OS working sets
- **Process restart** as a last-resort guarantees clean memory for subsequent generations

These layers prevent the out-of-memory failures common in diffusion-based 3-D generation workloads.

---

## Frequently Asked Questions

### Does Modly automatically select between CUDA and CPU?

Yes. The `detectGpuInfo()` function in [`electron/main/ipc-handlers.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts) automatically probes for NVIDIA GPUs via `nvidia-smi`, checks for MPS availability on Apple Silicon, and falls back to CPU if neither accelerator is found. This detection runs once at application startup and guides all subsequent model loading decisions.

### Why does Modly use `torch.cuda.empty_cache()` instead of relying on garbage collection?

PyTorch's CUDA memory allocator maintains a pool of cached blocks to speed up reallocation. Garbage collection only removes Python object references—it does not return GPU memory to the system. The `empty_cache()` call explicitly releases these pooled blocks, making memory available for other applications or subsequent generations with different tensor sizes.

### When does Modly restart the Python bridge?

The bridge restarts when memory cannot be reclaimed through normal cleanup, typically detected via memory threshold monitoring or explicit user action. The restart emits a log message "`[PythonBridge] Restarting to free memory…`" in [`electron/main/python-bridge.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/python-bridge.ts) before terminating and respawning the subprocess, ensuring no corrupted or fragmented state persists.