# How Modly Manages GPU Memory When Unloading and Reloading Models: A Deep Dive into the Source Code

> Explore how Modly manages GPU memory for unloading and reloading models. Discover its layered strategy using Python GC, PyTorch CUDA cache, and OS adjustments for efficient memory control.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: deep-dive
- Published: 2026-08-15

---

**Modly uses a layered memory management strategy that combines Python garbage collection, PyTorch CUDA cache clearing, and OS-level working set adjustments to fully release GPU memory when unloading models, then restores models on-demand through registry-managed loading.**

Efficient GPU memory management is critical for serving multiple large machine learning models from a single instance. The Modly inference server (`lightningpixel/modly`) implements a deterministic lifecycle for model weights that ensures VRAM is aggressively reclaimed when models are inactive and reliably reallocated upon request.

## The Unload Pipeline: From HTTP Request to VRAM Release

When a model is no longer needed, Modly triggers a cascading sequence of cleanup operations that target both Python heap memory and GPU allocations.

### Single Model Unload via REST API

The process begins at the FastAPI endpoint `POST /model/unload/{model_id}` defined in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) (lines 100-108). This route invokes the generator's `unload()` method:

```bash
curl -X POST http://localhost:8000/model/unload/sf3d

```

The core implementation resides in `BaseGenerator.unload()` within [`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py) (lines 77-98). This method performs three critical steps:

1. **Reference clearing**: Sets `self._model = None` to remove the Python reference to the model weights.
2. **Garbage collection**: Calls `gc.collect()` to force immediate freeing of RAM.
3. **CUDA cache purge**: If CUDA is available, executes `torch.cuda.empty_cache()` to release GPU allocations back to the driver.

On Windows systems, the implementation includes an additional OS-specific optimization using `ctypes` to call `SetProcessWorkingSetSizeEx`, forcing the operating system to shrink the process working set and drop cached pages.

### Global Unload for Complete Memory Release

For scenarios requiring total VRAM liberation—such as switching between incompatible models or preparing for a memory-intensive operation—Modly provides the `POST /model/unload-all` endpoint:

```bash
curl -X POST http://localhost:8000/model/unload-all

```

This endpoint triggers `GeneratorRegistry.unload_all()` in [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) (lines 41-47). The registry iterates over every registered generator and calls `unload()` on each instance. For extensions running in subprocesses (instances of `ExtensionProcess`), it instead calls `stop()` to terminate the child process. After the loop completes, the system runs `gc.collect()` again to ensure all circular references are broken.

## OS-Level Memory Reclamation on Windows

Following a bulk unload operation, Modly executes platform-specific memory reclamation logic in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) (lines 86-96):

```python
import ctypes, sys
if sys.platform == "win32":
    k32 = ctypes.windll.kernel32
    k32.SetProcessWorkingSetSizeEx(k32.GetCurrentProcess(), -1, -1, 0)

```

This Windows-specific call forces the operating system to immediately discard cached pages from the process working set, ensuring that freed VRAM becomes available to subsequent GPU allocations without delay.

## Reloading Models and Restoring GPU State

Modly's architecture supports on-demand reloading, allowing the server to maintain a minimal memory footprint while remaining responsive to inference requests.

### On-Demand Loading via the Registry

When a model is requested—typically via status checks or inference calls—the `GeneratorRegistry.get_active()` method in [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) (lines 42-55) validates the model state:

```python
if not gen.is_loaded():
    if not gen.is_downloaded():
        gen._auto_download()
    gen.load()

```

If `gen.is_loaded()` returns `False`, the registry checks for missing artifacts via `is_downloaded()`, triggers auto-download if necessary, and then invokes the concrete `load()` implementation. Each extension provides its own `load()` method (typically in `extensions/<name>/generator.py`) that handles transferring weights to the GPU, often using patterns like `torch.load(..., map_location='cuda')`.

### Automatic Model Switching

Modly automates the unload/reload cycle when switching between models through the `POST /model/switch` endpoint:

```bash
curl -X POST http://localhost:8000/model/switch -d '{"model_id":"mv-3d"}' -H "Content-Type: application/json"

```

The `GeneratorRegistry.switch_model()` method unloads the previously active generator before updating the active ID, ensuring that VRAM usage remains bounded to a single model at a time during transitions.

## Key Source Files and Implementation Details

Understanding Modly's memory management requires familiarity with these specific modules:

- **[`api/services/generators/base.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generators/base.py)**: Defines `BaseGenerator.unload()`, the core GPU-memory-release logic including CUDA cache emptying and Windows working set adjustments.
- **[`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py)**: Manages the lifecycle of all generators, implements bulk unload via `unload_all()`, and orchestrates on-demand loading through `get_active()`.
- **[`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py)**: Exposes REST endpoints (`/unload`, `/unload-all`) that trigger unload actions and contains the Windows-specific OS-level memory reclamation.
- **Extension-specific [`generator.py`](https://github.com/lightningpixel/modly/blob/main/generator.py)**: Contains concrete `load()` implementations that move model weights onto the GPU using framework-specific APIs (PyTorch, ONNX Runtime, etc.).
- **[`api/services/extension_process.py`](https://github.com/lightningpixel/modly/blob/main/api/services/extension_process.py)**: Handles subprocess-based extensions; its `stop()` method terminates child processes during unload operations for complete memory isolation.

## Summary

- **Reference hygiene**: Modly explicitly nullifies model references (`self._model = None`) before invoking garbage collection to ensure immediate Python memory freeing.
- **CUDA optimization**: Every unload operation triggers `torch.cuda.empty_cache()` to release GPU allocations back to the driver.
- **Windows enhancements**: On Windows platforms, Modly calls `SetProcessWorkingSetSizeEx` to force the OS to drop cached pages and return physical memory.
- **Registry orchestration**: The `GeneratorRegistry` class centralizes unload/load coordination, supporting both individual and bulk operations across diverse model backends.
- **Lazy loading**: Models reload automatically via `get_active()` when accessed, minimizing idle VRAM consumption.

## Frequently Asked Questions

### How do I completely free all GPU memory in Modly?

Send a `POST` request to the `/model/unload-all` endpoint. This triggers `GeneratorRegistry.unload_all()`, which iterates through all registered generators, calls `unload()` on each (or `stop()` for subprocess extensions), runs garbage collection, and executes Windows-specific working set reduction if applicable.

### Why does Modly include Windows-specific memory handling?

Windows typically retains freed memory in a process working set for potential reuse. Modly calls `SetProcessWorkingSetSizeEx` in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) to force the OS to discard these cached pages immediately, ensuring that VRAM becomes available to other processes or subsequent model loads without delay.

### What happens when I request inference from an unloaded model?

The `GeneratorRegistry.get_active()` method automatically detects the unloaded state via `is_loaded()`. If false, it verifies model artifacts, auto-downloads missing files if necessary, and invokes the extension's concrete `load()` method, which reallocates the model weights on the GPU before processing the request.

### Can I manually reload a specific model without calling the API?

Yes. You can interact with the registry directly in Python:

```python
from api.services.generator_registry import generator_registry

gen = generator_registry.get_active()
if not gen.is_loaded():
    gen.load()  # Concrete implementation moves weights to GPU

```

This pattern is useful for programmatic management or custom scheduling logic outside the standard REST interface.