How Modly Manages GPU Memory When Unloading and Reloading Models: A Deep Dive into the Source Code

Modly uses a layered memory management strategy that combines Python garbage collection, PyTorch CUDA cache clearing, and OS-level working set adjustments to fully release GPU memory when unloading models, then restores models on-demand through registry-managed loading.

Efficient GPU memory management is critical for serving multiple large machine learning models from a single instance. The Modly inference server (lightningpixel/modly) implements a deterministic lifecycle for model weights that ensures VRAM is aggressively reclaimed when models are inactive and reliably reallocated upon request.

The Unload Pipeline: From HTTP Request to VRAM Release

When a model is no longer needed, Modly triggers a cascading sequence of cleanup operations that target both Python heap memory and GPU allocations.

Single Model Unload via REST API

The process begins at the FastAPI endpoint POST /model/unload/{model_id} defined in api/routers/model.py (lines 100-108). This route invokes the generator's unload() method:

curl -X POST http://localhost:8000/model/unload/sf3d

The core implementation resides in BaseGenerator.unload() within api/services/generators/base.py (lines 77-98). This method performs three critical steps:

  1. Reference clearing: Sets self._model = None to remove the Python reference to the model weights.
  2. Garbage collection: Calls gc.collect() to force immediate freeing of RAM.
  3. CUDA cache purge: If CUDA is available, executes torch.cuda.empty_cache() to release GPU allocations back to the driver.

On Windows systems, the implementation includes an additional OS-specific optimization using ctypes to call SetProcessWorkingSetSizeEx, forcing the operating system to shrink the process working set and drop cached pages.

Global Unload for Complete Memory Release

For scenarios requiring total VRAM liberation—such as switching between incompatible models or preparing for a memory-intensive operation—Modly provides the POST /model/unload-all endpoint:

curl -X POST http://localhost:8000/model/unload-all

This endpoint triggers GeneratorRegistry.unload_all() in api/services/generator_registry.py (lines 41-47). The registry iterates over every registered generator and calls unload() on each instance. For extensions running in subprocesses (instances of ExtensionProcess), it instead calls stop() to terminate the child process. After the loop completes, the system runs gc.collect() again to ensure all circular references are broken.

OS-Level Memory Reclamation on Windows

Following a bulk unload operation, Modly executes platform-specific memory reclamation logic in api/routers/model.py (lines 86-96):

import ctypes, sys
if sys.platform == "win32":
    k32 = ctypes.windll.kernel32
    k32.SetProcessWorkingSetSizeEx(k32.GetCurrentProcess(), -1, -1, 0)

This Windows-specific call forces the operating system to immediately discard cached pages from the process working set, ensuring that freed VRAM becomes available to subsequent GPU allocations without delay.

Reloading Models and Restoring GPU State

Modly's architecture supports on-demand reloading, allowing the server to maintain a minimal memory footprint while remaining responsive to inference requests.

On-Demand Loading via the Registry

When a model is requested—typically via status checks or inference calls—the GeneratorRegistry.get_active() method in api/services/generator_registry.py (lines 42-55) validates the model state:

if not gen.is_loaded():
    if not gen.is_downloaded():
        gen._auto_download()
    gen.load()

If gen.is_loaded() returns False, the registry checks for missing artifacts via is_downloaded(), triggers auto-download if necessary, and then invokes the concrete load() implementation. Each extension provides its own load() method (typically in extensions/<name>/generator.py) that handles transferring weights to the GPU, often using patterns like torch.load(..., map_location='cuda').

Automatic Model Switching

Modly automates the unload/reload cycle when switching between models through the POST /model/switch endpoint:

curl -X POST http://localhost:8000/model/switch -d '{"model_id":"mv-3d"}' -H "Content-Type: application/json"

The GeneratorRegistry.switch_model() method unloads the previously active generator before updating the active ID, ensuring that VRAM usage remains bounded to a single model at a time during transitions.

Key Source Files and Implementation Details

Understanding Modly's memory management requires familiarity with these specific modules:

  • api/services/generators/base.py: Defines BaseGenerator.unload(), the core GPU-memory-release logic including CUDA cache emptying and Windows working set adjustments.
  • api/services/generator_registry.py: Manages the lifecycle of all generators, implements bulk unload via unload_all(), and orchestrates on-demand loading through get_active().
  • api/routers/model.py: Exposes REST endpoints (/unload, /unload-all) that trigger unload actions and contains the Windows-specific OS-level memory reclamation.
  • Extension-specific generator.py: Contains concrete load() implementations that move model weights onto the GPU using framework-specific APIs (PyTorch, ONNX Runtime, etc.).
  • api/services/extension_process.py: Handles subprocess-based extensions; its stop() method terminates child processes during unload operations for complete memory isolation.

Summary

  • Reference hygiene: Modly explicitly nullifies model references (self._model = None) before invoking garbage collection to ensure immediate Python memory freeing.
  • CUDA optimization: Every unload operation triggers torch.cuda.empty_cache() to release GPU allocations back to the driver.
  • Windows enhancements: On Windows platforms, Modly calls SetProcessWorkingSetSizeEx to force the OS to drop cached pages and return physical memory.
  • Registry orchestration: The GeneratorRegistry class centralizes unload/load coordination, supporting both individual and bulk operations across diverse model backends.
  • Lazy loading: Models reload automatically via get_active() when accessed, minimizing idle VRAM consumption.

Frequently Asked Questions

How do I completely free all GPU memory in Modly?

Send a POST request to the /model/unload-all endpoint. This triggers GeneratorRegistry.unload_all(), which iterates through all registered generators, calls unload() on each (or stop() for subprocess extensions), runs garbage collection, and executes Windows-specific working set reduction if applicable.

Why does Modly include Windows-specific memory handling?

Windows typically retains freed memory in a process working set for potential reuse. Modly calls SetProcessWorkingSetSizeEx in api/routers/model.py to force the OS to discard these cached pages immediately, ensuring that VRAM becomes available to other processes or subsequent model loads without delay.

What happens when I request inference from an unloaded model?

The GeneratorRegistry.get_active() method automatically detects the unloaded state via is_loaded(). If false, it verifies model artifacts, auto-downloads missing files if necessary, and invokes the extension's concrete load() method, which reallocates the model weights on the GPU before processing the request.

Can I manually reload a specific model without calling the API?

Yes. You can interact with the registry directly in Python:

from api.services.generator_registry import generator_registry

gen = generator_registry.get_active()
if not gen.is_loaded():
    gen.load()  # Concrete implementation moves weights to GPU

This pattern is useful for programmatic management or custom scheduling logic outside the standard REST interface.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →