How Modly Manages GPU Memory When Unloading and Reloading Models: A Deep Dive into the Source Code
Modly uses a layered memory management strategy that combines Python garbage collection, PyTorch CUDA cache clearing, and OS-level working set adjustments to fully release GPU memory when unloading models, then restores models on-demand through registry-managed loading.
Efficient GPU memory management is critical for serving multiple large machine learning models from a single instance. The Modly inference server (lightningpixel/modly) implements a deterministic lifecycle for model weights that ensures VRAM is aggressively reclaimed when models are inactive and reliably reallocated upon request.
The Unload Pipeline: From HTTP Request to VRAM Release
When a model is no longer needed, Modly triggers a cascading sequence of cleanup operations that target both Python heap memory and GPU allocations.
Single Model Unload via REST API
The process begins at the FastAPI endpoint POST /model/unload/{model_id} defined in api/routers/model.py (lines 100-108). This route invokes the generator's unload() method:
curl -X POST http://localhost:8000/model/unload/sf3d
The core implementation resides in BaseGenerator.unload() within api/services/generators/base.py (lines 77-98). This method performs three critical steps:
- Reference clearing: Sets
self._model = Noneto remove the Python reference to the model weights. - Garbage collection: Calls
gc.collect()to force immediate freeing of RAM. - CUDA cache purge: If CUDA is available, executes
torch.cuda.empty_cache()to release GPU allocations back to the driver.
On Windows systems, the implementation includes an additional OS-specific optimization using ctypes to call SetProcessWorkingSetSizeEx, forcing the operating system to shrink the process working set and drop cached pages.
Global Unload for Complete Memory Release
For scenarios requiring total VRAM liberation—such as switching between incompatible models or preparing for a memory-intensive operation—Modly provides the POST /model/unload-all endpoint:
curl -X POST http://localhost:8000/model/unload-all
This endpoint triggers GeneratorRegistry.unload_all() in api/services/generator_registry.py (lines 41-47). The registry iterates over every registered generator and calls unload() on each instance. For extensions running in subprocesses (instances of ExtensionProcess), it instead calls stop() to terminate the child process. After the loop completes, the system runs gc.collect() again to ensure all circular references are broken.
OS-Level Memory Reclamation on Windows
Following a bulk unload operation, Modly executes platform-specific memory reclamation logic in api/routers/model.py (lines 86-96):
import ctypes, sys
if sys.platform == "win32":
k32 = ctypes.windll.kernel32
k32.SetProcessWorkingSetSizeEx(k32.GetCurrentProcess(), -1, -1, 0)
This Windows-specific call forces the operating system to immediately discard cached pages from the process working set, ensuring that freed VRAM becomes available to subsequent GPU allocations without delay.
Reloading Models and Restoring GPU State
Modly's architecture supports on-demand reloading, allowing the server to maintain a minimal memory footprint while remaining responsive to inference requests.
On-Demand Loading via the Registry
When a model is requested—typically via status checks or inference calls—the GeneratorRegistry.get_active() method in api/services/generator_registry.py (lines 42-55) validates the model state:
if not gen.is_loaded():
if not gen.is_downloaded():
gen._auto_download()
gen.load()
If gen.is_loaded() returns False, the registry checks for missing artifacts via is_downloaded(), triggers auto-download if necessary, and then invokes the concrete load() implementation. Each extension provides its own load() method (typically in extensions/<name>/generator.py) that handles transferring weights to the GPU, often using patterns like torch.load(..., map_location='cuda').
Automatic Model Switching
Modly automates the unload/reload cycle when switching between models through the POST /model/switch endpoint:
curl -X POST http://localhost:8000/model/switch -d '{"model_id":"mv-3d"}' -H "Content-Type: application/json"
The GeneratorRegistry.switch_model() method unloads the previously active generator before updating the active ID, ensuring that VRAM usage remains bounded to a single model at a time during transitions.
Key Source Files and Implementation Details
Understanding Modly's memory management requires familiarity with these specific modules:
api/services/generators/base.py: DefinesBaseGenerator.unload(), the core GPU-memory-release logic including CUDA cache emptying and Windows working set adjustments.api/services/generator_registry.py: Manages the lifecycle of all generators, implements bulk unload viaunload_all(), and orchestrates on-demand loading throughget_active().api/routers/model.py: Exposes REST endpoints (/unload,/unload-all) that trigger unload actions and contains the Windows-specific OS-level memory reclamation.- Extension-specific
generator.py: Contains concreteload()implementations that move model weights onto the GPU using framework-specific APIs (PyTorch, ONNX Runtime, etc.). api/services/extension_process.py: Handles subprocess-based extensions; itsstop()method terminates child processes during unload operations for complete memory isolation.
Summary
- Reference hygiene: Modly explicitly nullifies model references (
self._model = None) before invoking garbage collection to ensure immediate Python memory freeing. - CUDA optimization: Every unload operation triggers
torch.cuda.empty_cache()to release GPU allocations back to the driver. - Windows enhancements: On Windows platforms, Modly calls
SetProcessWorkingSetSizeExto force the OS to drop cached pages and return physical memory. - Registry orchestration: The
GeneratorRegistryclass centralizes unload/load coordination, supporting both individual and bulk operations across diverse model backends. - Lazy loading: Models reload automatically via
get_active()when accessed, minimizing idle VRAM consumption.
Frequently Asked Questions
How do I completely free all GPU memory in Modly?
Send a POST request to the /model/unload-all endpoint. This triggers GeneratorRegistry.unload_all(), which iterates through all registered generators, calls unload() on each (or stop() for subprocess extensions), runs garbage collection, and executes Windows-specific working set reduction if applicable.
Why does Modly include Windows-specific memory handling?
Windows typically retains freed memory in a process working set for potential reuse. Modly calls SetProcessWorkingSetSizeEx in api/routers/model.py to force the OS to discard these cached pages immediately, ensuring that VRAM becomes available to other processes or subsequent model loads without delay.
What happens when I request inference from an unloaded model?
The GeneratorRegistry.get_active() method automatically detects the unloaded state via is_loaded(). If false, it verifies model artifacts, auto-downloads missing files if necessary, and invokes the extension's concrete load() method, which reallocates the model weights on the GPU before processing the request.
Can I manually reload a specific model without calling the API?
Yes. You can interact with the registry directly in Python:
from api.services.generator_registry import generator_registry
gen = generator_registry.get_active()
if not gen.is_loaded():
gen.load() # Concrete implementation moves weights to GPU
This pattern is useful for programmatic management or custom scheduling logic outside the standard REST interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →