Modly Memory Management Strategy: How MemoryIndicator Tracks GPU and RAM Usage
Modly uses three coordinated mechanisms—VRAM budgeting per generation, explicit model unloading APIs, and real-time RAM sampling via MemoryIndicator—to prevent out-of-memory crashes and give users live visibility into system resources.
Modly is an open-source AI image generation application built on Electron and Python. Its memory management strategy addresses both GPU (VRAM) and system RAM constraints, critical for running large diffusion models on consumer hardware. Understanding how MemoryIndicator tracks usage and how the backend enforces limits helps developers and power users optimize performance.
VRAM Budgeting: Per-Generation Limits in Performance Settings
Modly lets users cap GPU memory before each generation job. The PerformanceSection component in the settings UI exposes a dropdown for VRAM limits: 4 GB, 8 GB, 12 GB, or "no limit."
When a user selects a limit, the component stores it in React state:
// src/areas/settings/components/PerformanceSection.tsx
const [vram, setVram] = useState('8'); // Selected VRAM budget in GB
This value passes to the backend when spawning a generation subprocess. The cap prevents model initialization from exceeding available graphics memory, avoiding hard crashes on GPUs with limited VRAM.
Key file: [src/areas/settings/components/PerformanceSection.tsx](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx)
Explicit Model Unloading: Freeing GPU and System RAM
Modly provides API endpoints to reclaim memory on demand. The Model router exposes two routes for unloading:
| Endpoint | Purpose |
|---|---|
POST /model/unload |
Unload a specific model from GPU/CPU |
POST /model/unload_all |
Unload all active models |
Implementation in api/routers/model.py:
# api/routers/model.py
@router.post("/unload")
async def unload_model(model_id: str):
await model_service.unload(model_id)
@router.post("/unload_all")
async def unload_all_models():
await model_service.unload_all()
The service layer's unload() methods:
- Release GPU buffers via PyTorch CUDA or Metal APIs
- Run Python's garbage collector to return RAM to the OS
macOS Apple Silicon special case: Metal-allocated memory often remains "wired" even after deallocation. The source code notes that on Apple Silicon, the subprocess is terminated entirely to guarantee memory reclamation. This design decision is documented in [arch/decisions/APPLE-SILICON-SUPPORT.md](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md).
Key file: [api/routers/model.py](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py)
MemoryIndicator: Real-Time RAM Tracking Architecture
The MemoryIndicator is a React component in the top bar that visualizes system RAM usage. It updates every 2 seconds through a multi-process IPC flow.
Renderer Process: Polling and UI Rendering
In MemoryIndicator.tsx, a useEffect hook polls the main process and updates a progress bar:
// src/shared/components/layout/MemoryIndicator.tsx
useEffect(() => {
const tick = async () => {
const { total, used, available } = await window.electron.system.memory();
setMemory({ total, used, available });
};
tick();
const id = setInterval(tick, 2000); // 2-second polling interval
return () => clearInterval(id);
}, []);
The component applies color-coded thresholds:
- Green: Normal usage
- Amber (≥75%): Warning threshold
- Red (≥90%): Critical threshold
Main Process: Collecting System Statistics
The IPC handler in electron/main/ipc-handlers.ts (lines 630-661) implements system:memory:
// electron/main/ipc-handlers.ts
ipcMain.handle('system:memory', async () => {
const total = os.totalmem();
const free = os.freemem();
const used = total - free;
// macOS: mirror Activity Monitor by parsing vm_stat
// to include wired, active, and compressed pages
if (process.platform === 'darwin') {
const vmStats = parseVmStat(); // Executes vm_stat, parses output
return {
total,
used: vmStats.wired + vmStats.active + vmStats.compressed,
available: total - (vmStats.wired + vmStats.active + vmStats.compressed)
};
}
return { total, used, available: free };
});
Cross-platform behavior:
| Platform | Implementation Detail |
|---|---|
| Linux/Windows | Uses Node.js os.totalmem() / os.freemem() |
| macOS | Parses vm_stat output to match Activity Monitor's accounting of wired, active, and compressed memory |
Key files:
- Renderer: [
src/shared/components/layout/MemoryIndicator.tsx](https://github.com/lightningpixel/modly/blob/main/src/shared/components/layout/MemoryIndicator.tsx) - Main process:
electron/main/ipc-handlers.ts#L630-L661
Complete Memory Management Workflow
Modly's three mechanisms operate as a coordinated system:
- Prevention: Set VRAM limits before generation starts
- Intervention: Call unload APIs when memory runs low
- Visibility: Monitor
MemoryIndicatorto time interventions
Practical usage example:
// 1. Configure 8 GB VRAM budget in Settings
<PerformanceSection /> // Select "8 GB" from dropdown
// 2. Monitor RAM during heavy generation sessions
<MemoryIndicator /> // Watch for amber/red thresholds
// 3. Free resources when needed
await fetch('/api/model/unload_all', { method: 'POST' });
Summary
- VRAM budgeting via
PerformanceSectionprevents GPU over-allocation by passing user-defined limits to generation subprocesses. - Model unloading APIs in
model.pyrelease GPU buffers and trigger garbage collection; macOS Apple Silicon requires subprocess termination for full reclamation. - MemoryIndicator polls every 2 seconds through
ipc-handlers.ts, with macOS-specificvm_statparsing for accurate Activity Monitor parity. - Color-coded thresholds (75% amber, 90% red) give immediate visual feedback for memory pressure.
Frequently Asked Questions
How do I check Modly's RAM usage in real time?
The MemoryIndicator in the top toolbar updates every 2 seconds. It displays a progress bar showing used versus total system RAM, with color shifts to amber at 75% and red at 90% usage. On macOS, it specifically mirrors Activity Monitor's accounting by including wired and compressed memory via vm_stat parsing in [electron/main/ipc-handlers.ts](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts#L630-L661).
Why does Modly terminate subprocesses on Apple Silicon instead of just unloading models?
Metal's memory allocator on Apple Silicon marks deallocated GPU buffers as "wired," meaning they remain resident and unavailable to other processes. According to [arch/decisions/APPLE-SILICON-SUPPORT.md](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md), terminating the subprocess is the only reliable way to return this memory to the system pool.
Can I automate GPU memory cleanup in Modly?
Yes. Call the Model API endpoints programmatically. POST /api/model/unload_all unloads every active model and runs garbage collection. For targeted cleanup, use POST /api/model/unload with a specific model ID. Both routes are defined in [api/routers/model.py](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py).
Where is the VRAM limit stored and how does it affect generation?
The PerformanceSection component stores the selected limit in React state as a string (e.g., '8' for 8 GB). This value passes to the backend when launching a generation subprocess. The backend uses it to constrain PyTorch's CUDA or Metal memory planning, as implemented in [src/areas/settings/components/PerformanceSection.tsx](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →