Modly Memory Management Strategy: How MemoryIndicator Tracks GPU and RAM Usage

Modly uses three coordinated mechanisms—VRAM budgeting per generation, explicit model unloading APIs, and real-time RAM sampling via MemoryIndicator—to prevent out-of-memory crashes and give users live visibility into system resources.

Modly is an open-source AI image generation application built on Electron and Python. Its memory management strategy addresses both GPU (VRAM) and system RAM constraints, critical for running large diffusion models on consumer hardware. Understanding how MemoryIndicator tracks usage and how the backend enforces limits helps developers and power users optimize performance.


VRAM Budgeting: Per-Generation Limits in Performance Settings

Modly lets users cap GPU memory before each generation job. The PerformanceSection component in the settings UI exposes a dropdown for VRAM limits: 4 GB, 8 GB, 12 GB, or "no limit."

When a user selects a limit, the component stores it in React state:

// src/areas/settings/components/PerformanceSection.tsx
const [vram, setVram] = useState('8');   // Selected VRAM budget in GB

This value passes to the backend when spawning a generation subprocess. The cap prevents model initialization from exceeding available graphics memory, avoiding hard crashes on GPUs with limited VRAM.

Key file: [src/areas/settings/components/PerformanceSection.tsx](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx)


Explicit Model Unloading: Freeing GPU and System RAM

Modly provides API endpoints to reclaim memory on demand. The Model router exposes two routes for unloading:

Endpoint Purpose
POST /model/unload Unload a specific model from GPU/CPU
POST /model/unload_all Unload all active models

Implementation in api/routers/model.py:


# api/routers/model.py

@router.post("/unload")
async def unload_model(model_id: str):
    await model_service.unload(model_id)

@router.post("/unload_all")
async def unload_all_models():
    await model_service.unload_all()

The service layer's unload() methods:

  1. Release GPU buffers via PyTorch CUDA or Metal APIs
  2. Run Python's garbage collector to return RAM to the OS

macOS Apple Silicon special case: Metal-allocated memory often remains "wired" even after deallocation. The source code notes that on Apple Silicon, the subprocess is terminated entirely to guarantee memory reclamation. This design decision is documented in [arch/decisions/APPLE-SILICON-SUPPORT.md](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md).

Key file: [api/routers/model.py](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py)


MemoryIndicator: Real-Time RAM Tracking Architecture

The MemoryIndicator is a React component in the top bar that visualizes system RAM usage. It updates every 2 seconds through a multi-process IPC flow.

Renderer Process: Polling and UI Rendering

In MemoryIndicator.tsx, a useEffect hook polls the main process and updates a progress bar:

// src/shared/components/layout/MemoryIndicator.tsx
useEffect(() => {
  const tick = async () => {
    const { total, used, available } = await window.electron.system.memory();
    setMemory({ total, used, available });
  };
  
  tick();
  const id = setInterval(tick, 2000);   // 2-second polling interval
  return () => clearInterval(id);
}, []);

The component applies color-coded thresholds:

  • Green: Normal usage
  • Amber (≥75%): Warning threshold
  • Red (≥90%): Critical threshold

Main Process: Collecting System Statistics

The IPC handler in electron/main/ipc-handlers.ts (lines 630-661) implements system:memory:

// electron/main/ipc-handlers.ts
ipcMain.handle('system:memory', async () => {
  const total = os.totalmem();
  const free = os.freemem();
  const used = total - free;
  
  // macOS: mirror Activity Monitor by parsing vm_stat
  // to include wired, active, and compressed pages
  if (process.platform === 'darwin') {
    const vmStats = parseVmStat();   // Executes vm_stat, parses output
    return {
      total,
      used: vmStats.wired + vmStats.active + vmStats.compressed,
      available: total - (vmStats.wired + vmStats.active + vmStats.compressed)
    };
  }
  
  return { total, used, available: free };
});

Cross-platform behavior:

Platform Implementation Detail
Linux/Windows Uses Node.js os.totalmem() / os.freemem()
macOS Parses vm_stat output to match Activity Monitor's accounting of wired, active, and compressed memory

Key files:


Complete Memory Management Workflow

Modly's three mechanisms operate as a coordinated system:

  1. Prevention: Set VRAM limits before generation starts
  2. Intervention: Call unload APIs when memory runs low
  3. Visibility: Monitor MemoryIndicator to time interventions

Practical usage example:

// 1. Configure 8 GB VRAM budget in Settings
<PerformanceSection />   // Select "8 GB" from dropdown

// 2. Monitor RAM during heavy generation sessions
<MemoryIndicator />      // Watch for amber/red thresholds

// 3. Free resources when needed
await fetch('/api/model/unload_all', { method: 'POST' });

Summary

  • VRAM budgeting via PerformanceSection prevents GPU over-allocation by passing user-defined limits to generation subprocesses.
  • Model unloading APIs in model.py release GPU buffers and trigger garbage collection; macOS Apple Silicon requires subprocess termination for full reclamation.
  • MemoryIndicator polls every 2 seconds through ipc-handlers.ts, with macOS-specific vm_stat parsing for accurate Activity Monitor parity.
  • Color-coded thresholds (75% amber, 90% red) give immediate visual feedback for memory pressure.

Frequently Asked Questions

How do I check Modly's RAM usage in real time?

The MemoryIndicator in the top toolbar updates every 2 seconds. It displays a progress bar showing used versus total system RAM, with color shifts to amber at 75% and red at 90% usage. On macOS, it specifically mirrors Activity Monitor's accounting by including wired and compressed memory via vm_stat parsing in [electron/main/ipc-handlers.ts](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts#L630-L661).

Why does Modly terminate subprocesses on Apple Silicon instead of just unloading models?

Metal's memory allocator on Apple Silicon marks deallocated GPU buffers as "wired," meaning they remain resident and unavailable to other processes. According to [arch/decisions/APPLE-SILICON-SUPPORT.md](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md), terminating the subprocess is the only reliable way to return this memory to the system pool.

Can I automate GPU memory cleanup in Modly?

Yes. Call the Model API endpoints programmatically. POST /api/model/unload_all unloads every active model and runs garbage collection. For targeted cleanup, use POST /api/model/unload with a specific model ID. Both routes are defined in [api/routers/model.py](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py).

Where is the VRAM limit stored and how does it affect generation?

The PerformanceSection component stores the selected limit in React state as a string (e.g., '8' for 8 GB). This value passes to the backend when launching a generation subprocess. The backend uses it to constrain PyTorch's CUDA or Metal memory planning, as implemented in [src/areas/settings/components/PerformanceSection.tsx](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →