# Modly Memory Management Strategy: How MemoryIndicator Tracks GPU and RAM Usage

> Discover Modly's memory management strategy. Learn how MemoryIndicator tracks GPU and RAM usage to prevent crashes and provide live system resource visibility.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: internals
- Published: 2026-08-20

---

**Modly uses three coordinated mechanisms—VRAM budgeting per generation, explicit model unloading APIs, and real-time RAM sampling via MemoryIndicator—to prevent out-of-memory crashes and give users live visibility into system resources.**

Modly is an open-source AI image generation application built on Electron and Python. Its memory management strategy addresses both GPU (VRAM) and system RAM constraints, critical for running large diffusion models on consumer hardware. Understanding how `MemoryIndicator` tracks usage and how the backend enforces limits helps developers and power users optimize performance.

---

## VRAM Budgeting: Per-Generation Limits in Performance Settings

Modly lets users cap GPU memory before each generation job. The `PerformanceSection` component in the settings UI exposes a dropdown for **VRAM limits**: 4 GB, 8 GB, 12 GB, or "no limit."

When a user selects a limit, the component stores it in React state:

```tsx
// src/areas/settings/components/PerformanceSection.tsx
const [vram, setVram] = useState('8');   // Selected VRAM budget in GB

```

This value passes to the backend when spawning a generation **subprocess**. The cap prevents model initialization from exceeding available graphics memory, avoiding hard crashes on GPUs with limited VRAM.

**Key file:** [[`src/areas/settings/components/PerformanceSection.tsx`](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx)](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx)

---

## Explicit Model Unloading: Freeing GPU and System RAM

Modly provides **API endpoints** to reclaim memory on demand. The Model router exposes two routes for unloading:

| Endpoint | Purpose |
|----------|---------|
| `POST /model/unload` | Unload a specific model from GPU/CPU |
| `POST /model/unload_all` | Unload all active models |

**Implementation in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py):**

```python

# api/routers/model.py

@router.post("/unload")
async def unload_model(model_id: str):
    await model_service.unload(model_id)

@router.post("/unload_all")
async def unload_all_models():
    await model_service.unload_all()

```

The service layer's `unload()` methods:

1. **Release GPU buffers** via PyTorch CUDA or Metal APIs
2. **Run Python's garbage collector** to return RAM to the OS

**macOS Apple Silicon special case:** Metal-allocated memory often remains "wired" even after deallocation. The source code notes that on Apple Silicon, the subprocess is **terminated entirely** to guarantee memory reclamation. This design decision is documented in [[`arch/decisions/APPLE-SILICON-SUPPORT.md`](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md)](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md).

**Key file:** [[`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py)](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py)

---

## MemoryIndicator: Real-Time RAM Tracking Architecture

The **MemoryIndicator** is a React component in the top bar that visualizes system RAM usage. It updates every **2 seconds** through a multi-process IPC flow.

### Renderer Process: Polling and UI Rendering

In [`MemoryIndicator.tsx`](https://github.com/lightningpixel/modly/blob/main/MemoryIndicator.tsx), a `useEffect` hook polls the main process and updates a progress bar:

```tsx
// src/shared/components/layout/MemoryIndicator.tsx
useEffect(() => {
  const tick = async () => {
    const { total, used, available } = await window.electron.system.memory();
    setMemory({ total, used, available });
  };
  
  tick();
  const id = setInterval(tick, 2000);   // 2-second polling interval
  return () => clearInterval(id);
}, []);

```

The component applies **color-coded thresholds**:

- **Green:** Normal usage
- **Amber (≥75%):** Warning threshold
- **Red (≥90%):** Critical threshold

### Main Process: Collecting System Statistics

The IPC handler in [`electron/main/ipc-handlers.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts) (lines 630-661) implements `system:memory`:

```typescript
// electron/main/ipc-handlers.ts
ipcMain.handle('system:memory', async () => {
  const total = os.totalmem();
  const free = os.freemem();
  const used = total - free;
  
  // macOS: mirror Activity Monitor by parsing vm_stat
  // to include wired, active, and compressed pages
  if (process.platform === 'darwin') {
    const vmStats = parseVmStat();   // Executes vm_stat, parses output
    return {
      total,
      used: vmStats.wired + vmStats.active + vmStats.compressed,
      available: total - (vmStats.wired + vmStats.active + vmStats.compressed)
    };
  }
  
  return { total, used, available: free };
});

```

**Cross-platform behavior:**

| Platform | Implementation Detail |
|----------|----------------------|
| Linux/Windows | Uses Node.js `os.totalmem()` / `os.freemem()` |
| macOS | Parses `vm_stat` output to match Activity Monitor's accounting of wired, active, and compressed memory |

**Key files:**
- Renderer: [[`src/shared/components/layout/MemoryIndicator.tsx`](https://github.com/lightningpixel/modly/blob/main/src/shared/components/layout/MemoryIndicator.tsx)](https://github.com/lightningpixel/modly/blob/main/src/shared/components/layout/MemoryIndicator.tsx)
- Main process: [`electron/main/ipc-handlers.ts#L630-L661`](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts#L630-L661)

---

## Complete Memory Management Workflow

Modly's three mechanisms operate as a coordinated system:

1. **Prevention:** Set VRAM limits before generation starts
2. **Intervention:** Call unload APIs when memory runs low
3. **Visibility:** Monitor `MemoryIndicator` to time interventions

**Practical usage example:**

```tsx
// 1. Configure 8 GB VRAM budget in Settings
<PerformanceSection />   // Select "8 GB" from dropdown

// 2. Monitor RAM during heavy generation sessions
<MemoryIndicator />      // Watch for amber/red thresholds

// 3. Free resources when needed
await fetch('/api/model/unload_all', { method: 'POST' });

```

---

## Summary

- **VRAM budgeting** via `PerformanceSection` prevents GPU over-allocation by passing user-defined limits to generation subprocesses.
- **Model unloading APIs** in [`model.py`](https://github.com/lightningpixel/modly/blob/main/model.py) release GPU buffers and trigger garbage collection; macOS Apple Silicon requires subprocess termination for full reclamation.
- **MemoryIndicator** polls every 2 seconds through [`ipc-handlers.ts`](https://github.com/lightningpixel/modly/blob/main/ipc-handlers.ts), with macOS-specific `vm_stat` parsing for accurate Activity Monitor parity.
- Color-coded thresholds (75% amber, 90% red) give immediate visual feedback for memory pressure.

---

## Frequently Asked Questions

### How do I check Modly's RAM usage in real time?

The **MemoryIndicator** in the top toolbar updates every 2 seconds. It displays a progress bar showing used versus total system RAM, with color shifts to amber at 75% and red at 90% usage. On macOS, it specifically mirrors Activity Monitor's accounting by including wired and compressed memory via `vm_stat` parsing in [[`electron/main/ipc-handlers.ts`](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts)](https://github.com/lightningpixel/modly/blob/main/electron/main/ipc-handlers.ts#L630-L661).

### Why does Modly terminate subprocesses on Apple Silicon instead of just unloading models?

Metal's memory allocator on **Apple Silicon** marks deallocated GPU buffers as "wired," meaning they remain resident and unavailable to other processes. According to [[`arch/decisions/APPLE-SILICON-SUPPORT.md`](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md)](https://github.com/lightningpixel/modly/blob/main/arch/decisions/APPLE-SILICON-SUPPORT.md), terminating the subprocess is the only reliable way to return this memory to the system pool.

### Can I automate GPU memory cleanup in Modly?

Yes. Call the **Model API** endpoints programmatically. `POST /api/model/unload_all` unloads every active model and runs garbage collection. For targeted cleanup, use `POST /api/model/unload` with a specific model ID. Both routes are defined in [[`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py)](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py).

### Where is the VRAM limit stored and how does it affect generation?

The `PerformanceSection` component stores the selected limit in **React state** as a string (e.g., `'8'` for 8 GB). This value passes to the backend when launching a generation subprocess. The backend uses it to constrain PyTorch's CUDA or Metal memory planning, as implemented in [[`src/areas/settings/components/PerformanceSection.tsx`](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx)](https://github.com/lightningpixel/modly/blob/main/src/areas/settings/components/PerformanceSection.tsx).