# Performance Implications of Running Multiple Modly Extension Processes Concurrently

> Discover the performance implications of running multiple Modly extension processes. Understand memory scaling and crash isolation benefits to optimize your applications.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: performance
- Published: 2026-08-20

---

**Running multiple Modly extension processes concurrently provides crash isolation but scales linearly in memory usage, with each subprocess consuming 30–50 MB of RAM minimum plus GPU memory if applicable.**

Modly's architecture packages each extension in an isolated Python virtual environment and communicates through dedicated subprocesses. According to the `lightningpixel/modly` source code, this design prioritizes reliability over resource efficiency—understanding these trade-offs is essential for production deployments managing many extensions simultaneously.

## How Modly Isolates Extensions

### Subprocess-Based Architecture

The core isolation mechanism lives in [`api/services/extension_process.py`](https://github.com/lightningpixel/modly/blob/main/api/services/extension_process.py). The `ExtensionProcess` class spawns each extension via `subprocess.Popen` in its `_start` method (lines 93–135):

```python

# From api/services/extension_process.py

# Each extension gets its own Python interpreter in a separate venv

self._process = subprocess.Popen(
    [venv_python, "-m", "modly.api.runner"],
    cwd=self._extension_dir,
    env=env,
    stdin=subprocess.PIPE,
    stdout=subprocess.PIPE,
    stderr=subprocess.PIPE,
    text=True,
)

```

This guarantees that a **crash or memory leak** in one extension cannot corrupt others—a critical reliability feature for plugin systems.

### Inter-Process Communication Overhead

Communication uses **newline-delimited JSON** on standard streams. The `_read_loop` and `_stderr_loop` methods parse output via dedicated threads per process:

```python

# Each ExtensionProcess spawns reader threads

self._stdout_thread = threading.Thread(target=self._read_loop, daemon=True)
self._stderr_thread = threading.Thread(target=self._stderr_loop, daemon=True)

```

This per-process thread overhead is minimal for small deployments but compounds with scale.

## Memory Scaling Characteristics

### System RAM Consumption

Each extension carries **independent Python interpreter overhead** plus its virtual environment:

| Factor | Typical Impact |
|--------|--------------|
| Base interpreter + venv | 30–50 MB per process |
| Imported libraries | Additional 10–200+ MB depending on extensions |
| Model weights (if CPU-only) | Stored in process-local memory |

Running 10 extensions simultaneously therefore consumes **300–500 MB minimum** before any model loading.

### GPU VRAM Contention

Extensions declare GPU requirements via the `VRAM_GB` field in their manifest. The `ExtensionProcess.__init__` populates this value, but **Modly does not implement GPU memory sharing**:

```python

# Example: VRAM tracking in ExtensionProcess

self.vram_gb = manifest.get("VRAM_GB", 0)

```

Each GPU-heavy extension loads its model independently. Three extensions requiring 4 GB each will attempt to allocate **12 GB total**—exceeding most consumer GPUs and triggering out-of-memory errors or forced paging to system RAM.

## CPU and I/O Bottlenecks

### Context Switching Costs

With many active extensions, the operating system must rapidly switch between processes. This **context-switching overhead** becomes measurable when:

- Extensions produce high-frequency status messages
- The host has fewer CPU cores than active extensions
- Queue contention occurs in the thread-safe `queue.Queue` used for message passing

### Disk Bandwidth Saturation

Extensions launched from their own `venv` directories (`_venv_python`) may simultaneously load large model files from the shared `MODELS_DIR`. Concurrent reads of multi-gigabyte files can **saturate disk bandwidth**, extending startup latency proportionally.

## Startup Latency Factors

### First-Run Package Installation

The `_install_missing_package` method triggers **pip installs** when imports fail:

```python

# Inside the child process—adds startup latency

def _install_missing_package(self, package_name):
    subprocess.run([self._venv_pip, "install", package_name], check=True)

```

This penalty applies only on first startup, not to running processes.

## Practical Concurrency Management

### Hardware-Based Limits

Modly imposes **no hard limit** on active extensions. The effective ceiling derives from available:

- **CPU cores** (for context switching efficiency)
- **System RAM** (30–50 MB base × extension count)
- **GPU VRAM** (sum of all `VRAM_GB` declarations)

### Recommended Cap Calculation

The source analysis provides a GPU-aware loading pattern:

```python
import torch
from modly.api.services.generator_registry import GeneratorRegistry

registry = GeneratorRegistry()
available_vram = torch.cuda.get_device_properties(0).total_memory / (1024**3)  # GB

threshold = available_vram * 0.8  # keep 20% headroom

used = 0.0

for ext_dir in ["/ext1", "/ext2", "/ext3"]:
    manifest = registry.inspect_manifest(ext_dir)
    if used + manifest["vram_gb"] <= threshold:
        registry.register_extension(ext_dir)
        registry.load(manifest["id"])
        used += manifest["vram_gb"]

```

## Key Implementation Files

| File | Role |
|------|------|
| [`api/services/extension_process.py`](https://github.com/lightningpixel/modly/blob/main/api/services/extension_process.py) | Spawns and manages isolated subprocesses per extension |
| [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) | Tracks extensions and routes generation calls |
| [`api/runner.py`](https://github.com/lightningpixel/modly/blob/main/api/runner.py) | Entry point executed inside each extension's subprocess |
| [`api/routers/extensions.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/extensions.py) | HTTP API for extension lifecycle management |

## Summary

- **Process isolation** via `subprocess.Popen` prevents cross-extension crashes but multiplies memory overhead
- **RAM consumption** scales linearly at 30–50 MB minimum per extension
- **GPU VRAM** is not shared; sum all `VRAM_GB` declarations against hardware limits
- **CPU overhead** from threads and context switching grows with extension count
- **Disk I/O** contention occurs during concurrent model loading
- **No built-in limits** exist—administrators must enforce caps based on hardware profiles

## Frequently Asked Questions

### How much RAM does each Modly extension process use?

Each extension consumes 30–50 MB minimum for its isolated Python interpreter and virtual environment, plus additional memory for imported libraries and any loaded models. A deployment with 20 extensions therefore requires roughly 600 MB to 1 GB of system RAM before model weights.

### Does Modly share GPU memory between extensions?

No. According to the `ExtensionProcess` implementation in [`api/services/extension_process.py`](https://github.com/lightningpixel/modly/blob/main/api/services/extension_process.py), each extension loads models independently into GPU memory. Extensions declare requirements via the `VRAM_GB` manifest field, but the scheduler does not implement memory pooling or sharing.

### What happens if I exceed available GPU memory?

Running GPU-heavy extensions whose combined `VRAM_GB` exceeds hardware capacity triggers **out-of-memory errors** or forces the GPU driver to page memory to system RAM. This dramatically slows inference performance and may cause generation failures.

### Is there a maximum number of extensions Modly can run simultaneously?

Modly implements no hard limit. The practical maximum depends on host **CPU cores, system RAM, GPU VRAM, and disk I/O bandwidth**. Administrators should monitor these resources and implement application-level throttling based on the `VRAM_GB` calculation pattern shown above.