Performance Implications of Running Multiple Modly Extension Processes Concurrently
Running multiple Modly extension processes concurrently provides crash isolation but scales linearly in memory usage, with each subprocess consuming 30–50 MB of RAM minimum plus GPU memory if applicable.
Modly's architecture packages each extension in an isolated Python virtual environment and communicates through dedicated subprocesses. According to the lightningpixel/modly source code, this design prioritizes reliability over resource efficiency—understanding these trade-offs is essential for production deployments managing many extensions simultaneously.
How Modly Isolates Extensions
Subprocess-Based Architecture
The core isolation mechanism lives in api/services/extension_process.py. The ExtensionProcess class spawns each extension via subprocess.Popen in its _start method (lines 93–135):
# From api/services/extension_process.py
# Each extension gets its own Python interpreter in a separate venv
self._process = subprocess.Popen(
[venv_python, "-m", "modly.api.runner"],
cwd=self._extension_dir,
env=env,
stdin=subprocess.PIPE,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
text=True,
)
This guarantees that a crash or memory leak in one extension cannot corrupt others—a critical reliability feature for plugin systems.
Inter-Process Communication Overhead
Communication uses newline-delimited JSON on standard streams. The _read_loop and _stderr_loop methods parse output via dedicated threads per process:
# Each ExtensionProcess spawns reader threads
self._stdout_thread = threading.Thread(target=self._read_loop, daemon=True)
self._stderr_thread = threading.Thread(target=self._stderr_loop, daemon=True)
This per-process thread overhead is minimal for small deployments but compounds with scale.
Memory Scaling Characteristics
System RAM Consumption
Each extension carries independent Python interpreter overhead plus its virtual environment:
| Factor | Typical Impact |
|---|---|
| Base interpreter + venv | 30–50 MB per process |
| Imported libraries | Additional 10–200+ MB depending on extensions |
| Model weights (if CPU-only) | Stored in process-local memory |
Running 10 extensions simultaneously therefore consumes 300–500 MB minimum before any model loading.
GPU VRAM Contention
Extensions declare GPU requirements via the VRAM_GB field in their manifest. The ExtensionProcess.__init__ populates this value, but Modly does not implement GPU memory sharing:
# Example: VRAM tracking in ExtensionProcess
self.vram_gb = manifest.get("VRAM_GB", 0)
Each GPU-heavy extension loads its model independently. Three extensions requiring 4 GB each will attempt to allocate 12 GB total—exceeding most consumer GPUs and triggering out-of-memory errors or forced paging to system RAM.
CPU and I/O Bottlenecks
Context Switching Costs
With many active extensions, the operating system must rapidly switch between processes. This context-switching overhead becomes measurable when:
- Extensions produce high-frequency status messages
- The host has fewer CPU cores than active extensions
- Queue contention occurs in the thread-safe
queue.Queueused for message passing
Disk Bandwidth Saturation
Extensions launched from their own venv directories (_venv_python) may simultaneously load large model files from the shared MODELS_DIR. Concurrent reads of multi-gigabyte files can saturate disk bandwidth, extending startup latency proportionally.
Startup Latency Factors
First-Run Package Installation
The _install_missing_package method triggers pip installs when imports fail:
# Inside the child process—adds startup latency
def _install_missing_package(self, package_name):
subprocess.run([self._venv_pip, "install", package_name], check=True)
This penalty applies only on first startup, not to running processes.
Practical Concurrency Management
Hardware-Based Limits
Modly imposes no hard limit on active extensions. The effective ceiling derives from available:
- CPU cores (for context switching efficiency)
- System RAM (30–50 MB base × extension count)
- GPU VRAM (sum of all
VRAM_GBdeclarations)
Recommended Cap Calculation
The source analysis provides a GPU-aware loading pattern:
import torch
from modly.api.services.generator_registry import GeneratorRegistry
registry = GeneratorRegistry()
available_vram = torch.cuda.get_device_properties(0).total_memory / (1024**3) # GB
threshold = available_vram * 0.8 # keep 20% headroom
used = 0.0
for ext_dir in ["/ext1", "/ext2", "/ext3"]:
manifest = registry.inspect_manifest(ext_dir)
if used + manifest["vram_gb"] <= threshold:
registry.register_extension(ext_dir)
registry.load(manifest["id"])
used += manifest["vram_gb"]
Key Implementation Files
| File | Role |
|---|---|
api/services/extension_process.py |
Spawns and manages isolated subprocesses per extension |
api/services/generator_registry.py |
Tracks extensions and routes generation calls |
api/runner.py |
Entry point executed inside each extension's subprocess |
api/routers/extensions.py |
HTTP API for extension lifecycle management |
Summary
- Process isolation via
subprocess.Popenprevents cross-extension crashes but multiplies memory overhead - RAM consumption scales linearly at 30–50 MB minimum per extension
- GPU VRAM is not shared; sum all
VRAM_GBdeclarations against hardware limits - CPU overhead from threads and context switching grows with extension count
- Disk I/O contention occurs during concurrent model loading
- No built-in limits exist—administrators must enforce caps based on hardware profiles
Frequently Asked Questions
How much RAM does each Modly extension process use?
Each extension consumes 30–50 MB minimum for its isolated Python interpreter and virtual environment, plus additional memory for imported libraries and any loaded models. A deployment with 20 extensions therefore requires roughly 600 MB to 1 GB of system RAM before model weights.
Does Modly share GPU memory between extensions?
No. According to the ExtensionProcess implementation in api/services/extension_process.py, each extension loads models independently into GPU memory. Extensions declare requirements via the VRAM_GB manifest field, but the scheduler does not implement memory pooling or sharing.
What happens if I exceed available GPU memory?
Running GPU-heavy extensions whose combined VRAM_GB exceeds hardware capacity triggers out-of-memory errors or forces the GPU driver to page memory to system RAM. This dramatically slows inference performance and may cause generation failures.
Is there a maximum number of extensions Modly can run simultaneously?
Modly implements no hard limit. The practical maximum depends on host CPU cores, system RAM, GPU VRAM, and disk I/O bandwidth. Administrators should monitor these resources and implement application-level throttling based on the VRAM_GB calculation pattern shown above.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →