How FastAPI Lifespan Handles Model Loading and Unloading in Modly
FastAPI's lifespan context manager in Modly orchestrates model adapter initialization during startup and ensures clean resource release during shutdown via the generator_registry singleton.
Modly is an open-source generative AI platform that manages multiple model backends through a FastAPI-based API. Understanding its lifespan management is critical for developers building resource-intensive ML services, as it demonstrates how to prevent memory leaks and handle graceful shutdowns when running models in subprocesses.
The Lifespan Pattern in FastAPI
FastAPI's lifespan protocol replaces older startup/shutdown event hooks with a single async context manager. This pattern gives explicit control over the entire application lifecycle.
In api/main.py, Modly defines its lifespan coroutine at lines 15-22:
async def lifespan(app: FastAPI):
# Startup: initialize the registry (instantiates all adapters)
from services.generator_registry import generator_registry
generator_registry.initialize()
yield
# Shutdown: unload all models
generator_registry.unload_all()
The yield statement creates a clear boundary: everything before runs once at startup, everything after runs once at shutdown.
Startup Phase: Loading All Models
When the FastAPI application starts, generator_registry.initialize() executes. This method, implemented in services/generator_registry.py, performs three critical operations:
- Scans the
extensions/folder for generator implementations - Instantiates each adapter as either a concrete
BaseGeneratorsubclass or anExtensionProcesswrapper - Configures model-specific directories (
MODELS_DIR,WORKSPACE_DIR) and records manifest metadata
The registry stores all created instances in a singleton GeneratorRegistry, making them available throughout the application's lifetime without repeated loading overhead.
Shutdown Phase: Clean Resource Release
After the yield in the lifespan coroutine, Modly ensures no resources leak when the service terminates. The generator_registry.unload_all() call (lines 21-22 in api/main.py) handles two distinct generator types differently, as shown in services/generator_registry.py lines 41-47:
- ExtensionProcess instances — Subprocess-based generators receive
gen.stop()to terminate their isolated Python environments - In-process generators — Direct
BaseGeneratorinstances receivegen.unload(), which explicitly frees model weights, clears caches, and closes file handles
This dual-path cleanup prevents zombie processes and GPU memory fragmentation.
On-Demand Unload Endpoints
Beyond automatic lifespan management, Modly exposes HTTP endpoints for manual memory control in api/routers/model.py (lines 83-98):
| Endpoint | Function |
|---|---|
POST /model/unload-all |
Triggers generator_registry.unload_all() and forces Python to return memory to the OS |
POST /model/unload/{model_id} |
Unloads a specific model by ID while keeping others resident |
These endpoints are essential for long-running desktop deployments where users may need to free VRAM between different model workflows.
Practical Implementation Examples
Attaching Lifespan to Your FastAPI App
from fastapi import FastAPI
from api.main import lifespan
app = FastAPI(title="Modly API", version="0.4.1", lifespan=lifespan)
Checking Which Models Are Currently Loaded
from fastapi import APIRouter
from services.generator_registry import generator_registry
router = APIRouter()
@router.get("/models/loaded")
def loaded_models():
# Returns a list of model IDs that are currently loaded
return [mid for mid, gen in generator_registry._generators.items()
if gen.is_loaded()]
Programmatically Freeing VRAM
from services.generator_registry import generator_registry
def free_vram():
generator_registry.unload_all()
import gc; gc.collect() # optional: force GC
Key Source Files
| File | Purpose |
|---|---|
api/main.py |
Defines the FastAPI app and lifespan coroutine |
api/services/generator_registry.py |
Central registry with initialize() / unload_all() logic |
api/routers/model.py |
HTTP endpoints for manual unload operations |
Summary
- Lifespan guarantees — Modly's FastAPI lifespan ensures models load once at startup and release cleanly at shutdown
- Dual cleanup paths — The registry distinguishes between subprocess and in-process generators for proper termination
- Manual control API — Endpoints at
/model/unload-alland/model/unload/{model_id}enable runtime memory management - Production pattern — This implementation serves as a reference for ML services needing reliable resource lifecycle management
Frequently Asked Questions
What is the FastAPI lifespan and why use it instead of @app.on_event?
The lifespan is an async context manager that replaces the deprecated @app.on_event("startup") and @app.on_event("shutdown") decorators. It provides cleaner resource management with explicit yield-based separation between startup and shutdown phases, and integrates better with modern async Python patterns. Modly uses it to ensure the generator_registry initializes before any request handlers execute.
How does Modly handle generators running in separate processes?
Modly wraps external generators in an ExtensionProcess class. During shutdown, unload_all() detects these instances and calls gen.stop() rather than gen.unload(), which sends termination signals to the subprocess. This prevents orphaned Python processes when the Electron frontend closes or the API server restarts.
Can I unload models without stopping the entire FastAPI server?
Yes. The POST /model/unload-all and POST /model/unload/{model_id} endpoints in routers/model.py trigger the same cleanup methods used during lifespan shutdown, but leave the server running. After unloading, you can subsequently call generator_registry.initialize() or load specific models on demand without a full restart.
Where is model metadata stored during the lifespan?
The GeneratorRegistry singleton in services/generator_registry.py maintains the _generators dictionary mapping model IDs to adapter instances. It also populates manifest metadata and ensures MODELS_DIR and WORKSPACE_DIR exist before any generator attempts to access them.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →