How FastAPI Lifespan Handles Model Loading and Unloading in Modly

FastAPI's lifespan context manager in Modly orchestrates model adapter initialization during startup and ensures clean resource release during shutdown via the generator_registry singleton.

Modly is an open-source generative AI platform that manages multiple model backends through a FastAPI-based API. Understanding its lifespan management is critical for developers building resource-intensive ML services, as it demonstrates how to prevent memory leaks and handle graceful shutdowns when running models in subprocesses.


The Lifespan Pattern in FastAPI

FastAPI's lifespan protocol replaces older startup/shutdown event hooks with a single async context manager. This pattern gives explicit control over the entire application lifecycle.

In api/main.py, Modly defines its lifespan coroutine at lines 15-22:

async def lifespan(app: FastAPI):
    # Startup: initialize the registry (instantiates all adapters)

    from services.generator_registry import generator_registry
    generator_registry.initialize()
    yield
    # Shutdown: unload all models

    generator_registry.unload_all()

The yield statement creates a clear boundary: everything before runs once at startup, everything after runs once at shutdown.


Startup Phase: Loading All Models

When the FastAPI application starts, generator_registry.initialize() executes. This method, implemented in services/generator_registry.py, performs three critical operations:

  1. Scans the extensions/ folder for generator implementations
  2. Instantiates each adapter as either a concrete BaseGenerator subclass or an ExtensionProcess wrapper
  3. Configures model-specific directories (MODELS_DIR, WORKSPACE_DIR) and records manifest metadata

The registry stores all created instances in a singleton GeneratorRegistry, making them available throughout the application's lifetime without repeated loading overhead.


Shutdown Phase: Clean Resource Release

After the yield in the lifespan coroutine, Modly ensures no resources leak when the service terminates. The generator_registry.unload_all() call (lines 21-22 in api/main.py) handles two distinct generator types differently, as shown in services/generator_registry.py lines 41-47:

  • ExtensionProcess instances — Subprocess-based generators receive gen.stop() to terminate their isolated Python environments
  • In-process generators — Direct BaseGenerator instances receive gen.unload(), which explicitly frees model weights, clears caches, and closes file handles

This dual-path cleanup prevents zombie processes and GPU memory fragmentation.


On-Demand Unload Endpoints

Beyond automatic lifespan management, Modly exposes HTTP endpoints for manual memory control in api/routers/model.py (lines 83-98):

Endpoint Function
POST /model/unload-all Triggers generator_registry.unload_all() and forces Python to return memory to the OS
POST /model/unload/{model_id} Unloads a specific model by ID while keeping others resident

These endpoints are essential for long-running desktop deployments where users may need to free VRAM between different model workflows.


Practical Implementation Examples

Attaching Lifespan to Your FastAPI App

from fastapi import FastAPI
from api.main import lifespan

app = FastAPI(title="Modly API", version="0.4.1", lifespan=lifespan)

Checking Which Models Are Currently Loaded

from fastapi import APIRouter
from services.generator_registry import generator_registry

router = APIRouter()

@router.get("/models/loaded")
def loaded_models():
    # Returns a list of model IDs that are currently loaded

    return [mid for mid, gen in generator_registry._generators.items()
            if gen.is_loaded()]

Programmatically Freeing VRAM

from services.generator_registry import generator_registry

def free_vram():
    generator_registry.unload_all()
    import gc; gc.collect()  # optional: force GC

Key Source Files

File Purpose
api/main.py Defines the FastAPI app and lifespan coroutine
api/services/generator_registry.py Central registry with initialize() / unload_all() logic
api/routers/model.py HTTP endpoints for manual unload operations

Summary

  • Lifespan guarantees — Modly's FastAPI lifespan ensures models load once at startup and release cleanly at shutdown
  • Dual cleanup paths — The registry distinguishes between subprocess and in-process generators for proper termination
  • Manual control API — Endpoints at /model/unload-all and /model/unload/{model_id} enable runtime memory management
  • Production pattern — This implementation serves as a reference for ML services needing reliable resource lifecycle management

Frequently Asked Questions

What is the FastAPI lifespan and why use it instead of @app.on_event?

The lifespan is an async context manager that replaces the deprecated @app.on_event("startup") and @app.on_event("shutdown") decorators. It provides cleaner resource management with explicit yield-based separation between startup and shutdown phases, and integrates better with modern async Python patterns. Modly uses it to ensure the generator_registry initializes before any request handlers execute.

How does Modly handle generators running in separate processes?

Modly wraps external generators in an ExtensionProcess class. During shutdown, unload_all() detects these instances and calls gen.stop() rather than gen.unload(), which sends termination signals to the subprocess. This prevents orphaned Python processes when the Electron frontend closes or the API server restarts.

Can I unload models without stopping the entire FastAPI server?

Yes. The POST /model/unload-all and POST /model/unload/{model_id} endpoints in routers/model.py trigger the same cleanup methods used during lifespan shutdown, but leave the server running. After unloading, you can subsequently call generator_registry.initialize() or load specific models on demand without a full restart.

Where is model metadata stored during the lifespan?

The GeneratorRegistry singleton in services/generator_registry.py maintains the _generators dictionary mapping model IDs to adapter instances. It also populates manifest metadata and ensures MODELS_DIR and WORKSPACE_DIR exist before any generator attempts to access them.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →