# How FastAPI Lifespan Handles Model Loading and Unloading in Modly

> Learn how FastAPI lifespan manages model loading and unloading in Modly, orchestrating adapter initialization and resource release for efficient operations.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: internals
- Published: 2026-08-15

---

**FastAPI's lifespan context manager in Modly orchestrates model adapter initialization during startup and ensures clean resource release during shutdown via the `generator_registry` singleton.**

Modly is an open-source generative AI platform that manages multiple model backends through a FastAPI-based API. Understanding its **lifespan management** is critical for developers building resource-intensive ML services, as it demonstrates how to prevent memory leaks and handle graceful shutdowns when running models in subprocesses.

---

## The Lifespan Pattern in FastAPI

FastAPI's **lifespan** protocol replaces older startup/shutdown event hooks with a single async context manager. This pattern gives explicit control over the entire application lifecycle.

In [`api/main.py`](https://github.com/lightningpixel/modly/blob/main/api/main.py), Modly defines its lifespan coroutine at lines 15-22:

```python
async def lifespan(app: FastAPI):
    # Startup: initialize the registry (instantiates all adapters)

    from services.generator_registry import generator_registry
    generator_registry.initialize()
    yield
    # Shutdown: unload all models

    generator_registry.unload_all()

```

The `yield` statement creates a clear boundary: everything before runs once at startup, everything after runs once at shutdown.

---

## Startup Phase: Loading All Models

When the FastAPI application starts, `generator_registry.initialize()` executes. This method, implemented in [`services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/services/generator_registry.py), performs three critical operations:

1. **Scans** the `extensions/` folder for generator implementations
2. **Instantiates** each adapter as either a concrete `BaseGenerator` subclass or an `ExtensionProcess` wrapper
3. **Configures** model-specific directories (`MODELS_DIR`, `WORKSPACE_DIR`) and records manifest metadata

The registry stores all created instances in a singleton `GeneratorRegistry`, making them available throughout the application's lifetime without repeated loading overhead.

---

## Shutdown Phase: Clean Resource Release

After the `yield` in the lifespan coroutine, Modly ensures no resources leak when the service terminates. The `generator_registry.unload_all()` call (lines 21-22 in [`api/main.py`](https://github.com/lightningpixel/modly/blob/main/api/main.py)) handles two distinct generator types differently, as shown in [`services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/services/generator_registry.py) lines 41-47:

- **ExtensionProcess instances** — Subprocess-based generators receive `gen.stop()` to terminate their isolated Python environments
- **In-process generators** — Direct `BaseGenerator` instances receive `gen.unload()`, which explicitly frees model weights, clears caches, and closes file handles

This dual-path cleanup prevents zombie processes and GPU memory fragmentation.

---

## On-Demand Unload Endpoints

Beyond automatic lifespan management, Modly exposes HTTP endpoints for manual memory control in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) (lines 83-98):

| Endpoint | Function |
|----------|----------|
| `POST /model/unload-all` | Triggers `generator_registry.unload_all()` and forces Python to return memory to the OS |
| `POST /model/unload/{model_id}` | Unloads a specific model by ID while keeping others resident |

These endpoints are essential for long-running desktop deployments where users may need to free VRAM between different model workflows.

---

## Practical Implementation Examples

### Attaching Lifespan to Your FastAPI App

```python
from fastapi import FastAPI
from api.main import lifespan

app = FastAPI(title="Modly API", version="0.4.1", lifespan=lifespan)

```

### Checking Which Models Are Currently Loaded

```python
from fastapi import APIRouter
from services.generator_registry import generator_registry

router = APIRouter()

@router.get("/models/loaded")
def loaded_models():
    # Returns a list of model IDs that are currently loaded

    return [mid for mid, gen in generator_registry._generators.items()
            if gen.is_loaded()]

```

### Programmatically Freeing VRAM

```python
from services.generator_registry import generator_registry

def free_vram():
    generator_registry.unload_all()
    import gc; gc.collect()  # optional: force GC

```

---

## Key Source Files

| File | Purpose |
|------|---------|
| [`api/main.py`](https://github.com/lightningpixel/modly/blob/main/api/main.py) | Defines the FastAPI app and `lifespan` coroutine |
| [`api/services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/api/services/generator_registry.py) | Central registry with `initialize()` / `unload_all()` logic |
| [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) | HTTP endpoints for manual unload operations |

---

## Summary

- **Lifespan guarantees** — Modly's FastAPI lifespan ensures models load once at startup and release cleanly at shutdown
- **Dual cleanup paths** — The registry distinguishes between subprocess and in-process generators for proper termination
- **Manual control API** — Endpoints at `/model/unload-all` and `/model/unload/{model_id}` enable runtime memory management
- **Production pattern** — This implementation serves as a reference for ML services needing reliable resource lifecycle management

---

## Frequently Asked Questions

### What is the FastAPI lifespan and why use it instead of `@app.on_event`?

The **lifespan** is an async context manager that replaces the deprecated `@app.on_event("startup")` and `@app.on_event("shutdown")` decorators. It provides cleaner resource management with explicit `yield`-based separation between startup and shutdown phases, and integrates better with modern async Python patterns. Modly uses it to ensure the `generator_registry` initializes before any request handlers execute.

### How does Modly handle generators running in separate processes?

Modly wraps external generators in an `ExtensionProcess` class. During shutdown, `unload_all()` detects these instances and calls `gen.stop()` rather than `gen.unload()`, which sends termination signals to the subprocess. This prevents orphaned Python processes when the Electron frontend closes or the API server restarts.

### Can I unload models without stopping the entire FastAPI server?

Yes. The `POST /model/unload-all` and `POST /model/unload/{model_id}` endpoints in [`routers/model.py`](https://github.com/lightningpixel/modly/blob/main/routers/model.py) trigger the same cleanup methods used during lifespan shutdown, but leave the server running. After unloading, you can subsequently call `generator_registry.initialize()` or load specific models on demand without a full restart.

### Where is model metadata stored during the lifespan?

The `GeneratorRegistry` singleton in [`services/generator_registry.py`](https://github.com/lightningpixel/modly/blob/main/services/generator_registry.py) maintains the `_generators` dictionary mapping model IDs to adapter instances. It also populates manifest metadata and ensures `MODELS_DIR` and `WORKSPACE_DIR` exist before any generator attempts to access them.