# FastAPI Backend Router Architecture in Modly: Generation, Model, Optimize, and Workflow_Runs Explained

> Explore the Modly FastAPI backend router architecture: generation, model, optimize, and workflow_runs. Learn how it handles state and delegates tasks efficiently.

- Repository: [lightningpixel/modly](https://github.com/lightningpixel/modly)
- Tags: architecture
- Published: 2026-08-15

---

**The Modly FastAPI backend uses a single FastAPI instance with four functional routers—generation, model, optimize, and workflow_runs—mounted at distinct prefixes, each delegating heavy work to service classes while managing state through thread-safe dictionaries and BackgroundTasks.**

This article examines the **FastAPI backend router architecture** in [lightningpixel/modly](https://github.com/lightningpixel/modly), an Electron-based desktop application for image-to-3D generation. The server architecture follows a service-oriented pattern where HTTP routers expose endpoints, delegate processing to service classes, and maintain job state through module-level dictionaries protected by threading primitives.

## Overview of the Router Architecture

Modly creates one **FastAPI application** in [`api/main.py`](https://github.com/lightningpixel/modly/blob/main/api/main.py) and mounts four specialized routers:

| Router | Prefix | Core Purpose |
|--------|--------|--------------|
| Generation | `/generate` | Image-to-3D jobs with progress tracking and cancellation |
| Model | `/model` | Model lifecycle, HuggingFace downloads, memory management |
| Optimize | `/optimize` | Mesh post-processing and format conversion |
| Workflow-Runs | `/workflow-runs` | UI-friendly wrapper around generation pipelines |

```python

# From api/main.py

app.include_router(generation.router, prefix="/generate")
app.include_router(model.router,      prefix="/model")
app.include_router(optimize.router,   prefix="/optimize")
app.include_router(workflow_runs.router, prefix="/workflow-runs")

```

All routers share a **workspace directory** (`WORKSPACE_DIR` from `generator_registry`) for file storage and rely on the **`generator_registry`** service for model management.

## Generation Router: Asynchronous Job Processing

The **generation router** in [`api/routers/generation.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/generation.py) handles the core image-to-3D pipeline. Its architecture centers on **background task execution** with robust cancellation support.

### State Management Design

The router maintains four module-level data structures for **thread-safe job tracking**:

- `_jobs: Dict[str, JobStatus]` — stores job metadata
- `_cancel_events: Dict[str, threading.Event]` — per-job cancellation signals
- `_cancelled: Set[str]` — completed cancellations
- `_completed_at: Dict[str, datetime]` — timestamps for cleanup

### Key Endpoint Flow

The `POST /from-image` endpoint demonstrates the typical flow:

1. Validate input parameters and image
2. Call `generator_registry.switch_model()` to ensure the requested model is active
3. Create a `JobStatus` instance and store it in `_jobs`
4. Launch `_run_generation()` via `background_tasks.add_task()`
5. Return the `job_id` immediately

```python

# POST /generate/from-image

@router.post("/from-image")
async def generate_from_image(
    image: UploadFile,
    model_id: str,
    collection: str = "output",
    background_tasks: BackgroundTasks = BackgroundTasks(),
    ...
):
    # Switch model if needed

    await generator_registry.switch_model(model_id)
    
    # Create job record

    job_id = str(uuid.uuid4())
    _jobs[job_id] = JobStatus(
        job_id=job_id,
        status="pending",
        progress=0,
        step="Initializing"
    )
    
    # Launch background work

    background_tasks.add_task(
        _run_generation,
        job_id=job_id,
        image_bytes=await image.read(),
        params=params,
        collection=collection
    )
    
    return {"job_id": job_id}

```

### Background Execution and Progress

The `_run_generation()` async function runs in a **thread pool** and reports progress through callbacks:

```python
async def _run_generation(job_id: str, image_bytes: bytes, params: dict, collection: str):
    _jobs[job_id].status = "running"
    
    # Get active generator (may block during model load)

    generator = await generator_registry.get_active()
    
    # Progress callback updates shared state

    def progress_cb(percent: int, step: str):
        _jobs[job_id].progress = percent
        _jobs[job_id].step = step
    
    # Execute generation in thread pool

    loop = asyncio.get_event_loop()
    result = await loop.run_in_executor(
        None,
        lambda: generator.generate(
            image=image_bytes,
            params=params,
            progress_cb=progress_cb,
            cancel_event=_cancel_events.get(job_id)
        )
    )
    
    # Store result and mark complete

    output_path = WORKSPACE_DIR / collection / f"{job_id}.glb"
    _jobs[job_id].output_url = f"/workspace/{collection}/{job_id}.glb"
    _jobs[job_id].status = "completed"

```

### Cancellation Mechanism

The **cancellation endpoint** (`POST /cancel/{job_id}`) coordinates between threading primitives and subprocess termination:

```python
@router.post("/cancel/{job_id}")
async def cancel_generation(job_id: str):
    if job_id not in _jobs:
        raise HTTPException(404, "Job not found")
    
    # Signal cancellation

    _cancelled.add(job_id)
    if job_id in _cancel_events:
        _cancel_events[job_id].set()
    
    # Kill subprocess if running

    generator = generator_registry.get_active_sync()
    if hasattr(generator, '_proc') and generator._proc:
        generator._proc.kill()
    
    _jobs[job_id].status = "cancelled"
    return {"status": "cancelled"}

```

## Model Router: Lifecycle and Download Management

The **model router** in [`api/routers/model.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/model.py) exposes comprehensive model management, with particular sophistication around **HuggingFace Hub integration**.

### Core Endpoints

| Endpoint | Method | Purpose |
|----------|--------|---------|
| `/status` | GET | Active model status (loaded, VRAM usage) |
| `/all` | GET | All known models with availability |
| `/params` | GET | Parameter schema for UI generation |
| `/switch` | POST | Change active model (triggers unload) |
| `/unload-all` | POST | Free all GPU memory |
| `/hf-download/*` | POST/GET | Streaming download with SSE progress |

### HuggingFace Download with Server-Sent Events

The **download endpoint** streams progress to the client using SSE:

```python

# POST /model/hf-download

@router.post("/hf-download")
async def hf_download(repo_id: str, model_id: str):
    _download_controls[model_id] = {
        "pause": threading.Event(),
        "cancel": threading.Event()
    }
    
    async def generate():
        async for chunk in _download_file_streamed(repo_id, model_id):
            yield f"data: {json.dumps(chunk)}\n\n"
        yield "data: [DONE]\n\n"
    
    return StreamingResponse(
        generate(),
        media_type="text/event-stream",
        headers={"Cache-Control": "no-cache"}
    )

```

The download implementation includes **resumable range requests**, exponential backoff retries, and pause/cancel support through the `_download_controls` dictionary.

### Model Registry Integration

All model endpoints interact directly with `generator_registry`:

```python

# From GET /model/status

@router.get("/status")
async def model_status():
    active = generator_registry.get_active_model_id()
    info = generator_registry.get_model_info(active)
    return {
        "model_id": active,
        "downloaded": info.downloaded,
        "loaded": info.loaded,
        "vram_mb": info.vram_usage
    }

```

## Optimize Router: Mesh Post-Processing

The **optimize router** in [`api/routers/optimize.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/optimize.py) provides **synchronous CPU-bound operations** for mesh refinement. Unlike the generation router, it relies on FastAPI's automatic thread-pool execution for blocking operations.

### Processing Capabilities

| Operation | Dependencies | Input/Output |
|-----------|--------------|------------|
| Decimation | pymeshlab | OBJ/PLY → reduced face count |
| Smoothing | pymeshlab | Laplacian smoothing |
| Transform | trimesh | 4×4 matrix application |
| Format conversion | trimesh/pymeshlab | PLY → SPLAT, GLB, OBJ |

### Decimation with Texture Preservation

The `_decimate()` function handles two code paths based on texture presence:

```python
def _decimate(input_path: Path, target_faces: int, output_path: Path):
    has_tex = _has_texture(input_path)
    
    # Load into MeshSet

    ms = pymeshlab.MeshSet()
    ms.load_new_mesh(str(input_path))
    
    if has_tex:
        # Quadric edge collapse decimation with texture coordinate preservation

        ms.apply_filter(
            "meshing_decimation_quadric_edge_collapse",
            targetfacenum=target_faces,
            preservetopology=True,
            texture=True
        )
        # Save as OBJ to retain texture reference

        ms.save_current_mesh(str(output_path), save_textures=True)
    else:
        # PLY path: no texture handling needed

        ms.apply_filter(
            "meshing_decimation_quadric_edge_collapse",
            targetfacenum=target_faces
        )
        ms.save_current_mesh(str(output_path))

```

### Workspace Integration

All optimized outputs are written to paths under `WORKSPACE_DIR`, with the `/serve-file` endpoint returning proper MIME types:

```python
@router.get("/serve-file")
async def serve_file(path: str):
    full_path = WORKSPACE_DIR / path
    mime = "model/gltf-binary" if path.endswith(".glb") else "application/octet-stream"
    return FileResponse(full_path, media_type=mime)

```

## Workflow_Runs Router: UI-Friendly Abstraction

The **workflow_runs router** in [`api/routers/workflow_runs.py`](https://github.com/lightningpixel/modly/blob/main/api/routers/workflow_runs.py) provides a **thin façade** over generation functionality, designed for frontend workflow orchestration.

### Architectural Relationship

The router **reuses generation infrastructure directly**:

```python

# From workflow_runs.py

from routers.generation import (
    _jobs,
    _cancel_events,
    _run_generation,
    WORKSPACE_DIR
)

```

### WorkflowRunStatus Model

The key distinction is the **response schema**, which includes UI-specific fields:

```python
class WorkflowRunStatus(BaseModel):
    run_id: str
    status: Literal["pending", "running", "completed", "cancelled", "error"]
    progress: int
    step: Optional[str]
    scene_candidate: Optional[SceneCandidate]  # UI navigation data

    error: Optional[str]

```

### Endpoint Mapping

| Workflow Endpoint | Delegates To | Added Value |
|-------------------|------------|-------------|
| `POST /from-image` | generation logic | Returns `run_id`, initializes `scene_candidate` |
| `GET /{run_id}` | `_jobs` lookup | Transforms `JobStatus` → `WorkflowRunStatus` |
| `POST /{run_id}/cancel` | cancellation logic | Same subprocess-kill behavior |

The `GET /{run_id}` endpoint demonstrates the transformation:

```python
@router.get("/{run_id}")
async def get_workflow_run(run_id: str):
    if run_id not in _jobs:
        raise HTTPException(404, "Run not found")
    
    job = _jobs[run_id]
    scene = None
    if job.output_url:
        # Parse workspace path from output URL

        relative_path = job.output_url.replace("/workspace/", "")
        scene = SceneCandidate(
            path=relative_path,
            format="glb" if relative_path.endswith(".glb") else "splat"
        )
    
    return WorkflowRunStatus(
        run_id=run_id,
        status=job.status,
        progress=job.progress,
        step=job.step,
        scene_candidate=scene,
        error=job.error
    )

```

## Central Application Configuration

The **FastAPI application factory** in [`api/main.py`](https://github.com/lightningpixel/modly/blob/main/api/main.py) orchestrates router registration and shared concerns:

### Lifespan Management

```python
@asynccontextmanager
async def lifespan(app: FastAPI):
    # Startup: initialize all generator adapters

    await generator_registry.initialize()
    yield
    # Shutdown: clean release of GPU memory

    await generator_registry.unload_all()

```

### Static File Serving

A catch-all endpoint serves workspace assets:

```python
@app.get("/workspace/{full_path:path}")
async def serve_workspace(full_path: str):
    target = WORKSPACE_DIR / full_path
    # Security: ensure path stays within WORKSPACE_DIR

    if not target.resolve().is_relative_to(WORKSPACE_DIR.resolve()):
        raise HTTPException(403, "Access denied")
    return FileResponse(target)

```

## API Usage Examples

### Submit and Monitor Generation

```bash

# Start job

curl -X POST "http://localhost:8000/generate/from-image" \
  -F "image=@product.jpg" \
  -F "model_id=sf3d" \
  -F "collection=BatchA"

# Poll status

curl "http://localhost:8000/generate/status/{job_id}"

```

### Stream Model Download

```bash
curl "http://localhost:8000/model/hf-download?repo_id=stabilityai/stable-fast-3d&model_id=sf3d"

```

### Optimize Result

```bash
curl -X POST "http://localhost:8000/optimize/mesh" \
  -H "Content-Type: application/json" \
  -d '{"path":"BatchA/job_xxx.glb","target_faces":3000}'

```

## Summary

The **Modly FastAPI backend router architecture** demonstrates several production patterns for desktop-embedded AI services:

- **Router-service separation** — HTTP concerns isolated from heavy processing
- **Thread-safe state** — `threading.Event` and module-level dictionaries coordinate across async/await and thread-pool boundaries
- **Streaming downloads** — SSE for real-time progress on long-running HuggingFace operations
- **Façade pattern** — workflow_runs router reuses generation internals with UI-optimized contracts
- **Workspace sandbox** — unified file storage with path-traversal-protected static serving

All four routers mount on a single FastAPI instance with predictable prefixes, sharing the `generator_registry` service for model lifecycle management and `WORKSPACE_DIR` for asset persistence.

## Frequently Asked Questions

### How does the generation router handle cancellation?

The generation router uses a **two-stage cancellation mechanism**. First, setting a `threading.Event` in `_cancel_events[job_id]` signals the generator to stop at the next checkpoint. Second, if the generator spawned a subprocess, the `/cancel/{job_id}` endpoint calls `gen._proc.kill()` for immediate termination. The job status transitions to `"cancelled"` regardless of which mechanism succeeds.

### What is the difference between the generation and workflow_runs routers?

The **workflow_runs router** is a thin abstraction layer over **generation router internals**. Both use the same `_jobs` dictionary, `_run_generation()` function, and cancellation logic. The workflow router transforms `JobStatus` into `WorkflowRunStatus` to include UI-specific fields like `scene_candidate`, while the generation router exposes raw job metadata optimized for polling clients.

### How does the model router stream HuggingFace downloads?

The **model router** implements **Server-Sent Events (SSE)** through FastAPI's `StreamingResponse`. The `_download_file_streamed()` async generator yields JSON fragments with download progress, which the endpoint wraps in `data:` protocol format. The client receives real-time updates without WebSocket overhead, and the download supports pause/resume through module-level `_download_controls` events.

### Why are optimize router endpoints synchronous?

The **optimize router** uses **synchronous functions** because mesh processing through `pymeshlab` and `trimesh` releases the GIL but performs CPU-intensive work unsuitable for async/await. FastAPI automatically runs these endpoints in a thread pool via `run_in_threadpool`, preventing event loop blocking while leveraging multi-core processing.