FastAPI Backend Router Architecture in Modly: Generation, Model, Optimize, and Workflow_Runs Explained
The Modly FastAPI backend uses a single FastAPI instance with four functional routers—generation, model, optimize, and workflow_runs—mounted at distinct prefixes, each delegating heavy work to service classes while managing state through thread-safe dictionaries and BackgroundTasks.
This article examines the FastAPI backend router architecture in lightningpixel/modly, an Electron-based desktop application for image-to-3D generation. The server architecture follows a service-oriented pattern where HTTP routers expose endpoints, delegate processing to service classes, and maintain job state through module-level dictionaries protected by threading primitives.
Overview of the Router Architecture
Modly creates one FastAPI application in api/main.py and mounts four specialized routers:
| Router | Prefix | Core Purpose |
|---|---|---|
| Generation | /generate |
Image-to-3D jobs with progress tracking and cancellation |
| Model | /model |
Model lifecycle, HuggingFace downloads, memory management |
| Optimize | /optimize |
Mesh post-processing and format conversion |
| Workflow-Runs | /workflow-runs |
UI-friendly wrapper around generation pipelines |
# From api/main.py
app.include_router(generation.router, prefix="/generate")
app.include_router(model.router, prefix="/model")
app.include_router(optimize.router, prefix="/optimize")
app.include_router(workflow_runs.router, prefix="/workflow-runs")
All routers share a workspace directory (WORKSPACE_DIR from generator_registry) for file storage and rely on the generator_registry service for model management.
Generation Router: Asynchronous Job Processing
The generation router in api/routers/generation.py handles the core image-to-3D pipeline. Its architecture centers on background task execution with robust cancellation support.
State Management Design
The router maintains four module-level data structures for thread-safe job tracking:
_jobs: Dict[str, JobStatus]— stores job metadata_cancel_events: Dict[str, threading.Event]— per-job cancellation signals_cancelled: Set[str]— completed cancellations_completed_at: Dict[str, datetime]— timestamps for cleanup
Key Endpoint Flow
The POST /from-image endpoint demonstrates the typical flow:
- Validate input parameters and image
- Call
generator_registry.switch_model()to ensure the requested model is active - Create a
JobStatusinstance and store it in_jobs - Launch
_run_generation()viabackground_tasks.add_task() - Return the
job_idimmediately
# POST /generate/from-image
@router.post("/from-image")
async def generate_from_image(
image: UploadFile,
model_id: str,
collection: str = "output",
background_tasks: BackgroundTasks = BackgroundTasks(),
...
):
# Switch model if needed
await generator_registry.switch_model(model_id)
# Create job record
job_id = str(uuid.uuid4())
_jobs[job_id] = JobStatus(
job_id=job_id,
status="pending",
progress=0,
step="Initializing"
)
# Launch background work
background_tasks.add_task(
_run_generation,
job_id=job_id,
image_bytes=await image.read(),
params=params,
collection=collection
)
return {"job_id": job_id}
Background Execution and Progress
The _run_generation() async function runs in a thread pool and reports progress through callbacks:
async def _run_generation(job_id: str, image_bytes: bytes, params: dict, collection: str):
_jobs[job_id].status = "running"
# Get active generator (may block during model load)
generator = await generator_registry.get_active()
# Progress callback updates shared state
def progress_cb(percent: int, step: str):
_jobs[job_id].progress = percent
_jobs[job_id].step = step
# Execute generation in thread pool
loop = asyncio.get_event_loop()
result = await loop.run_in_executor(
None,
lambda: generator.generate(
image=image_bytes,
params=params,
progress_cb=progress_cb,
cancel_event=_cancel_events.get(job_id)
)
)
# Store result and mark complete
output_path = WORKSPACE_DIR / collection / f"{job_id}.glb"
_jobs[job_id].output_url = f"/workspace/{collection}/{job_id}.glb"
_jobs[job_id].status = "completed"
Cancellation Mechanism
The cancellation endpoint (POST /cancel/{job_id}) coordinates between threading primitives and subprocess termination:
@router.post("/cancel/{job_id}")
async def cancel_generation(job_id: str):
if job_id not in _jobs:
raise HTTPException(404, "Job not found")
# Signal cancellation
_cancelled.add(job_id)
if job_id in _cancel_events:
_cancel_events[job_id].set()
# Kill subprocess if running
generator = generator_registry.get_active_sync()
if hasattr(generator, '_proc') and generator._proc:
generator._proc.kill()
_jobs[job_id].status = "cancelled"
return {"status": "cancelled"}
Model Router: Lifecycle and Download Management
The model router in api/routers/model.py exposes comprehensive model management, with particular sophistication around HuggingFace Hub integration.
Core Endpoints
| Endpoint | Method | Purpose |
|---|---|---|
/status |
GET | Active model status (loaded, VRAM usage) |
/all |
GET | All known models with availability |
/params |
GET | Parameter schema for UI generation |
/switch |
POST | Change active model (triggers unload) |
/unload-all |
POST | Free all GPU memory |
/hf-download/* |
POST/GET | Streaming download with SSE progress |
HuggingFace Download with Server-Sent Events
The download endpoint streams progress to the client using SSE:
# POST /model/hf-download
@router.post("/hf-download")
async def hf_download(repo_id: str, model_id: str):
_download_controls[model_id] = {
"pause": threading.Event(),
"cancel": threading.Event()
}
async def generate():
async for chunk in _download_file_streamed(repo_id, model_id):
yield f"data: {json.dumps(chunk)}\n\n"
yield "data: [DONE]\n\n"
return StreamingResponse(
generate(),
media_type="text/event-stream",
headers={"Cache-Control": "no-cache"}
)
The download implementation includes resumable range requests, exponential backoff retries, and pause/cancel support through the _download_controls dictionary.
Model Registry Integration
All model endpoints interact directly with generator_registry:
# From GET /model/status
@router.get("/status")
async def model_status():
active = generator_registry.get_active_model_id()
info = generator_registry.get_model_info(active)
return {
"model_id": active,
"downloaded": info.downloaded,
"loaded": info.loaded,
"vram_mb": info.vram_usage
}
Optimize Router: Mesh Post-Processing
The optimize router in api/routers/optimize.py provides synchronous CPU-bound operations for mesh refinement. Unlike the generation router, it relies on FastAPI's automatic thread-pool execution for blocking operations.
Processing Capabilities
| Operation | Dependencies | Input/Output |
|---|---|---|
| Decimation | pymeshlab | OBJ/PLY → reduced face count |
| Smoothing | pymeshlab | Laplacian smoothing |
| Transform | trimesh | 4×4 matrix application |
| Format conversion | trimesh/pymeshlab | PLY → SPLAT, GLB, OBJ |
Decimation with Texture Preservation
The _decimate() function handles two code paths based on texture presence:
def _decimate(input_path: Path, target_faces: int, output_path: Path):
has_tex = _has_texture(input_path)
# Load into MeshSet
ms = pymeshlab.MeshSet()
ms.load_new_mesh(str(input_path))
if has_tex:
# Quadric edge collapse decimation with texture coordinate preservation
ms.apply_filter(
"meshing_decimation_quadric_edge_collapse",
targetfacenum=target_faces,
preservetopology=True,
texture=True
)
# Save as OBJ to retain texture reference
ms.save_current_mesh(str(output_path), save_textures=True)
else:
# PLY path: no texture handling needed
ms.apply_filter(
"meshing_decimation_quadric_edge_collapse",
targetfacenum=target_faces
)
ms.save_current_mesh(str(output_path))
Workspace Integration
All optimized outputs are written to paths under WORKSPACE_DIR, with the /serve-file endpoint returning proper MIME types:
@router.get("/serve-file")
async def serve_file(path: str):
full_path = WORKSPACE_DIR / path
mime = "model/gltf-binary" if path.endswith(".glb") else "application/octet-stream"
return FileResponse(full_path, media_type=mime)
Workflow_Runs Router: UI-Friendly Abstraction
The workflow_runs router in api/routers/workflow_runs.py provides a thin façade over generation functionality, designed for frontend workflow orchestration.
Architectural Relationship
The router reuses generation infrastructure directly:
# From workflow_runs.py
from routers.generation import (
_jobs,
_cancel_events,
_run_generation,
WORKSPACE_DIR
)
WorkflowRunStatus Model
The key distinction is the response schema, which includes UI-specific fields:
class WorkflowRunStatus(BaseModel):
run_id: str
status: Literal["pending", "running", "completed", "cancelled", "error"]
progress: int
step: Optional[str]
scene_candidate: Optional[SceneCandidate] # UI navigation data
error: Optional[str]
Endpoint Mapping
| Workflow Endpoint | Delegates To | Added Value |
|---|---|---|
POST /from-image |
generation logic | Returns run_id, initializes scene_candidate |
GET /{run_id} |
_jobs lookup |
Transforms JobStatus → WorkflowRunStatus |
POST /{run_id}/cancel |
cancellation logic | Same subprocess-kill behavior |
The GET /{run_id} endpoint demonstrates the transformation:
@router.get("/{run_id}")
async def get_workflow_run(run_id: str):
if run_id not in _jobs:
raise HTTPException(404, "Run not found")
job = _jobs[run_id]
scene = None
if job.output_url:
# Parse workspace path from output URL
relative_path = job.output_url.replace("/workspace/", "")
scene = SceneCandidate(
path=relative_path,
format="glb" if relative_path.endswith(".glb") else "splat"
)
return WorkflowRunStatus(
run_id=run_id,
status=job.status,
progress=job.progress,
step=job.step,
scene_candidate=scene,
error=job.error
)
Central Application Configuration
The FastAPI application factory in api/main.py orchestrates router registration and shared concerns:
Lifespan Management
@asynccontextmanager
async def lifespan(app: FastAPI):
# Startup: initialize all generator adapters
await generator_registry.initialize()
yield
# Shutdown: clean release of GPU memory
await generator_registry.unload_all()
Static File Serving
A catch-all endpoint serves workspace assets:
@app.get("/workspace/{full_path:path}")
async def serve_workspace(full_path: str):
target = WORKSPACE_DIR / full_path
# Security: ensure path stays within WORKSPACE_DIR
if not target.resolve().is_relative_to(WORKSPACE_DIR.resolve()):
raise HTTPException(403, "Access denied")
return FileResponse(target)
API Usage Examples
Submit and Monitor Generation
# Start job
curl -X POST "http://localhost:8000/generate/from-image" \
-F "image=@product.jpg" \
-F "model_id=sf3d" \
-F "collection=BatchA"
# Poll status
curl "http://localhost:8000/generate/status/{job_id}"
Stream Model Download
curl "http://localhost:8000/model/hf-download?repo_id=stabilityai/stable-fast-3d&model_id=sf3d"
Optimize Result
curl -X POST "http://localhost:8000/optimize/mesh" \
-H "Content-Type: application/json" \
-d '{"path":"BatchA/job_xxx.glb","target_faces":3000}'
Summary
The Modly FastAPI backend router architecture demonstrates several production patterns for desktop-embedded AI services:
- Router-service separation — HTTP concerns isolated from heavy processing
- Thread-safe state —
threading.Eventand module-level dictionaries coordinate across async/await and thread-pool boundaries - Streaming downloads — SSE for real-time progress on long-running HuggingFace operations
- Façade pattern — workflow_runs router reuses generation internals with UI-optimized contracts
- Workspace sandbox — unified file storage with path-traversal-protected static serving
All four routers mount on a single FastAPI instance with predictable prefixes, sharing the generator_registry service for model lifecycle management and WORKSPACE_DIR for asset persistence.
Frequently Asked Questions
How does the generation router handle cancellation?
The generation router uses a two-stage cancellation mechanism. First, setting a threading.Event in _cancel_events[job_id] signals the generator to stop at the next checkpoint. Second, if the generator spawned a subprocess, the /cancel/{job_id} endpoint calls gen._proc.kill() for immediate termination. The job status transitions to "cancelled" regardless of which mechanism succeeds.
What is the difference between the generation and workflow_runs routers?
The workflow_runs router is a thin abstraction layer over generation router internals. Both use the same _jobs dictionary, _run_generation() function, and cancellation logic. The workflow router transforms JobStatus into WorkflowRunStatus to include UI-specific fields like scene_candidate, while the generation router exposes raw job metadata optimized for polling clients.
How does the model router stream HuggingFace downloads?
The model router implements Server-Sent Events (SSE) through FastAPI's StreamingResponse. The _download_file_streamed() async generator yields JSON fragments with download progress, which the endpoint wraps in data: protocol format. The client receives real-time updates without WebSocket overhead, and the download supports pause/resume through module-level _download_controls events.
Why are optimize router endpoints synchronous?
The optimize router uses synchronous functions because mesh processing through pymeshlab and trimesh releases the GIL but performs CPU-intensive work unsuitable for async/await. FastAPI automatically runs these endpoints in a thread pool via run_in_threadpool, preventing event loop blocking while leveraging multi-core processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →