How Modly Handles API Requests and Responses: A Deep Dive into the FastAPI Backend
Modly handles API requests and responses using FastAPI routers with automatic Pydantic validation, offloads heavy inference to background workers tracked via in-memory job dictionaries, and streams large file downloads through Server-Sent Events while managing model lifecycle through a singleton registry.
Modly is an open-source AI model serving platform built on FastAPI. Understanding how Modly handles API requests and responses reveals a sophisticated architecture designed for asynchronous, long-running inference workloads. The implementation spans multiple modules under api/routers/ and api/services/, utilizing background workers, real-time progress streaming, and a centralized generator registry to coordinate between HTTP endpoints and model execution.
FastAPI Router Architecture
Modly organizes all HTTP endpoints into dedicated router modules under api/routers/. Each router focuses on a specific domain, such as generation or model management, and registers path operations using FastAPI’s APIRouter class.
The generation router (api/routers/generation.py) handles image-to-3D generation endpoints, while the model router (api/routers/model.py) manages model switching, status checks, and downloads. Both routers rely on FastAPI’s dependency injection system to access the singleton generator_registry instance, ensuring consistent state across the application.
Request Validation and Input Handling
All incoming requests undergo strict validation through FastAPI’s automatic parsing. In api/routers/generation.py, the /generation/from-image endpoint declares parameters with type hints like UploadFile, Form, and Optional[str], allowing FastAPI to automatically parse multipart data, JSON bodies, and query strings.
The router performs explicit validation beyond type checking. For image uploads, it verifies content types:
if not image.content_type.startswith("image/"):
raise HTTPException(400, "File must be an image")
For collection names, it sanitizes inputs to prevent path traversal attacks using regex validation:
if not collection or _re.search(r'[/:*?"<>|\\]', collection):
collection = "Default"
Background Job Execution and Tracking
Heavy inference workloads are offloaded from the HTTP thread to prevent timeouts. When a client submits a generation request, the endpoint creates a background task and returns immediately with a job identifier.
The Generation Flow
In api/routers/generation.py, the _run_generation function (lines 25-190) handles the actual inference:
- Job Creation – The endpoint initializes a
JobStatusPydantic model and stores it in the module-level_jobsdictionary - Background Task –
background_tasks.add_task(_run_generation, job_id, image_bytes, full_params, collection)defers execution - Model Loading – The function calls
generator_registry.get_active()to load the model into memory if not already present - Execution – Inference runs inside an
asyncioexecutor to prevent blocking - Progress Updates – A separate progress thread (
smooth_progress) updates theJobStatusin_jobswith completion percentages and output URLs
The heavy lifting remains non-blocking, allowing the API to accept concurrent requests while GPUs process previous jobs.
Submitting and Polling Jobs
Clients receive a UUID to track progress:
import requests
url = "http://localhost:8000/generation/from-image"
files = {"image": open("my_photo.png", "rb")}
data = {
"model_id": "sf3d",
"collection": "MyWorks",
"remesh": "quad",
"enable_texture": "true",
"texture_resolution": "1024",
"params": '{"some_param": 42}'
}
resp = requests.post(url, files=files, data=data)
job_id = resp.json()["job_id"]
Poll the status endpoint to retrieve progress:
import time, requests
status_url = f"http://localhost:8000/generation/status/{job_id}"
while True:
r = requests.get(status_url)
info = r.json()
print(info["status"], info["progress"])
if info["status"] in ("done", "cancelled"):
print("Result URL:", info.get("output_url"))
break
time.sleep(1)
Streaming Responses for Large Downloads
For operations like HuggingFace model downloads, Modly uses StreamingResponse with async generators to emit Server-Sent Events (SSE). The /model/hf-download endpoint in api/routers/model.py creates a per-model control object (_download_control) containing pause and cancel events.
The endpoint streams progress as JSON-formatted SSE:
import sseclient, requests, json
params = {
"repo_id": "myorg/my-model",
"model_id": "my_new_model"
}
response = requests.get("http://localhost:8000/model/hf-download", params=params, stream=True)
client = sseclient.SSEClient(response)
for event in client.events():
data = json.loads(event.data)
print(f"{data['percent']}% – {data.get('status', '')}")
This approach prevents memory exhaustion on large files while providing real-time feedback to clients.
Cancellation and Cleanup Mechanisms
Modly implements graceful job cancellation through the /cancel/{job_id} endpoint (lines 100-122 in api/routers/generation.py). The system maintains three module-level data structures:
_jobs– Dictionary mapping job IDs toJobStatusinstances_cancel_events– Dictionary ofasyncio.Eventobjects for signaling cancellation_cancelled– Set of cancelled job IDs
When cancellation is requested:
- The endpoint adds the job ID to
_cancelled - It triggers the corresponding event in
_cancel_events - If the generator spawned a subprocess, it calls
gen._proc.kill()to terminate the process - The
_run_generationcoroutine checks these flags periodically and raisesGenerationCancelled(defined inservices/generators/base.py) to abort gracefully
A periodic purge removes completed jobs from memory after a TTL expires, preventing unbounded memory growth.
Generator Registry and Model Management
The api/services/generator_registry.py file implements a singleton pattern (generator_registry = GeneratorRegistry()) that acts as the bridge between HTTP routers and model implementations.
Discovery – The _discover_extensions() method scans EXTENSIONS_DIR for manifest.json and generator.py files, building a mapping of model IDs to generator classes.
Instantiation – The initialize() method creates generator instances. Extensions requiring virtual environments are wrapped in ExtensionProcess (subprocess mode), while others load directly (legacy mode).
State Management – The registry tracks the active model via _active_id (defaulting to "sf3d") and exposes helper methods used by routers:
active_status()– Returns current model load stateswitch_model(model_id)– Changes the active modelget_active()– Returns the current generator instanceunload_all()– Frees VRAM/RAM
These methods enable the model router endpoints (/model/status, /model/switch, /model/unload-all) to control model lifecycle without directly accessing generator internals.
Response Schemas and Error Handling
Modly uses Pydantic models defined in api/schemas/ for response serialization. The JobStatus schema includes fields for job_id, status, progress, step, output_url, and timestamps, ensuring clients receive typed JSON payloads without manual serialization.
Error handling relies on FastAPI’s HTTPException for validation failures (400), missing resources (404), and server errors. The generation router catches GenerationCancelled specifically to distinguish user-initiated aborts from system failures, returning appropriate status codes to the client.
Summary
- Modly uses FastAPI with domain-specific routers in
api/routers/to organize endpoints for generation and model management. - Heavy inference runs in background tasks via
BackgroundTasksandasyncioexecutors, with job state tracked in module-level dictionaries (_jobs,_cancel_events). - Real-time updates use Server-Sent Events for streaming downloads and JSON endpoints for job status polling.
- Graceful cancellation is implemented through event flags and subprocess termination in the
/cancel/{job_id}endpoint. - Model lifecycle is managed by a singleton
GeneratorRegistryinapi/services/generator_registry.pythat handles discovery, instantiation, and switching.
Frequently Asked Questions
How does Modly prevent HTTP timeouts during long model inference?
Modly offloads heavy work to background tasks using FastAPI’s BackgroundTasks class. When a client calls /generation/from-image, the endpoint immediately returns a job ID while _run_generation executes asynchronously. This prevents the HTTP connection from remaining open during GPU-intensive operations that might take minutes to complete.
What data structures does Modly use to track job progress?
Modly maintains three module-level dictionaries in api/routers/generation.py: _jobs stores JobStatus Pydantic models containing progress and output URLs, _cancel_events holds asyncio.Event objects for cancellation signaling, and _completed_at tracks finish timestamps. These structures enable real-time polling and graceful shutdown of running jobs.
How does Modly handle large file downloads without memory issues?
Rather than buffering entire files, Modly uses StreamingResponse with async generators in the /model/hf-download endpoint. This streams HuggingFace model files chunk-by-chunk while emitting Server-Sent Events to report download percentages, keeping memory usage constant regardless of file size.
Can clients cancel a generation job after submission?
Yes, clients can POST to /generation/cancel/{job_id} to abort running jobs. The endpoint sets a cancellation flag in the _cancelled set and triggers the corresponding event in _cancel_events. If the job spawned a subprocess, the system calls kill() on the process handle, and the generation coroutine checks these flags to raise GenerationCancelled and exit cleanly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →