Key Configuration Options for Modal Serverless GPU Workloads: A Complete Guide
Configure Modal serverless GPU workloads using the @app.function decorator with parameters like gpu, cpu, memory, and autoscaling settings to control hardware allocation, concurrency, and cost across distributed training and inference tasks.
Modal provides a serverless platform for running Python code on GPUs with fine-grained control over hardware resources. When defining GPU-enabled functions in the K-Dense-AI/scientific-agent-skills repository, you specify configuration options through the @app.function decorator to optimize for performance, cost, and scalability. Understanding these configuration parameters is essential for deploying production machine learning workloads efficiently.
Selecting and Configuring GPU Types
Modal offers a catalog of GPU types including T4 (16 GB), L40S (48 GB), A100-80GB (80 GB), H100 (80 GB), and B200+ (192 GB). You specify the GPU type using the gpu parameter in @app.function, with options ranging from budget inference to large-scale training configurations.
Basic GPU Requests
Request a single GPU by passing the type identifier as a string. According to scientific-skills/modal/references/gpu.md (lines 33‑38), a basic configuration looks like this:
@app.function(gpu="H100")
def train():
import torch
assert torch.cuda.is_available()
print(f"Using: {torch.cuda.get_device_name(0)}")
Multi-GPU Configurations
For distributed workloads, append :count to request multiple GPUs on the same node. As shown in lines 83‑90 of the GPU reference:
@app.function(gpu="H100:4")
def distributed():
import torch
print(f"GPUs available: {torch.cuda.device_count()}")
Fallback Chains and Version Locking
Define fallback chains using an ordered list so Modal tries alternative GPUs if your preferred type is unavailable (lines 101‑106):
@app.function(gpu=["H100", "A100-80GB", "L40S"])
def flexible():
...
Prevent auto-upgrades to newer GPU hardware by appending an exclamation mark (lines 15‑20):
@app.function(gpu="H100!")
def must_use_h100():
...
Allocating CPU, Memory, and Storage Resources
Combine GPU requests with CPU cores, RAM, and ephemeral disk to support data preprocessing and model loading. As documented in scientific-skills/modal/references/resources.md (lines 5‑15), specify physical CPU cores with cpu (float values), RAM in MiB with memory, and temporary storage with ephemeral_disk:
@app.function(
cpu=8.0, # 8 physical cores
memory=32768, # 32 GiB RAM
gpu="L40S", # L40S GPU
ephemeral_disk=204800, # 200 GiB temp storage
timeout=7200, # 2 h max runtime
max_containers=50,
min_containers=1,
)
def full_pipeline(data_path: str):
...
Autoscaling and Container Pool Management
Control serverless scaling behavior using parameters defined in scientific-skills/modal/references/scaling.md (lines 26‑33). The container pool configuration determines latency and cost trade-offs:
@app.function(
max_containers=100, # Upper limit
min_containers=2, # Keep 2 warm containers
buffer_containers=5, # Extra for burst traffic
scaledown_window=300, # 5 min idle before shutdown
)
def handle_request(data):
...
Key trade-offs include:
min_containers: Lower latency but higher idle costbuffer_containers: Reduces queuing but increases costscaledown_window: Longer windows save cost but may increase cold-start times
Concurrency, Batching, and Throughput Optimization
Handle multiple concurrent requests on a single GPU using decorators that maximize hardware utilization without oversubscription.
Concurrent Input Processing
Apply @modal.concurrent to serve multiple I/O-bound requests simultaneously (lines 9‑13):
@app.function(gpu="L40S")
@modal.concurrent(max_inputs=10)
async def predict(text: str):
result = await model.predict_async(text)
return result
Dynamic Batching
Use @modal.batched to consolidate inputs before GPU invocation, improving throughput for inference workloads (lines 33‑38):
@app.function(gpu="L40S")
@modal.batched(max_batch_size=32, wait_ms=100)
async def batch_predict(texts: list[str]):
embeddings = model.encode(texts)
return list(embeddings)
ASGI Web Endpoints
Deploy GPU-accelerated models behind FastAPI interfaces using @modal.asgi_app():
import modal
app = modal.App()
@app.function(gpu="L40S")
@modal.asgi_app()
def inference():
import fastapi, torch, torchvision
model = torchvision.models.resnet50(pretrained=True).eval().cuda()
api = fastapi.FastAPI()
@api.post("/predict")
async def predict(image: bytes):
tensor = torchvision.io.decode_image(image).float().cuda()
with torch.no_grad():
logits = model(tensor.unsqueeze(0))
return {"label": logits.argmax().item()}
return api
Multi-GPU Training Patterns
Modal integrates with distributed training frameworks through multi-GPU function configurations.
PyTorch Distributed Data Parallel (DDP)
Initialize process groups using NCCL backend across allocated GPUs (lines 37‑46):
@app.function(gpu="H100:4", image=image, timeout=86400)
def train_distributed():
import torch, torch.distributed as dist, os
dist.init_process_group(backend="nccl")
local_rank = int(os.getenv("LOCAL_RANK", 0))
device = torch.device(f"cuda:{local_rank}")
# … training loop …
PyTorch Lightning Spawn
For frameworks requiring subprocess spawning, execute training scripts directly (lines 56‑60):
@app.function(gpu="H100:4", image=image)
def train():
import subprocess
subprocess.run(["python", "train_script.py"], check=True)
Summary
- GPU Selection: Use
gpu="TYPE"for single GPUs,gpu="TYPE:count"for multi-GPU nodes, and lists for fallback chains; append!to prevent auto-upgrades - Resource Allocation: Combine
cpu(float cores),memory(MiB),ephemeral_disk(MiB), andtimeout(seconds) to support your workload - Autoscaling: Tune
min_containers,max_containers,buffer_containers, andscaledown_windowto balance latency against cost - Throughput Optimization: Apply
@modal.concurrentfor parallel I/O and@modal.batchedfor inference batching - Distributed Training: Configure
gpu="TYPE:count"and initialize NCCL process groups for PyTorch DDP workloads
Frequently Asked Questions
How do I prevent Modal from automatically upgrading my GPU hardware?
Append an exclamation mark to your GPU string, such as gpu="H100!". According to lines 15‑20 of scientific-skills/modal/references/gpu.md, this configuration locks the function to the specific GPU version and prevents automatic upgrades to newer hardware generations.
What is the difference between min_containers and buffer_containers in Modal autoscaling?
min_containers maintains a baseline number of warm containers ready to accept requests, reducing cold-start latency but incurring idle costs. buffer_containers provisions additional containers beyond current demand to handle burst traffic, reducing queuing time during load spikes. Configure both in the @app.function decorator as shown in scientific-skills/modal/references/scaling.md (lines 26‑33).
When should I use @modal.batched instead of @modal.concurrent for GPU inference?
Use @modal.batched when your model benefits from processing multiple inputs simultaneously in a single forward pass, such as embedding generation or image classification batches. Use @modal.concurrent when handling I/O-bound operations or independent requests that don't benefit from tensor concatenation. Batching improves GPU utilization for compute-heavy inference, while concurrency optimizes throughput for latency-sensitive, asynchronous workloads.
How do I configure a fallback chain for GPU availability?
Pass an ordered list to the gpu parameter, such as gpu=["H100", "A100-80GB", "L40S"]. Modal attempts to allocate the first GPU type, falling back to subsequent options if the preferred type is unavailable. This pattern, documented in lines 101‑106 of the GPU reference, ensures workload execution during high-demand periods when specific GPU types face capacity constraints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →