# Key Configuration Options for Modal Serverless GPU Workloads: A Complete Guide

> Master Modal serverless GPU workloads. Explore key configuration options like gpu, cpu, and memory with the @app.function decorator for efficient distributed training and inference. Optimize cost and scale.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: how-to-guide
- Published: 2026-05-14

---

**Configure Modal serverless GPU workloads using the `@app.function` decorator with parameters like `gpu`, `cpu`, `memory`, and autoscaling settings to control hardware allocation, concurrency, and cost across distributed training and inference tasks.**

Modal provides a serverless platform for running Python code on GPUs with fine-grained control over hardware resources. When defining GPU-enabled functions in the `K-Dense-AI/scientific-agent-skills` repository, you specify configuration options through the `@app.function` decorator to optimize for performance, cost, and scalability. Understanding these configuration parameters is essential for deploying production machine learning workloads efficiently.

## Selecting and Configuring GPU Types

Modal offers a catalog of GPU types including **T4** (16 GB), **L40S** (48 GB), **A100-80GB** (80 GB), **H100** (80 GB), and **B200+** (192 GB). You specify the GPU type using the `gpu` parameter in `@app.function`, with options ranging from budget inference to large-scale training configurations.

### Basic GPU Requests

Request a single GPU by passing the type identifier as a string. According to [`scientific-skills/modal/references/gpu.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/modal/references/gpu.md) (lines 33‑38), a basic configuration looks like this:

```python
@app.function(gpu="H100")
def train():
    import torch
    assert torch.cuda.is_available()
    print(f"Using: {torch.cuda.get_device_name(0)}")

```

### Multi-GPU Configurations

For distributed workloads, append `:count` to request multiple GPUs on the same node. As shown in lines 83‑90 of the GPU reference:

```python
@app.function(gpu="H100:4")
def distributed():
    import torch
    print(f"GPUs available: {torch.cuda.device_count()}")

```

### Fallback Chains and Version Locking

Define fallback chains using an ordered list so Modal tries alternative GPUs if your preferred type is unavailable (lines 101‑106):

```python
@app.function(gpu=["H100", "A100-80GB", "L40S"])
def flexible():
    ...

```

Prevent auto-upgrades to newer GPU hardware by appending an exclamation mark (lines 15‑20):

```python
@app.function(gpu="H100!")
def must_use_h100():
    ...

```

## Allocating CPU, Memory, and Storage Resources

Combine GPU requests with CPU cores, RAM, and ephemeral disk to support data preprocessing and model loading. As documented in [`scientific-skills/modal/references/resources.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/modal/references/resources.md) (lines 5‑15), specify physical CPU cores with `cpu` (float values), RAM in MiB with `memory`, and temporary storage with `ephemeral_disk`:

```python
@app.function(
    cpu=8.0,                # 8 physical cores

    memory=32768,           # 32 GiB RAM

    gpu="L40S",             # L40S GPU

    ephemeral_disk=204800,  # 200 GiB temp storage

    timeout=7200,           # 2 h max runtime

    max_containers=50,
    min_containers=1,
)
def full_pipeline(data_path: str):
    ...

```

## Autoscaling and Container Pool Management

Control serverless scaling behavior using parameters defined in [`scientific-skills/modal/references/scaling.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/modal/references/scaling.md) (lines 26‑33). The container pool configuration determines latency and cost trade-offs:

```python
@app.function(
    max_containers=100,    # Upper limit

    min_containers=2,      # Keep 2 warm containers

    buffer_containers=5,   # Extra for burst traffic

    scaledown_window=300, # 5 min idle before shutdown

)
def handle_request(data):
    ...

```

**Key trade-offs** include:

- **`min_containers`**: Lower latency but higher idle cost
- **`buffer_containers`**: Reduces queuing but increases cost
- **`scaledown_window`**: Longer windows save cost but may increase cold-start times

## Concurrency, Batching, and Throughput Optimization

Handle multiple concurrent requests on a single GPU using decorators that maximize hardware utilization without oversubscription.

### Concurrent Input Processing

Apply `@modal.concurrent` to serve multiple I/O-bound requests simultaneously (lines 9‑13):

```python
@app.function(gpu="L40S")
@modal.concurrent(max_inputs=10)
async def predict(text: str):
    result = await model.predict_async(text)
    return result

```

### Dynamic Batching

Use `@modal.batched` to consolidate inputs before GPU invocation, improving throughput for inference workloads (lines 33‑38):

```python
@app.function(gpu="L40S")
@modal.batched(max_batch_size=32, wait_ms=100)
async def batch_predict(texts: list[str]):
    embeddings = model.encode(texts)
    return list(embeddings)

```

### ASGI Web Endpoints

Deploy GPU-accelerated models behind FastAPI interfaces using `@modal.asgi_app()`:

```python
import modal

app = modal.App()

@app.function(gpu="L40S")
@modal.asgi_app()
def inference():
    import fastapi, torch, torchvision
    model = torchvision.models.resnet50(pretrained=True).eval().cuda()
    api = fastapi.FastAPI()

    @api.post("/predict")
    async def predict(image: bytes):
        tensor = torchvision.io.decode_image(image).float().cuda()
        with torch.no_grad():
            logits = model(tensor.unsqueeze(0))
        return {"label": logits.argmax().item()}

    return api

```

## Multi-GPU Training Patterns

Modal integrates with distributed training frameworks through multi-GPU function configurations.

### PyTorch Distributed Data Parallel (DDP)

Initialize process groups using NCCL backend across allocated GPUs (lines 37‑46):

```python
@app.function(gpu="H100:4", image=image, timeout=86400)
def train_distributed():
    import torch, torch.distributed as dist, os
    dist.init_process_group(backend="nccl")
    local_rank = int(os.getenv("LOCAL_RANK", 0))
    device = torch.device(f"cuda:{local_rank}")
    # … training loop …

```

### PyTorch Lightning Spawn

For frameworks requiring subprocess spawning, execute training scripts directly (lines 56‑60):

```python
@app.function(gpu="H100:4", image=image)
def train():
    import subprocess
    subprocess.run(["python", "train_script.py"], check=True)

```

## Summary

- **GPU Selection**: Use `gpu="TYPE"` for single GPUs, `gpu="TYPE:count"` for multi-GPU nodes, and lists for fallback chains; append `!` to prevent auto-upgrades
- **Resource Allocation**: Combine `cpu` (float cores), `memory` (MiB), `ephemeral_disk` (MiB), and `timeout` (seconds) to support your workload
- **Autoscaling**: Tune `min_containers`, `max_containers`, `buffer_containers`, and `scaledown_window` to balance latency against cost
- **Throughput Optimization**: Apply `@modal.concurrent` for parallel I/O and `@modal.batched` for inference batching
- **Distributed Training**: Configure `gpu="TYPE:count"` and initialize NCCL process groups for PyTorch DDP workloads

## Frequently Asked Questions

### How do I prevent Modal from automatically upgrading my GPU hardware?

Append an exclamation mark to your GPU string, such as `gpu="H100!"`. According to lines 15‑20 of [`scientific-skills/modal/references/gpu.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/modal/references/gpu.md), this configuration locks the function to the specific GPU version and prevents automatic upgrades to newer hardware generations.

### What is the difference between `min_containers` and `buffer_containers` in Modal autoscaling?

**`min_containers`** maintains a baseline number of warm containers ready to accept requests, reducing cold-start latency but incurring idle costs. **`buffer_containers`** provisions additional containers beyond current demand to handle burst traffic, reducing queuing time during load spikes. Configure both in the `@app.function` decorator as shown in [`scientific-skills/modal/references/scaling.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/modal/references/scaling.md) (lines 26‑33).

### When should I use `@modal.batched` instead of `@modal.concurrent` for GPU inference?

Use **`@modal.batched`** when your model benefits from processing multiple inputs simultaneously in a single forward pass, such as embedding generation or image classification batches. Use **`@modal.concurrent`** when handling I/O-bound operations or independent requests that don't benefit from tensor concatenation. Batching improves GPU utilization for compute-heavy inference, while concurrency optimizes throughput for latency-sensitive, asynchronous workloads.

### How do I configure a fallback chain for GPU availability?

Pass an ordered list to the `gpu` parameter, such as `gpu=["H100", "A100-80GB", "L40S"]`. Modal attempts to allocate the first GPU type, falling back to subsequent options if the preferred type is unavailable. This pattern, documented in lines 101‑106 of the GPU reference, ensures workload execution during high-demand periods when specific GPU types face capacity constraints.