# Requirements for GPU Pools in Agent Substrate: Architecture, Configuration, and Limitations

> Configure GPU pools in Agent Substrate by meeting requirements for gvisor, ateom-gvisor binary, nvidia.com/gpu limits, and host node NVIDIA toolkit setup for actor workloads.

- Repository: [Agent Substrate/substrate](https://github.com/agent-substrate/substrate)
- Tags: architecture
- Published: 2026-08-22

---

**GPU pools in Agent Substrate require the `gvisor` sandbox class, a glibc-compiled `ateom-gvisor` binary, explicit `nvidia.com/gpu` resource limits, and proper NVIDIA toolkit configuration on host nodes to enable NVIDIA GPU passthrough for actor workloads.**

Agent Substrate is an open-source platform for running sandboxed actors at scale, and GPU pools represent a specialized **WorkerPool** implementation designed for compute-intensive workloads. Unlike standard CPU pools, GPU pools enforce strict architectural constraints to ensure safe GPU device passthrough through gVisor. This guide examines the technical requirements derived from the `agent-substrate/substrate` source code, including CRD validation rules in [`pkg/api/v1alpha1/workerpool_types.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_types.go), runtime dependencies, and configuration patterns necessary for production deployment.

## Core Architectural Requirements

Before deploying a GPU-enabled workload, the underlying infrastructure and binary artifacts must satisfy specific architectural constraints enforced by the Substrate controller and runtime.

### Sandbox Class Constraint

Every GPU pool must specify `sandboxClass: gvisor` in the WorkerPool specification. According to the **XValidation rules** in [`pkg/api/v1alpha1/workerpool_types.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_types.go), the controller explicitly validates that GPU pools use the gVisor sandbox implementation, as GPU passthrough is only implemented for this runtime. The validation prevents creation of GPU pools with alternative sandbox classes like `runc` or `wasm`, ensuring the specialized gVisor device forwarding paths are available.

### glibc-Compiled ateom-gvisor Binary

The GPU device plugin expects a glibc runtime environment, making the standard musl build of **ateom-gvisor** incompatible with NVIDIA driver libraries. As documented in [`docs/api-guide.md`](https://github.com/agent-substrate/substrate/blob/main/docs/api-guide.md), production GPU pools require a specifically compiled glibc variant of the `ateom-gvisor` binary to properly load `libcuda.so` and other NVIDIA driver components. This requirement stems from underlying CUDA toolkit dependencies that link against glibc symbols unavailable in musl libc implementations.

### Node Scheduling Requirements

Worker pods must be scheduled onto nodes physically equipped with NVIDIA GPUs. This requires:

- **Node selectors** targeting GPU-enabled nodes (e.g., `cloud.google.com/gke-accelerator: nvidia-tesla-t4`)
- **Tolerations** allowing pods to schedule on nodes tainted with `nvidia.com/gpu: NoSchedule`
- **Atelet deployment** on GPU nodes with matching tolerations to handle the actual GPU passthrough operations

The `atelet` process performs the low-level GPU device attachment and must co-reside on the same node as the GPU hardware, requiring you to add matching tolerations to its DaemonSet configuration.

## Resource and Environment Configuration

Beyond architectural prerequisites, GPU pools require specific Kubernetes resource definitions and environment variables to trigger the NVIDIA device plugin correctly.

### GPU Resource Limits and Requests

Each GPU pool must define explicit GPU resource limits using the `nvidia.com/gpu` extended resource. The **XValidation rules** in [`pkg/api/v1alpha1/workerpool_types.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_types.go) enforce that GPU pools specify both requests and limits for GPU resources, as Kubernetes forbids an extended-resource request without a matching limit. A typical configuration specifies:

- `limits.nvidia.com/gpu: "1"` (required)
- `requests.nvidia.com/gpu: "1"` (recommended)

Omitting the GPU limit prevents the Kubernetes scheduler from assigning the pod to a GPU-enabled node and blocks the NVIDIA device plugin from allocating the physical device.

### NVIDIA Toolkit Host Path

The environment variable **ATE_NVIDIA_TOOLKIT_HOST_PATH** must point to the host-installed NVIDIA toolkit directory, typically `/usr/local/nvidia` or `/usr/lib/nvidia`. This path provides the driver libraries and CUDA binaries that the **nvidia-ctk** (NVIDIA Container Toolkit) injects into the gVisor sandbox. According to [`docs/api-guide.md`](https://github.com/agent-substrate/substrate/blob/main/docs/api-guide.md), this variable ensures that `libcuda.so`, device nodes, and driver headers are properly mounted into the sandbox filesystem before actor execution begins.

### Container Device Interface (CDI) Integration

The runtime relies on `nvidia-ctk` to generate CDI specifications that inject GPU device nodes, driver libraries, and environment variables into the sandbox. As implemented in [`cmd/ateom-gvisor/gpu.go`](https://github.com/agent-substrate/substrate/blob/main/cmd/ateom-gvisor/gpu.go), the runtime detects host GPU devices using the glob pattern `/dev/nvidia[0-9]*` and coordinates with the toolkit to establish the passthrough path. The [`gpu.go`](https://github.com/agent-substrate/substrate/blob/main/gpu.go) file defines `gpuDeviceGlob = "/dev/nvidia[0-9]*"` to identify available GPU devices on the host node before initialization.

## Runtime Limitations and Constraints

GPU pools impose specific operational constraints that differ from standard CPU-based actor pools.

### CUDA Context Serialization Limitations

Due to gVisor's inability to serialize GPU state, actors utilizing GPU pools **cannot be suspended while a CUDA context remains open**. This limitation, documented in [`docs/api-guide.md`](https://github.com/agent-substrate/substrate/blob/main/docs/api-guide.md), means that checkpointing and restore operations are only safe when the actor has released all CUDA contexts and GPU memory. Attempting to suspend an actor during active GPU computation results in undefined behavior and potential data corruption because the runtime cannot capture the GPU's volatile state.

### Device Detection Mechanism

The runtime automatically detects GPU presence by scanning for `/dev/nvidia*` device nodes before initializing the sandbox. The detection logic in [`cmd/ateom-gvisor/gpu.go`](https://github.com/agent-substrate/substrate/blob/main/cmd/ateom-gvisor/gpu.go) uses the `gpuDeviceGlob` constant to validate that the host node actually exposes NVIDIA devices before attempting passthrough, preventing scheduling failures on non-GPU nodes.

## Practical Configuration Examples

The following YAML configurations demonstrate a compliant GPU WorkerPool and corresponding ActorTemplate, incorporating the validation requirements from [`pkg/api/v1alpha1/workerpool_validation_test.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_validation_test.go).

### GPU-Enabled WorkerPool

```yaml
apiVersion: substrate.dev/v1alpha1
kind: WorkerPool
metadata:
  name: gpu-pool
spec:
  sandboxClass: gvisor
  nodeSelector:
    cloud.google.com/gke-accelerator: nvidia-tesla-t4
  tolerations:
  - key: nvidia.com/gpu
    operator: Exists
    effect: NoSchedule
  template:
    resources:
      limits:
        nvidia.com/gpu: "1"
      requests:
        nvidia.com/gpu: "1"
    env:
    - name: ATE_NVIDIA_TOOLKIT_HOST_PATH
      value: /usr/local/nvidia

```

This configuration satisfies all CRD validation rules, including the mandatory gVisor sandbox class and GPU resource specifications enforced by the Webhook and XValidation logic in [`pkg/api/v1alpha1/workerpool_types.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_types.go).

### GPU ActorTemplate

```yaml
apiVersion: substrate.dev/v1alpha1
kind: ActorTemplate
metadata:
  name: gpu-worker
spec:
  containers:
  - name: main
    image: myregistry.com/gpu-app:latest
    command: ["python", "train.py"]
    env:
    - name: CUDA_VISIBLE_DEVICES
      value: "0"

```

The ActorTemplate requires no additional GPU-specific configuration; the WorkerPool's resource limits and `ATE_NVIDIA_TOOLKIT_HOST_PATH` environment variable handle device injection automatically via the `nvidia-ctk` integration.

## Summary

- **GPU pools require the `gvisor` sandbox class**, enforced by XValidation rules in [`pkg/api/v1alpha1/workerpool_types.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_types.go).
- **A glibc-compiled `ateom-gvisor` binary is mandatory** for loading NVIDIA driver libraries, as musl builds are incompatible with CUDA.
- **Explicit `nvidia.com/gpu` resource limits and requests** are required to trigger Kubernetes GPU scheduling and satisfy the extended-resource validation constraints.
- **The `ATE_NVIDIA_TOOLKIT_HOST_PATH` environment variable** must reference the host NVIDIA toolkit installation for proper driver injection via `nvidia-ctk`.
- **GPU actors cannot be suspended** while holding open CUDA contexts due to gVisor's inability to serialize GPU state, as documented in [`docs/api-guide.md`](https://github.com/agent-substrate/substrate/blob/main/docs/api-guide.md).
- **Node tolerations and selectors** are required to ensure `atelet` and worker pods schedule exclusively on GPU-equipped nodes, with the runtime detecting devices via `/dev/nvidia[0-9]*` in [`cmd/ateom-gvisor/gpu.go`](https://github.com/agent-substrate/substrate/blob/main/cmd/ateom-gvisor/gpu.go).

## Frequently Asked Questions

### Can I use GPU pools with the default musl build of ateom-gvisor?

No. GPU pools explicitly require a glibc-compiled `ateom-gvisor` binary because the NVIDIA driver libraries (`libcuda.so`) link against glibc symbols. The default musl build cannot load these proprietary NVIDIA libraries, resulting in runtime failures when actors attempt CUDA initialization. You must compile or obtain the glibc variant of the binary for GPU workloads to satisfy the requirements documented in [`docs/api-guide.md`](https://github.com/agent-substrate/substrate/blob/main/docs/api-guide.md).

### Why does my GPU WorkerPool fail validation without explicit resource requests?

The XValidation rules in [`pkg/api/v1alpha1/workerpool_types.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_types.go) enforce that GPU pools specify both resource limits and requests for `nvidia.com/gpu`. Kubernetes extended resources require matching limits for proper scheduling, and the Substrate controller validates this constraint to prevent scheduling failures on nodes lacking GPU capacity. The validation explicitly checks that requests do not exceed limits while ensuring both fields are present.

### What happens if I attempt to suspend a GPU actor during CUDA execution?

Suspending a GPU actor while a CUDA context remains active is unsupported and will likely cause data corruption or runtime crashes. According to the implementation in [`cmd/ateom-gvisor/gpu.go`](https://github.com/agent-substrate/substrate/blob/main/cmd/ateom-gvisor/gpu.go) and documentation in [`docs/api-guide.md`](https://github.com/agent-substrate/substrate/blob/main/docs/api-guide.md), gVisor cannot serialize GPU state or CUDA contexts. You must ensure the actor releases all GPU resources and CUDA contexts before initiating a suspend operation to maintain snapshot integrity.

### How does the runtime detect available GPUs on the host node?

The `ateom-gvisor` runtime detects GPUs by scanning for device nodes matching the glob pattern `/dev/nvidia[0-9]*`, as defined by the `gpuDeviceGlob` constant in [`cmd/ateom-gvisor/gpu.go`](https://github.com/agent-substrate/substrate/blob/main/cmd/ateom-gvisor/gpu.go). This detection occurs during sandbox initialization to validate that the host node actually exposes NVIDIA devices before attempting passthrough, ensuring compatibility with the `nvidia-ctk` injection process and preventing failures on non-GPU nodes.