Requirements for GPU Pools in Agent Substrate: Architecture, Configuration, and Limitations
GPU pools in Agent Substrate require the gvisor sandbox class, a glibc-compiled ateom-gvisor binary, explicit nvidia.com/gpu resource limits, and proper NVIDIA toolkit configuration on host nodes to enable NVIDIA GPU passthrough for actor workloads.
Agent Substrate is an open-source platform for running sandboxed actors at scale, and GPU pools represent a specialized WorkerPool implementation designed for compute-intensive workloads. Unlike standard CPU pools, GPU pools enforce strict architectural constraints to ensure safe GPU device passthrough through gVisor. This guide examines the technical requirements derived from the agent-substrate/substrate source code, including CRD validation rules in pkg/api/v1alpha1/workerpool_types.go, runtime dependencies, and configuration patterns necessary for production deployment.
Core Architectural Requirements
Before deploying a GPU-enabled workload, the underlying infrastructure and binary artifacts must satisfy specific architectural constraints enforced by the Substrate controller and runtime.
Sandbox Class Constraint
Every GPU pool must specify sandboxClass: gvisor in the WorkerPool specification. According to the XValidation rules in pkg/api/v1alpha1/workerpool_types.go, the controller explicitly validates that GPU pools use the gVisor sandbox implementation, as GPU passthrough is only implemented for this runtime. The validation prevents creation of GPU pools with alternative sandbox classes like runc or wasm, ensuring the specialized gVisor device forwarding paths are available.
glibc-Compiled ateom-gvisor Binary
The GPU device plugin expects a glibc runtime environment, making the standard musl build of ateom-gvisor incompatible with NVIDIA driver libraries. As documented in docs/api-guide.md, production GPU pools require a specifically compiled glibc variant of the ateom-gvisor binary to properly load libcuda.so and other NVIDIA driver components. This requirement stems from underlying CUDA toolkit dependencies that link against glibc symbols unavailable in musl libc implementations.
Node Scheduling Requirements
Worker pods must be scheduled onto nodes physically equipped with NVIDIA GPUs. This requires:
- Node selectors targeting GPU-enabled nodes (e.g.,
cloud.google.com/gke-accelerator: nvidia-tesla-t4) - Tolerations allowing pods to schedule on nodes tainted with
nvidia.com/gpu: NoSchedule - Atelet deployment on GPU nodes with matching tolerations to handle the actual GPU passthrough operations
The atelet process performs the low-level GPU device attachment and must co-reside on the same node as the GPU hardware, requiring you to add matching tolerations to its DaemonSet configuration.
Resource and Environment Configuration
Beyond architectural prerequisites, GPU pools require specific Kubernetes resource definitions and environment variables to trigger the NVIDIA device plugin correctly.
GPU Resource Limits and Requests
Each GPU pool must define explicit GPU resource limits using the nvidia.com/gpu extended resource. The XValidation rules in pkg/api/v1alpha1/workerpool_types.go enforce that GPU pools specify both requests and limits for GPU resources, as Kubernetes forbids an extended-resource request without a matching limit. A typical configuration specifies:
limits.nvidia.com/gpu: "1"(required)requests.nvidia.com/gpu: "1"(recommended)
Omitting the GPU limit prevents the Kubernetes scheduler from assigning the pod to a GPU-enabled node and blocks the NVIDIA device plugin from allocating the physical device.
NVIDIA Toolkit Host Path
The environment variable ATE_NVIDIA_TOOLKIT_HOST_PATH must point to the host-installed NVIDIA toolkit directory, typically /usr/local/nvidia or /usr/lib/nvidia. This path provides the driver libraries and CUDA binaries that the nvidia-ctk (NVIDIA Container Toolkit) injects into the gVisor sandbox. According to docs/api-guide.md, this variable ensures that libcuda.so, device nodes, and driver headers are properly mounted into the sandbox filesystem before actor execution begins.
Container Device Interface (CDI) Integration
The runtime relies on nvidia-ctk to generate CDI specifications that inject GPU device nodes, driver libraries, and environment variables into the sandbox. As implemented in cmd/ateom-gvisor/gpu.go, the runtime detects host GPU devices using the glob pattern /dev/nvidia[0-9]* and coordinates with the toolkit to establish the passthrough path. The gpu.go file defines gpuDeviceGlob = "/dev/nvidia[0-9]*" to identify available GPU devices on the host node before initialization.
Runtime Limitations and Constraints
GPU pools impose specific operational constraints that differ from standard CPU-based actor pools.
CUDA Context Serialization Limitations
Due to gVisor's inability to serialize GPU state, actors utilizing GPU pools cannot be suspended while a CUDA context remains open. This limitation, documented in docs/api-guide.md, means that checkpointing and restore operations are only safe when the actor has released all CUDA contexts and GPU memory. Attempting to suspend an actor during active GPU computation results in undefined behavior and potential data corruption because the runtime cannot capture the GPU's volatile state.
Device Detection Mechanism
The runtime automatically detects GPU presence by scanning for /dev/nvidia* device nodes before initializing the sandbox. The detection logic in cmd/ateom-gvisor/gpu.go uses the gpuDeviceGlob constant to validate that the host node actually exposes NVIDIA devices before attempting passthrough, preventing scheduling failures on non-GPU nodes.
Practical Configuration Examples
The following YAML configurations demonstrate a compliant GPU WorkerPool and corresponding ActorTemplate, incorporating the validation requirements from pkg/api/v1alpha1/workerpool_validation_test.go.
GPU-Enabled WorkerPool
apiVersion: substrate.dev/v1alpha1
kind: WorkerPool
metadata:
name: gpu-pool
spec:
sandboxClass: gvisor
nodeSelector:
cloud.google.com/gke-accelerator: nvidia-tesla-t4
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
template:
resources:
limits:
nvidia.com/gpu: "1"
requests:
nvidia.com/gpu: "1"
env:
- name: ATE_NVIDIA_TOOLKIT_HOST_PATH
value: /usr/local/nvidia
This configuration satisfies all CRD validation rules, including the mandatory gVisor sandbox class and GPU resource specifications enforced by the Webhook and XValidation logic in pkg/api/v1alpha1/workerpool_types.go.
GPU ActorTemplate
apiVersion: substrate.dev/v1alpha1
kind: ActorTemplate
metadata:
name: gpu-worker
spec:
containers:
- name: main
image: myregistry.com/gpu-app:latest
command: ["python", "train.py"]
env:
- name: CUDA_VISIBLE_DEVICES
value: "0"
The ActorTemplate requires no additional GPU-specific configuration; the WorkerPool's resource limits and ATE_NVIDIA_TOOLKIT_HOST_PATH environment variable handle device injection automatically via the nvidia-ctk integration.
Summary
- GPU pools require the
gvisorsandbox class, enforced by XValidation rules inpkg/api/v1alpha1/workerpool_types.go. - A glibc-compiled
ateom-gvisorbinary is mandatory for loading NVIDIA driver libraries, as musl builds are incompatible with CUDA. - Explicit
nvidia.com/gpuresource limits and requests are required to trigger Kubernetes GPU scheduling and satisfy the extended-resource validation constraints. - The
ATE_NVIDIA_TOOLKIT_HOST_PATHenvironment variable must reference the host NVIDIA toolkit installation for proper driver injection vianvidia-ctk. - GPU actors cannot be suspended while holding open CUDA contexts due to gVisor's inability to serialize GPU state, as documented in
docs/api-guide.md. - Node tolerations and selectors are required to ensure
ateletand worker pods schedule exclusively on GPU-equipped nodes, with the runtime detecting devices via/dev/nvidia[0-9]*incmd/ateom-gvisor/gpu.go.
Frequently Asked Questions
Can I use GPU pools with the default musl build of ateom-gvisor?
No. GPU pools explicitly require a glibc-compiled ateom-gvisor binary because the NVIDIA driver libraries (libcuda.so) link against glibc symbols. The default musl build cannot load these proprietary NVIDIA libraries, resulting in runtime failures when actors attempt CUDA initialization. You must compile or obtain the glibc variant of the binary for GPU workloads to satisfy the requirements documented in docs/api-guide.md.
Why does my GPU WorkerPool fail validation without explicit resource requests?
The XValidation rules in pkg/api/v1alpha1/workerpool_types.go enforce that GPU pools specify both resource limits and requests for nvidia.com/gpu. Kubernetes extended resources require matching limits for proper scheduling, and the Substrate controller validates this constraint to prevent scheduling failures on nodes lacking GPU capacity. The validation explicitly checks that requests do not exceed limits while ensuring both fields are present.
What happens if I attempt to suspend a GPU actor during CUDA execution?
Suspending a GPU actor while a CUDA context remains active is unsupported and will likely cause data corruption or runtime crashes. According to the implementation in cmd/ateom-gvisor/gpu.go and documentation in docs/api-guide.md, gVisor cannot serialize GPU state or CUDA contexts. You must ensure the actor releases all GPU resources and CUDA contexts before initiating a suspend operation to maintain snapshot integrity.
How does the runtime detect available GPUs on the host node?
The ateom-gvisor runtime detects GPUs by scanning for device nodes matching the glob pattern /dev/nvidia[0-9]*, as defined by the gpuDeviceGlob constant in cmd/ateom-gvisor/gpu.go. This detection occurs during sandbox initialization to validate that the host node actually exposes NVIDIA devices before attempting passthrough, ensuring compatibility with the nvidia-ctk injection process and preventing failures on non-GPU nodes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →