How to Configure WorkerPool for GPU Passthrough in Agent Substrate

Configure GPU passthrough by defining nvidia.com/gpu in the WorkerPool resourceRequests field, ensuring underlying nodes run the NVIDIA device plugin, and optionally setting runtimeClassName: nvidia for direct CUDA access.

Agent Substrate manages ephemeral workloads through WorkerPools, custom Kubernetes resources defined in the agent-substrate/substrate repository that map short-lived actors to ready Pods. To expose physical GPU devices to workers, you configure the WorkerPool specification to request GPU resources while ensuring the node supplies the corresponding device capacity.

Prerequisites for GPU Node Configuration

Before configuring the WorkerPool, prepare your cluster nodes to expose GPU devices.

  • Install NVIDIA drivers on the host operating system to enable hardware access.
  • Deploy the NVIDIA GPU device plugin as a DaemonSet. The official Helm chart (nvidia-device-plugin) registers the nvidia.com/gpu resource with the kubelet.
  • Verify node capacity demonstrates GPU availability:
kubectl describe node <gpu-node-name> | grep nvidia.com/gpu

The output must show allocatable GPU resources before the WorkerPool can schedule GPU-bound workers.

Defining GPU Resources in the WorkerPool Spec

The WorkerPool custom resource definition (CRD) accepts a resourceRequests map that translates to container resource requirements for every worker spawned from the pool. As defined in the API types, this field allows you to request specific hardware devices like GPUs.

YAML Configuration for GPU Passthrough

Create a WorkerPool that binds to GPU-equipped nodes and requests one GPU per worker:

apiVersion: substrate.agent-substrate.github.com/v1alpha1
kind: WorkerPool
metadata:
  name: gpu-pool
  namespace: default
spec:
  nodeSelector:
    accelerator: nvidia-gpu
  resourceRequests:
    nvidia.com/gpu: "1"
  maxWorkers: 10

Key configuration details:

  • nodeSelector: Targets nodes labeled with GPU hardware (e.g., accelerator: nvidia-gpu).
  • resourceRequests: Requests the nvidia.com/gpu resource; the value represents the count of GPUs per worker.
  • maxWorkers: Limits concurrent workers to prevent GPU oversubscription (unless using MPS).

Programmatic Configuration Using the Go Client

For dynamic provisioning, use the generated typed client defined in pkg/client/clientset/versioned/typed/api/v1alpha1/workerpool.go:

import (
    "context"
    apiv1alpha1 "github.com/agent-substrate/substrate/pkg/apis/api/v1alpha1"
    clientset "github.com/agent-substrate/substrate/pkg/client/clientset/versioned"
    metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
)

func createGPUWorkerPool(cs clientset.Interface) error {
    pool := &apiv1alpha1.WorkerPool{
        ObjectMeta: metav1.ObjectMeta{
            Name:      "gpu-pool",
            Namespace: "default",
        },
        Spec: apiv1alpha1.WorkerPoolSpec{
            NodeSelector: map[string]string{
                "accelerator": "nvidia-gpu",
            },
            ResourceRequests: map[string]string{
                "nvidia.com/gpu": "1",
            },
            MaxWorkers: 10,
        },
    }
    _, err := cs.ApiV1alpha1().WorkerPools("default").Create(context.Background(), pool, metav1.CreateOptions{})
    return err
}

The client implementation handles serialization of the WorkerPool spec according to the protobuf definitions found in pkg/proto/ateapipb/ateapi.pb.go, where the WorkerPool field tracks the pool assignment for individual workers.

Advanced GPU Passthrough Options

For workloads requiring direct CUDA access or shared GPU contexts, configure additional runtime parameters.

  • NVIDIA Container Runtime: Set runtimeClassName: nvidia in the WorkerPool spec (if your cluster supports the nvidia runtime class) to ensure the container uses nvidia-container-runtime instead of the default runc.
  • Multi-Process Service (MPS): For time-sharing GPUs across multiple workers, reference the GPUSharingConfig in the NodeConfig (as hinted in the protobuf definitions) to enable MPS strategy instead of exclusive GPU assignment.
  • Network Annotations: If using SR-IOV or specific GPU network overlays, add the annotation k8s.v1.cni.cncf.io/networks: gpu-net to the WorkerPool metadata to attach specialized network interfaces alongside the GPU device.

Validation and Testing

Deploy the configuration and verify GPU allocation:

kubectl apply -f gpu-pool.yaml

Inspect running workers to confirm GPU device injection:


# List workers belonging to the pool

kubectl get workers -n default -l substrate.io/worker-pool=gpu-pool

# Verify nvidia-smi inside a worker pod

kubectl exec -it <worker-pod-name> -- nvidia-smi

Successful output from nvidia-smi inside the container confirms that the WorkerPool configuration correctly passed through the GPU device from the node to the worker runtime.

Key Source Files

Understanding these implementation files helps when debugging GPU scheduling issues or extending the API:

Summary

  • WorkerPools in Agent Substrate require the resourceRequests field to specify nvidia.com/gpu for GPU passthrough.
  • Nodes must run NVIDIA drivers and the NVIDIA device plugin to advertise GPU capacity to the cluster.
  • Use runtimeClassName: nvidia for CUDA-compatible container runtimes.
  • The generated Go client in pkg/client/clientset/versioned supports programmatic WorkerPool creation with GPU resource constraints.
  • Validate configurations by executing nvidia-smi inside running worker pods to confirm device visibility.

Frequently Asked Questions

How do I enable GPU sharing for multiple concurrent workers?

Configure the MPS (Multi-Process Service) strategy in the NodeConfig or WorkerPool annotations. MPS allows multiple CUDA processes to share a single GPU context, though you must ensure the maxWorkers value does not exceed the GPU's memory and compute capacity. Without MPS, each worker typically requires a dedicated GPU to avoid resource conflicts.

What runtime class should I specify for NVIDIA GPU containers?

Set runtimeClassName: nvidia in the WorkerPool specification if your cluster has registered the nvidia runtime class. This ensures the containerd or CRI-O runtime invokes nvidia-container-runtime, which configures the GPU device hooks and library mounts required for CUDA applications. If using standard containerd without the NVIDIA runtime class, the GPU device plugin alone handles device injection via the Device Plugin API.

How do I verify that a worker actually received the GPU device?

Execute kubectl exec -it <worker-pod-name> -- nvidia-smi to check if the NVIDIA System Management Interface detects the hardware. If the command returns GPU details (name, temperature, memory usage), the passthrough succeeded. If the command fails with "command not found" or "no devices found," verify that the node reports nvidia.com/gpu capacity and that the WorkerPool nodeSelector targets the correct GPU-equipped nodes.

Can I use AMD or Intel GPUs with Agent Substrate WorkerPools?

Yes, provided the respective device plugins (AMD GPU device plugin or Intel GPU plugin) are installed on the nodes and advertise resources (e.g., amd.com/gpu or gpu.intel.com/i915). Replace nvidia.com/gpu with the appropriate vendor resource name in the WorkerPool resourceRequests field. Ensure the container image includes the necessary vendor-specific drivers and runtime libraries.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →