How Agent Substrate Manages WorkerPool CRDs: Controller Architecture and Lifecycle

Agent Substrate manages WorkerPool CRDs through a dedicated controller that translates declarative specifications into Kubernetes Deployments while maintaining status synchronization and exporting OpenTelemetry metrics.

The Agent Substrate project provides a Kubernetes-native platform for orchestrating sandboxed worker pools. WorkerPool Custom Resource Definitions (CRDs) serve as the primary interface for declaring collections of ateom worker pods, with the controller handling complex lifecycle operations including Deployment generation, sandbox configuration, and telemetry propagation.

WorkerPool CRD Schema and API Definition

The WorkerPool schema resides in pkg/api/v1alpha1/workerpool_types.go and defines the contract between user intent and system state.

Spec Structure (WorkerPoolSpec) captures the desired configuration:

  • replicas: Target number of worker pods
  • ateomImage: Container image for the ateom binary
  • template: Optional pod-template overrides for resources, node selectors, and tolerations
  • sandboxClass: Sandbox backend selection (gVisor or micro-VM)

Status Structure (WorkerPoolStatus) reports observed state:

  • Active replica count and ready replicas
  • Pod selector for service discovery
  • Current deployment reference

The type registers with the Kubernetes API machinery through SchemeBuilder, enabling the API server to serve the ate.dev/v1alpha1 version of the resource.

The Reconciliation Loop

The core control logic lives in cmd/atecontroller/internal/controllers/workerpool_controller.go. The controller implements the standard Kubernetes controller pattern with a Reconcile method that executes the following sequence:

  1. Fetch: Retrieves the WorkerPool object by namespace and name
  2. Apply: Calls applyDeployment to generate or mutate the backing Deployment
  3. Sync: Reads the resulting Deployment state to update the WorkerPool's status fields with actual replica counts and selector information

The controller uses server-side apply to manage the Deployment resource, ensuring that field ownership remains clear between the controller and other Kubernetes components.

Deployment Construction and Pod Configuration

Heavy lifting for resource generation occurs in workerpool_apply.go. This file constructs the actual Kubernetes objects that realize the WorkerPool specification.

Building the Deployment Apply Configuration

The buildDeploymentApplyConfig function constructs an appsv1ac.DeploymentApplyConfiguration that includes:

  • Identity labels: The ate.dev/worker-pool label links every pod back to its parent WorkerPool
  • Container specification: Configures the ateom binary with arguments, ports, security contexts, and environment variables
  • Storage volumes: Mounts for the ateom runtime, tunnel identity credentials, and egress trust bundles
  • User overrides: Merges custom labels and annotations from the optional template field

Sandbox-Specific Configurations

Helper functions handle specialized hardware and security requirements:

  • maybeApplyMicroVMPodShape: Adds /dev/kvm device access and nested virtualization resources for micro-VM sandbox pools
  • maybeApplyGPUPodShape: Configures NVIDIA toolkit mounts and GPU resource limits when GPU acceleration is requested
  • Environment injection: The ateomContainerEnv function populates OpenTelemetry environment variables (OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_RESOURCE_ATTRIBUTES) when the controller is configured with a telemetry endpoint

OpenTelemetry Integration and Metrics

The controller exposes operational visibility through two OpenTelemetry counters registered during initialization:

  • ate.workerpool.desired_workers: Reports spec.replicas values across all pools
  • ate.workerpool.ready_workers: Reports status.readyReplicas for availability tracking

Both metrics include workerpoolNamespace and workerpoolName labels for dimensional filtering. The callback enumerates all WorkerPool instances in the cluster and records current values, enabling real-time capacity planning and alerting.

Telemetry configuration propagates from controller flags down to individual worker pods through the ateomContainerEnv helper in workerpool_apply.go (lines 91-108), ensuring each ateom instance can export traces and metrics to the configured collector.

CLI Operations and Worker Inspection

Operators interact with WorkerPools through the kubectl-ate CLI plugin. The workers.go file in cmd/kubectl-ate/internal/cmd/ implements the listAllWorkers function, which:

  • Invokes the ListWorkers RPC from the ateapi service
  • Supports filtering by namespace, atespace, label selector, or sandbox class
  • Returns live pod information from the actual Kubernetes resources created by the WorkerPool controller

This abstraction allows administrators to inspect worker health without directly querying Deployments or Pods, maintaining the CRD as the single source of truth.

Practical Examples

Create a basic WorkerPool with three gVisor-sandboxed replicas:

apiVersion: ate.dev/v1alpha1
kind: WorkerPool
metadata:
  name: demo-pool
spec:
  replicas: 3
  ateomImage: ghcr.io/agent-substrate/ateom:latest
  sandboxClass: gvisor
  template:
    resources:
      limits:
        cpu: "1"
        memory: "512Mi"
      requests:
        cpu: "0.5"
        memory: "256Mi"

Apply the configuration:

kubectl apply -f demo-pool.yaml

Verify status synchronization:

kubectl get workerpool demo-pool -o yaml

Query workers through the CLI:

kubectl ate workers --selector=ate.dev/worker-pool=demo-pool

Configure a micro-VM pool with node selection:

spec:
  sandboxClass: microvm
  template:
    nodeSelector:
      ate.dev/sandboxClass: microvm
    tolerations:
    - key: ate.dev/sandboxClass
      operator: Equal
      value: microvm
      effect: NoSchedule

Summary

  • WorkerPool CRDs in Agent Substrate act as declarative interfaces for managing collections of sandboxed worker pods
  • The controller in workerpool_controller.go maintains the reconciliation loop, ensuring Deployments match the desired specification
  • workerpool_apply.go handles complex Deployment construction, including sandbox-specific resource allocation for micro-VMs and GPUs
  • OpenTelemetry metrics (ate.workerpool.desired_workers, ate.workerpool.ready_workers) provide observability into pool capacity and health
  • The kubectl-ate CLI abstracts pod queries through the ateapi service, enabling filtered worker inspection by labels or sandbox class

Frequently Asked Questions

What is the relationship between a WorkerPool CRD and the Kubernetes Deployment?

The WorkerPool CRD serves as a user-facing abstraction while the controller manages an underlying Kubernetes Deployment that actually creates and manages the worker pods. The controller in workerpool_controller.go generates this Deployment through server-side apply and continuously updates the WorkerPool status to reflect the Deployment's actual state, including replica counts and pod selectors.

How does Agent Substrate handle different sandbox technologies like gVisor and micro-VMs?

The sandboxClass field in the WorkerPool spec determines the runtime configuration. When set to microvm, the controller invokes maybeApplyMicroVMPodShape in workerpool_apply.go to add /dev/kvm device access and nested virtualization resources. For GPU workloads, maybeApplyGPUPodShape injects NVIDIA toolkit mounts and resource limits. The default gvisor configuration applies standard security contexts without additional hardware resources.

Can I customize resource limits and node placement for worker pods?

Yes, the template field in WorkerPoolSpec accepts standard Kubernetes pod template configurations. You can specify CPU and memory limits, node selectors, tolerations, and affinity rules that the controller merges into the generated Deployment via applyWorkerPoolPodTemplate. This allows targeting specific hardware profiles or isolating micro-VM workers to nodes with virtualization support.

How do I monitor the health and capacity of my WorkerPools?

The controller exports two OpenTelemetry metrics: ate.workerpool.desired_workers (from spec.replicas) and ate.workerpool.ready_workers (from status.readyReplicas). These metrics include namespace and name labels for filtering. Additionally, you can inspect individual worker status using kubectl ate workers with label selectors, or check the WorkerPool status field directly via kubectl get workerpool to see current replica counts and readiness states.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →