How Agent Substrate Manages WorkerPool CRDs: Controller Architecture and Lifecycle
Agent Substrate manages WorkerPool CRDs through a dedicated controller that translates declarative specifications into Kubernetes Deployments while maintaining status synchronization and exporting OpenTelemetry metrics.
The Agent Substrate project provides a Kubernetes-native platform for orchestrating sandboxed worker pools. WorkerPool Custom Resource Definitions (CRDs) serve as the primary interface for declaring collections of ateom worker pods, with the controller handling complex lifecycle operations including Deployment generation, sandbox configuration, and telemetry propagation.
WorkerPool CRD Schema and API Definition
The WorkerPool schema resides in pkg/api/v1alpha1/workerpool_types.go and defines the contract between user intent and system state.
Spec Structure (WorkerPoolSpec) captures the desired configuration:
replicas: Target number of worker podsateomImage: Container image for theateombinarytemplate: Optional pod-template overrides for resources, node selectors, and tolerationssandboxClass: Sandbox backend selection (gVisor or micro-VM)
Status Structure (WorkerPoolStatus) reports observed state:
- Active replica count and ready replicas
- Pod selector for service discovery
- Current deployment reference
The type registers with the Kubernetes API machinery through SchemeBuilder, enabling the API server to serve the ate.dev/v1alpha1 version of the resource.
The Reconciliation Loop
The core control logic lives in cmd/atecontroller/internal/controllers/workerpool_controller.go. The controller implements the standard Kubernetes controller pattern with a Reconcile method that executes the following sequence:
- Fetch: Retrieves the WorkerPool object by namespace and name
- Apply: Calls
applyDeploymentto generate or mutate the backing Deployment - Sync: Reads the resulting Deployment state to update the WorkerPool's status fields with actual replica counts and selector information
The controller uses server-side apply to manage the Deployment resource, ensuring that field ownership remains clear between the controller and other Kubernetes components.
Deployment Construction and Pod Configuration
Heavy lifting for resource generation occurs in workerpool_apply.go. This file constructs the actual Kubernetes objects that realize the WorkerPool specification.
Building the Deployment Apply Configuration
The buildDeploymentApplyConfig function constructs an appsv1ac.DeploymentApplyConfiguration that includes:
- Identity labels: The
ate.dev/worker-poollabel links every pod back to its parent WorkerPool - Container specification: Configures the
ateombinary with arguments, ports, security contexts, and environment variables - Storage volumes: Mounts for the ateom runtime, tunnel identity credentials, and egress trust bundles
- User overrides: Merges custom labels and annotations from the optional
templatefield
Sandbox-Specific Configurations
Helper functions handle specialized hardware and security requirements:
maybeApplyMicroVMPodShape: Adds/dev/kvmdevice access and nested virtualization resources for micro-VM sandbox poolsmaybeApplyGPUPodShape: Configures NVIDIA toolkit mounts and GPU resource limits when GPU acceleration is requested- Environment injection: The
ateomContainerEnvfunction populates OpenTelemetry environment variables (OTEL_EXPORTER_OTLP_ENDPOINT,OTEL_RESOURCE_ATTRIBUTES) when the controller is configured with a telemetry endpoint
OpenTelemetry Integration and Metrics
The controller exposes operational visibility through two OpenTelemetry counters registered during initialization:
ate.workerpool.desired_workers: Reportsspec.replicasvalues across all poolsate.workerpool.ready_workers: Reportsstatus.readyReplicasfor availability tracking
Both metrics include workerpoolNamespace and workerpoolName labels for dimensional filtering. The callback enumerates all WorkerPool instances in the cluster and records current values, enabling real-time capacity planning and alerting.
Telemetry configuration propagates from controller flags down to individual worker pods through the ateomContainerEnv helper in workerpool_apply.go (lines 91-108), ensuring each ateom instance can export traces and metrics to the configured collector.
CLI Operations and Worker Inspection
Operators interact with WorkerPools through the kubectl-ate CLI plugin. The workers.go file in cmd/kubectl-ate/internal/cmd/ implements the listAllWorkers function, which:
- Invokes the
ListWorkersRPC from theateapiservice - Supports filtering by namespace, atespace, label selector, or sandbox class
- Returns live pod information from the actual Kubernetes resources created by the WorkerPool controller
This abstraction allows administrators to inspect worker health without directly querying Deployments or Pods, maintaining the CRD as the single source of truth.
Practical Examples
Create a basic WorkerPool with three gVisor-sandboxed replicas:
apiVersion: ate.dev/v1alpha1
kind: WorkerPool
metadata:
name: demo-pool
spec:
replicas: 3
ateomImage: ghcr.io/agent-substrate/ateom:latest
sandboxClass: gvisor
template:
resources:
limits:
cpu: "1"
memory: "512Mi"
requests:
cpu: "0.5"
memory: "256Mi"
Apply the configuration:
kubectl apply -f demo-pool.yaml
Verify status synchronization:
kubectl get workerpool demo-pool -o yaml
Query workers through the CLI:
kubectl ate workers --selector=ate.dev/worker-pool=demo-pool
Configure a micro-VM pool with node selection:
spec:
sandboxClass: microvm
template:
nodeSelector:
ate.dev/sandboxClass: microvm
tolerations:
- key: ate.dev/sandboxClass
operator: Equal
value: microvm
effect: NoSchedule
Summary
- WorkerPool CRDs in Agent Substrate act as declarative interfaces for managing collections of sandboxed worker pods
- The controller in
workerpool_controller.gomaintains the reconciliation loop, ensuring Deployments match the desired specification workerpool_apply.gohandles complex Deployment construction, including sandbox-specific resource allocation for micro-VMs and GPUs- OpenTelemetry metrics (
ate.workerpool.desired_workers,ate.workerpool.ready_workers) provide observability into pool capacity and health - The
kubectl-ateCLI abstracts pod queries through theateapiservice, enabling filtered worker inspection by labels or sandbox class
Frequently Asked Questions
What is the relationship between a WorkerPool CRD and the Kubernetes Deployment?
The WorkerPool CRD serves as a user-facing abstraction while the controller manages an underlying Kubernetes Deployment that actually creates and manages the worker pods. The controller in workerpool_controller.go generates this Deployment through server-side apply and continuously updates the WorkerPool status to reflect the Deployment's actual state, including replica counts and pod selectors.
How does Agent Substrate handle different sandbox technologies like gVisor and micro-VMs?
The sandboxClass field in the WorkerPool spec determines the runtime configuration. When set to microvm, the controller invokes maybeApplyMicroVMPodShape in workerpool_apply.go to add /dev/kvm device access and nested virtualization resources. For GPU workloads, maybeApplyGPUPodShape injects NVIDIA toolkit mounts and resource limits. The default gvisor configuration applies standard security contexts without additional hardware resources.
Can I customize resource limits and node placement for worker pods?
Yes, the template field in WorkerPoolSpec accepts standard Kubernetes pod template configurations. You can specify CPU and memory limits, node selectors, tolerations, and affinity rules that the controller merges into the generated Deployment via applyWorkerPoolPodTemplate. This allows targeting specific hardware profiles or isolating micro-VM workers to nodes with virtualization support.
How do I monitor the health and capacity of my WorkerPools?
The controller exports two OpenTelemetry metrics: ate.workerpool.desired_workers (from spec.replicas) and ate.workerpool.ready_workers (from status.readyReplicas). These metrics include namespace and name labels for filtering. Additionally, you can inspect individual worker status using kubectl ate workers with label selectors, or check the WorkerPool status field directly via kubectl get workerpool to see current replica counts and readiness states.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →