# How Agent Substrate Manages WorkerPool CRDs: Controller Architecture and Lifecycle

> Learn how Agent Substrate manages WorkerPool CRDs with its controller architecture. Discover how it translates specs, syncs status, and exports metrics for Kubernetes Deployments.

- Repository: [Agent Substrate/substrate](https://github.com/agent-substrate/substrate)
- Tags: architecture
- Published: 2026-08-22

---

**Agent Substrate manages WorkerPool CRDs through a dedicated controller that translates declarative specifications into Kubernetes Deployments while maintaining status synchronization and exporting OpenTelemetry metrics.**

The **Agent Substrate** project provides a Kubernetes-native platform for orchestrating sandboxed worker pools. WorkerPool Custom Resource Definitions (CRDs) serve as the primary interface for declaring collections of `ateom` worker pods, with the controller handling complex lifecycle operations including Deployment generation, sandbox configuration, and telemetry propagation.

## WorkerPool CRD Schema and API Definition

The WorkerPool schema resides in [`pkg/api/v1alpha1/workerpool_types.go`](https://github.com/agent-substrate/substrate/blob/main/pkg/api/v1alpha1/workerpool_types.go) and defines the contract between user intent and system state.

**Spec Structure** (`WorkerPoolSpec`) captures the desired configuration:
- `replicas`: Target number of worker pods
- `ateomImage`: Container image for the `ateom` binary
- `template`: Optional pod-template overrides for resources, node selectors, and tolerations
- `sandboxClass`: Sandbox backend selection (gVisor or micro-VM)

**Status Structure** (`WorkerPoolStatus`) reports observed state:
- Active replica count and ready replicas
- Pod selector for service discovery
- Current deployment reference

The type registers with the Kubernetes API machinery through `SchemeBuilder`, enabling the API server to serve the `ate.dev/v1alpha1` version of the resource.

## The Reconciliation Loop

The core control logic lives in [`cmd/atecontroller/internal/controllers/workerpool_controller.go`](https://github.com/agent-substrate/substrate/blob/main/cmd/atecontroller/internal/controllers/workerpool_controller.go). The controller implements the standard Kubernetes controller pattern with a `Reconcile` method that executes the following sequence:

1. **Fetch**: Retrieves the WorkerPool object by namespace and name
2. **Apply**: Calls `applyDeployment` to generate or mutate the backing Deployment
3. **Sync**: Reads the resulting Deployment state to update the WorkerPool's status fields with actual replica counts and selector information

The controller uses server-side apply to manage the Deployment resource, ensuring that field ownership remains clear between the controller and other Kubernetes components.

## Deployment Construction and Pod Configuration

Heavy lifting for resource generation occurs in [`workerpool_apply.go`](https://github.com/agent-substrate/substrate/blob/main/workerpool_apply.go). This file constructs the actual Kubernetes objects that realize the WorkerPool specification.

### Building the Deployment Apply Configuration

The `buildDeploymentApplyConfig` function constructs an `appsv1ac.DeploymentApplyConfiguration` that includes:

- **Identity labels**: The `ate.dev/worker-pool` label links every pod back to its parent WorkerPool
- **Container specification**: Configures the `ateom` binary with arguments, ports, security contexts, and environment variables
- **Storage volumes**: Mounts for the ateom runtime, tunnel identity credentials, and egress trust bundles
- **User overrides**: Merges custom labels and annotations from the optional `template` field

### Sandbox-Specific Configurations

Helper functions handle specialized hardware and security requirements:

- **`maybeApplyMicroVMPodShape`**: Adds `/dev/kvm` device access and nested virtualization resources for micro-VM sandbox pools
- **`maybeApplyGPUPodShape`**: Configures NVIDIA toolkit mounts and GPU resource limits when GPU acceleration is requested
- **Environment injection**: The `ateomContainerEnv` function populates OpenTelemetry environment variables (`OTEL_EXPORTER_OTLP_ENDPOINT`, `OTEL_RESOURCE_ATTRIBUTES`) when the controller is configured with a telemetry endpoint

## OpenTelemetry Integration and Metrics

The controller exposes operational visibility through two OpenTelemetry counters registered during initialization:

- **`ate.workerpool.desired_workers`**: Reports `spec.replicas` values across all pools
- **`ate.workerpool.ready_workers`**: Reports `status.readyReplicas` for availability tracking

Both metrics include `workerpoolNamespace` and `workerpoolName` labels for dimensional filtering. The callback enumerates all WorkerPool instances in the cluster and records current values, enabling real-time capacity planning and alerting.

Telemetry configuration propagates from controller flags down to individual worker pods through the `ateomContainerEnv` helper in [`workerpool_apply.go`](https://github.com/agent-substrate/substrate/blob/main/workerpool_apply.go) (lines 91-108), ensuring each `ateom` instance can export traces and metrics to the configured collector.

## CLI Operations and Worker Inspection

Operators interact with WorkerPools through the `kubectl-ate` CLI plugin. The [`workers.go`](https://github.com/agent-substrate/substrate/blob/main/workers.go) file in `cmd/kubectl-ate/internal/cmd/` implements the `listAllWorkers` function, which:

- Invokes the `ListWorkers` RPC from the `ateapi` service
- Supports filtering by namespace, atespace, label selector, or sandbox class
- Returns live pod information from the actual Kubernetes resources created by the WorkerPool controller

This abstraction allows administrators to inspect worker health without directly querying Deployments or Pods, maintaining the CRD as the single source of truth.

## Practical Examples

Create a basic WorkerPool with three gVisor-sandboxed replicas:

```yaml
apiVersion: ate.dev/v1alpha1
kind: WorkerPool
metadata:
  name: demo-pool
spec:
  replicas: 3
  ateomImage: ghcr.io/agent-substrate/ateom:latest
  sandboxClass: gvisor
  template:
    resources:
      limits:
        cpu: "1"
        memory: "512Mi"
      requests:
        cpu: "0.5"
        memory: "256Mi"

```

Apply the configuration:

```bash
kubectl apply -f demo-pool.yaml

```

Verify status synchronization:

```bash
kubectl get workerpool demo-pool -o yaml

```

Query workers through the CLI:

```bash
kubectl ate workers --selector=ate.dev/worker-pool=demo-pool

```

Configure a micro-VM pool with node selection:

```yaml
spec:
  sandboxClass: microvm
  template:
    nodeSelector:
      ate.dev/sandboxClass: microvm
    tolerations:
    - key: ate.dev/sandboxClass
      operator: Equal
      value: microvm
      effect: NoSchedule

```

## Summary

- **WorkerPool CRDs** in Agent Substrate act as declarative interfaces for managing collections of sandboxed worker pods
- The controller in [`workerpool_controller.go`](https://github.com/agent-substrate/substrate/blob/main/workerpool_controller.go) maintains the reconciliation loop, ensuring Deployments match the desired specification
- [`workerpool_apply.go`](https://github.com/agent-substrate/substrate/blob/main/workerpool_apply.go) handles complex Deployment construction, including sandbox-specific resource allocation for micro-VMs and GPUs
- OpenTelemetry metrics (`ate.workerpool.desired_workers`, `ate.workerpool.ready_workers`) provide observability into pool capacity and health
- The `kubectl-ate` CLI abstracts pod queries through the `ateapi` service, enabling filtered worker inspection by labels or sandbox class

## Frequently Asked Questions

### What is the relationship between a WorkerPool CRD and the Kubernetes Deployment?

The WorkerPool CRD serves as a user-facing abstraction while the controller manages an underlying Kubernetes Deployment that actually creates and manages the worker pods. The controller in [`workerpool_controller.go`](https://github.com/agent-substrate/substrate/blob/main/workerpool_controller.go) generates this Deployment through server-side apply and continuously updates the WorkerPool status to reflect the Deployment's actual state, including replica counts and pod selectors.

### How does Agent Substrate handle different sandbox technologies like gVisor and micro-VMs?

The `sandboxClass` field in the WorkerPool spec determines the runtime configuration. When set to `microvm`, the controller invokes `maybeApplyMicroVMPodShape` in [`workerpool_apply.go`](https://github.com/agent-substrate/substrate/blob/main/workerpool_apply.go) to add `/dev/kvm` device access and nested virtualization resources. For GPU workloads, `maybeApplyGPUPodShape` injects NVIDIA toolkit mounts and resource limits. The default `gvisor` configuration applies standard security contexts without additional hardware resources.

### Can I customize resource limits and node placement for worker pods?

Yes, the `template` field in `WorkerPoolSpec` accepts standard Kubernetes pod template configurations. You can specify CPU and memory limits, node selectors, tolerations, and affinity rules that the controller merges into the generated Deployment via `applyWorkerPoolPodTemplate`. This allows targeting specific hardware profiles or isolating micro-VM workers to nodes with virtualization support.

### How do I monitor the health and capacity of my WorkerPools?

The controller exports two OpenTelemetry metrics: `ate.workerpool.desired_workers` (from `spec.replicas`) and `ate.workerpool.ready_workers` (from `status.readyReplicas`). These metrics include namespace and name labels for filtering. Additionally, you can inspect individual worker status using `kubectl ate workers` with label selectors, or check the WorkerPool status field directly via `kubectl get workerpool` to see current replica counts and readiness states.