# How to Scale AxonHub Horizontally with Kubernetes Helm Charts

> Scale AxonHub horizontally using Kubernetes Helm charts and the looplj/axonhub repository. Configure replica counts or enable autoscaling for efficient resource management.

- Repository: [Loop/axonhub](https://github.com/looplj/axonhub)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Scale AxonHub horizontally by configuring the `axonhub.replicaCount` value for manual scaling or enabling `autoscaling.enabled` in the Helm chart to deploy a Horizontal Pod Autoscaler that adjusts replicas based on CPU and memory metrics.**

AxonHub from the `looplj/axonhub` repository supports horizontal scaling through its official Helm chart located in `deploy/helm`. Because AxonHub pods are stateless and persist data in PostgreSQL, Kubernetes can distribute traffic across multiple instances without session affinity concerns. This guide explains how to scale AxonHub horizontally with Kubernetes Helm charts using both static replica configuration and dynamic autoscaling strategies.

## Understanding AxonHub's Stateless Architecture

Horizontal scaling works because AxonHub stores all persistent data in PostgreSQL rather than local storage. As implemented in the source code, AxonHub pods only require access to the database connection string (`AXONHUB_DB_DSN`) and do not maintain local state. This stateless design means any replica can serve traffic, allowing Kubernetes to safely distribute requests across multiple pods. The Helm chart leverages this architecture through templates in `deploy/helm/templates/` that deploy stateless pod replicas fronted by a single Service.

## Configuring Manual Horizontal Scaling

### Setting Replica Count in values.yaml

For predictable workloads, manually configure the number of AxonHub replicas by editing [`deploy/helm/values.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/values.yaml). The `axonhub.replicaCount` value controls how many pods the Deployment creates when autoscaling is disabled.

```yaml

# deploy/helm/values.yaml

axonhub:
  replicaCount: 3  # Deploy 3 AxonHub pods

  resources:
    limits:
      cpu: "2000m"
      memory: "2Gi"
    requests:
      cpu: "1000m"
      memory: "1Gi"

```

### Deployment Template Logic

The [`deploy/helm/templates/deployment.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/templates/deployment.yaml) template conditionally renders the `replicas` field only when autoscaling is disabled. This prevents conflicts between manual replica management and HPA control.

```yaml

# deploy/helm/templates/deployment.yaml

spec:
  {{- if not .Values.autoscaling.enabled }}
  replicas: {{ .Values.axonhub.replicaCount }}
  {{- end }}
  selector:
    matchLabels:
      {{- include "axonhub.selectorLabels" . | nindent 6 }}
      app.kubernetes.io/component: axonhub
  template:
    spec:
      containers:
        - name: {{ .Chart.Name }}
          resources:
            {{- toYaml .Values.axonhub.resources | nindent 12 }}

```

## Enabling Horizontal Pod Autoscaler (HPA)

### Autoscaling Configuration Parameters

For dynamic workloads, enable the Horizontal Pod Autoscaler by setting `autoscaling.enabled: true` in [`values.yaml`](https://github.com/looplj/axonhub/blob/main/values.yaml). This instructs Helm to render the HPA resource defined in [`deploy/helm/templates/hpa.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/templates/hpa.yaml).

```yaml

# deploy/helm/values.yaml

autoscaling:
  enabled: true
  minReplicas: 2
  maxReplicas: 10
  targetCPUUtilizationPercentage: 70
  targetMemoryUtilizationPercentage: 80

```

### HPA Template Implementation

The [`deploy/helm/templates/hpa.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/templates/hpa.yaml) template creates an `autoscaling/v2` HorizontalPodAutoscaler resource that targets the AxonHub Deployment. It supports both CPU and memory utilization metrics, scaling the pod count between the configured minimum and maximum replicas.

```yaml

# deploy/helm/templates/hpa.yaml

{{- if .Values.autoscaling.enabled }}
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: {{ include "axonhub.fullname" . }}
spec:
  minReplicas: {{ .Values.autoscaling.minReplicas }}
  maxReplicas: {{ .Values.autoscaling.maxReplicas }}
  metrics:
    {{- if .Values.autoscaling.targetCPUUtilizationPercentage }}
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: {{ .Values.autoscaling.targetCPUUtilizationPercentage }}
    {{- end }}
    {{- if .Values.autoscaling.targetMemoryUtilizationPercentage }}
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: {{ .Values.autoscaling.targetMemoryUtilizationPercentage }}
    {{- end }}
{{- end }}

```

## Service Discovery and Load Balancing

When scaling horizontally, Kubernetes handles traffic distribution automatically. The Helm chart creates a `ClusterIP` Service defined in [`deploy/helm/templates/service.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/templates/service.yaml) that exposes AxonHub on port `8090`. This Service fronts all pod replicas, using Kubernetes' built-in load balancing to distribute requests across available AxonHub instances. If you enable Ingress via `ingress.enabled: true`, external traffic routes through the Ingress controller to this Service, maintaining balanced distribution across all replicas.

## Resource Requirements for Effective Scaling

To ensure the Horizontal Pod Autoscaler calculates utilization percentages correctly, you must define resource requests and limits in [`values.yaml`](https://github.com/looplj/axonhub/blob/main/values.yaml). The HPA compares current usage against these defined values to determine when to scale. Without proper resource specifications, the autoscaler cannot make informed scaling decisions.

Example configuration:

```yaml
axonhub:
  resources:
    limits:
      cpu: "2000m"
      memory: "2Gi"
    requests:
      cpu: "1000m"
      memory: "1Gi"

```

## Step-by-Step Deployment Workflow

Follow these steps to deploy and scale AxonHub in your Kubernetes cluster.

1. **Install with default settings** – Deploy AxonHub with a single replica for initial testing:
   ```bash
   helm install axonhub ./deploy/helm
   ```

2. **Configure autoscaling** – Create a custom [`scaling-values.yaml`](https://github.com/looplj/axonhub/blob/main/scaling-values.yaml) file enabling HPA with your desired thresholds:
   ```yaml
   autoscaling:
     enabled: true
     minReplicas: 2
     maxReplicas: 10
     targetCPUUtilizationPercentage: 70
   ```

3. **Apply the configuration** – Upgrade the Helm release to apply scaling settings:
   ```bash
   helm upgrade axonhub ./deploy/helm -f scaling-values.yaml
   ```

4. **Verify the deployment** – Confirm the HPA is active and pods are distributed:
   ```bash
   kubectl get hpa
   kubectl get pods -l app.kubernetes.io/component=axonhub
   ```

## Summary

- AxonHub supports horizontal scaling through the official Helm chart in `deploy/helm` due to its stateless architecture and external PostgreSQL persistence.
- Manual scaling uses `axonhub.replicaCount` in [`values.yaml`](https://github.com/looplj/axonhub/blob/main/values.yaml), rendered in [`deploy/helm/templates/deployment.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/templates/deployment.yaml) when `autoscaling.enabled` is false.
- Automated scaling deploys a Horizontal Pod Autoscaler via [`deploy/helm/templates/hpa.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/templates/hpa.yaml), configured through `autoscaling.minReplicas`, `maxReplicas`, and target utilization percentages.
- Kubernetes Services automatically load-balance traffic across AxonHub pods on port 8090.
- Proper resource requests and limits are required for the HPA to calculate utilization metrics correctly.

## Frequently Asked Questions

### What makes AxonHub suitable for horizontal scaling?

AxonHub is stateless by design, storing all persistent data in PostgreSQL rather than local storage. This means each pod replica operates independently without requiring session affinity or shared state, allowing Kubernetes to distribute traffic across any number of AxonHub instances safely.

### How do I disable autoscaling and set a fixed number of replicas?

Set `autoscaling.enabled: false` in your [`values.yaml`](https://github.com/looplj/axonhub/blob/main/values.yaml) file and specify the desired count using `axonhub.replicaCount`. The Helm template in [`deploy/helm/templates/deployment.yaml`](https://github.com/looplj/axonhub/blob/main/deploy/helm/templates/deployment.yaml) will render the static `replicas` field only when autoscaling is disabled, preventing conflicts with HPA controllers.

### Why is the Horizontal Pod Autoscaler not scaling my AxonHub deployment?

The HPA requires properly defined resource requests and limits to calculate utilization percentages. If you haven't specified `axonhub.resources.requests` and `limits` in [`values.yaml`](https://github.com/looplj/axonhub/blob/main/values.yaml), the autoscaler cannot determine when CPU or memory thresholds are exceeded. Additionally, ensure `autoscaling.enabled` is set to `true` and that metrics-server is running in your cluster.

### Can I expose multiple AxonHub replicas externally using an Ingress?

Yes, the Helm chart supports Ingress configuration via `ingress.enabled: true`. When enabled, the Ingress routes external traffic to the AxonHub Service, which automatically load-balances across all pod replicas. This works seamlessly whether you have 2 replicas or 20, with the Service distributing requests on port 8090.