How to Scale AxonHub Horizontally with Kubernetes Helm Charts

Scale AxonHub horizontally by configuring the axonhub.replicaCount value for manual scaling or enabling autoscaling.enabled in the Helm chart to deploy a Horizontal Pod Autoscaler that adjusts replicas based on CPU and memory metrics.

AxonHub from the looplj/axonhub repository supports horizontal scaling through its official Helm chart located in deploy/helm. Because AxonHub pods are stateless and persist data in PostgreSQL, Kubernetes can distribute traffic across multiple instances without session affinity concerns. This guide explains how to scale AxonHub horizontally with Kubernetes Helm charts using both static replica configuration and dynamic autoscaling strategies.

Understanding AxonHub's Stateless Architecture

Horizontal scaling works because AxonHub stores all persistent data in PostgreSQL rather than local storage. As implemented in the source code, AxonHub pods only require access to the database connection string (AXONHUB_DB_DSN) and do not maintain local state. This stateless design means any replica can serve traffic, allowing Kubernetes to safely distribute requests across multiple pods. The Helm chart leverages this architecture through templates in deploy/helm/templates/ that deploy stateless pod replicas fronted by a single Service.

Configuring Manual Horizontal Scaling

Setting Replica Count in values.yaml

For predictable workloads, manually configure the number of AxonHub replicas by editing deploy/helm/values.yaml. The axonhub.replicaCount value controls how many pods the Deployment creates when autoscaling is disabled.


# deploy/helm/values.yaml

axonhub:
  replicaCount: 3  # Deploy 3 AxonHub pods

  resources:
    limits:
      cpu: "2000m"
      memory: "2Gi"
    requests:
      cpu: "1000m"
      memory: "1Gi"

Deployment Template Logic

The deploy/helm/templates/deployment.yaml template conditionally renders the replicas field only when autoscaling is disabled. This prevents conflicts between manual replica management and HPA control.


# deploy/helm/templates/deployment.yaml

spec:
  {{- if not .Values.autoscaling.enabled }}
  replicas: {{ .Values.axonhub.replicaCount }}
  {{- end }}
  selector:
    matchLabels:
      {{- include "axonhub.selectorLabels" . | nindent 6 }}
      app.kubernetes.io/component: axonhub
  template:
    spec:
      containers:
        - name: {{ .Chart.Name }}
          resources:
            {{- toYaml .Values.axonhub.resources | nindent 12 }}

Enabling Horizontal Pod Autoscaler (HPA)

Autoscaling Configuration Parameters

For dynamic workloads, enable the Horizontal Pod Autoscaler by setting autoscaling.enabled: true in values.yaml. This instructs Helm to render the HPA resource defined in deploy/helm/templates/hpa.yaml.


# deploy/helm/values.yaml

autoscaling:
  enabled: true
  minReplicas: 2
  maxReplicas: 10
  targetCPUUtilizationPercentage: 70
  targetMemoryUtilizationPercentage: 80

HPA Template Implementation

The deploy/helm/templates/hpa.yaml template creates an autoscaling/v2 HorizontalPodAutoscaler resource that targets the AxonHub Deployment. It supports both CPU and memory utilization metrics, scaling the pod count between the configured minimum and maximum replicas.


# deploy/helm/templates/hpa.yaml

{{- if .Values.autoscaling.enabled }}
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: {{ include "axonhub.fullname" . }}
spec:
  minReplicas: {{ .Values.autoscaling.minReplicas }}
  maxReplicas: {{ .Values.autoscaling.maxReplicas }}
  metrics:
    {{- if .Values.autoscaling.targetCPUUtilizationPercentage }}
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: {{ .Values.autoscaling.targetCPUUtilizationPercentage }}
    {{- end }}
    {{- if .Values.autoscaling.targetMemoryUtilizationPercentage }}
    - type: Resource
      resource:
        name: memory
        target:
          type: Utilization
          averageUtilization: {{ .Values.autoscaling.targetMemoryUtilizationPercentage }}
    {{- end }}
{{- end }}

Service Discovery and Load Balancing

When scaling horizontally, Kubernetes handles traffic distribution automatically. The Helm chart creates a ClusterIP Service defined in deploy/helm/templates/service.yaml that exposes AxonHub on port 8090. This Service fronts all pod replicas, using Kubernetes' built-in load balancing to distribute requests across available AxonHub instances. If you enable Ingress via ingress.enabled: true, external traffic routes through the Ingress controller to this Service, maintaining balanced distribution across all replicas.

Resource Requirements for Effective Scaling

To ensure the Horizontal Pod Autoscaler calculates utilization percentages correctly, you must define resource requests and limits in values.yaml. The HPA compares current usage against these defined values to determine when to scale. Without proper resource specifications, the autoscaler cannot make informed scaling decisions.

Example configuration:

axonhub:
  resources:
    limits:
      cpu: "2000m"
      memory: "2Gi"
    requests:
      cpu: "1000m"
      memory: "1Gi"

Step-by-Step Deployment Workflow

Follow these steps to deploy and scale AxonHub in your Kubernetes cluster.

  1. Install with default settings – Deploy AxonHub with a single replica for initial testing:

    helm install axonhub ./deploy/helm
  2. Configure autoscaling – Create a custom scaling-values.yaml file enabling HPA with your desired thresholds:

    autoscaling:
      enabled: true
      minReplicas: 2
      maxReplicas: 10
      targetCPUUtilizationPercentage: 70
  3. Apply the configuration – Upgrade the Helm release to apply scaling settings:

    helm upgrade axonhub ./deploy/helm -f scaling-values.yaml
  4. Verify the deployment – Confirm the HPA is active and pods are distributed:

    kubectl get hpa
    kubectl get pods -l app.kubernetes.io/component=axonhub

Summary

  • AxonHub supports horizontal scaling through the official Helm chart in deploy/helm due to its stateless architecture and external PostgreSQL persistence.
  • Manual scaling uses axonhub.replicaCount in values.yaml, rendered in deploy/helm/templates/deployment.yaml when autoscaling.enabled is false.
  • Automated scaling deploys a Horizontal Pod Autoscaler via deploy/helm/templates/hpa.yaml, configured through autoscaling.minReplicas, maxReplicas, and target utilization percentages.
  • Kubernetes Services automatically load-balance traffic across AxonHub pods on port 8090.
  • Proper resource requests and limits are required for the HPA to calculate utilization metrics correctly.

Frequently Asked Questions

What makes AxonHub suitable for horizontal scaling?

AxonHub is stateless by design, storing all persistent data in PostgreSQL rather than local storage. This means each pod replica operates independently without requiring session affinity or shared state, allowing Kubernetes to distribute traffic across any number of AxonHub instances safely.

How do I disable autoscaling and set a fixed number of replicas?

Set autoscaling.enabled: false in your values.yaml file and specify the desired count using axonhub.replicaCount. The Helm template in deploy/helm/templates/deployment.yaml will render the static replicas field only when autoscaling is disabled, preventing conflicts with HPA controllers.

Why is the Horizontal Pod Autoscaler not scaling my AxonHub deployment?

The HPA requires properly defined resource requests and limits to calculate utilization percentages. If you haven't specified axonhub.resources.requests and limits in values.yaml, the autoscaler cannot determine when CPU or memory thresholds are exceeded. Additionally, ensure autoscaling.enabled is set to true and that metrics-server is running in your cluster.

Can I expose multiple AxonHub replicas externally using an Ingress?

Yes, the Helm chart supports Ingress configuration via ingress.enabled: true. When enabled, the Ingress routes external traffic to the AxonHub Service, which automatically load-balances across all pod replicas. This works seamlessly whether you have 2 replicas or 20, with the Service distributing requests on port 8090.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →