How to Scale AxonHub Horizontally with Kubernetes Helm Charts
Scale AxonHub horizontally by configuring the axonhub.replicaCount value for manual scaling or enabling autoscaling.enabled in the Helm chart to deploy a Horizontal Pod Autoscaler that adjusts replicas based on CPU and memory metrics.
AxonHub from the looplj/axonhub repository supports horizontal scaling through its official Helm chart located in deploy/helm. Because AxonHub pods are stateless and persist data in PostgreSQL, Kubernetes can distribute traffic across multiple instances without session affinity concerns. This guide explains how to scale AxonHub horizontally with Kubernetes Helm charts using both static replica configuration and dynamic autoscaling strategies.
Understanding AxonHub's Stateless Architecture
Horizontal scaling works because AxonHub stores all persistent data in PostgreSQL rather than local storage. As implemented in the source code, AxonHub pods only require access to the database connection string (AXONHUB_DB_DSN) and do not maintain local state. This stateless design means any replica can serve traffic, allowing Kubernetes to safely distribute requests across multiple pods. The Helm chart leverages this architecture through templates in deploy/helm/templates/ that deploy stateless pod replicas fronted by a single Service.
Configuring Manual Horizontal Scaling
Setting Replica Count in values.yaml
For predictable workloads, manually configure the number of AxonHub replicas by editing deploy/helm/values.yaml. The axonhub.replicaCount value controls how many pods the Deployment creates when autoscaling is disabled.
# deploy/helm/values.yaml
axonhub:
replicaCount: 3 # Deploy 3 AxonHub pods
resources:
limits:
cpu: "2000m"
memory: "2Gi"
requests:
cpu: "1000m"
memory: "1Gi"
Deployment Template Logic
The deploy/helm/templates/deployment.yaml template conditionally renders the replicas field only when autoscaling is disabled. This prevents conflicts between manual replica management and HPA control.
# deploy/helm/templates/deployment.yaml
spec:
{{- if not .Values.autoscaling.enabled }}
replicas: {{ .Values.axonhub.replicaCount }}
{{- end }}
selector:
matchLabels:
{{- include "axonhub.selectorLabels" . | nindent 6 }}
app.kubernetes.io/component: axonhub
template:
spec:
containers:
- name: {{ .Chart.Name }}
resources:
{{- toYaml .Values.axonhub.resources | nindent 12 }}
Enabling Horizontal Pod Autoscaler (HPA)
Autoscaling Configuration Parameters
For dynamic workloads, enable the Horizontal Pod Autoscaler by setting autoscaling.enabled: true in values.yaml. This instructs Helm to render the HPA resource defined in deploy/helm/templates/hpa.yaml.
# deploy/helm/values.yaml
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 10
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: 80
HPA Template Implementation
The deploy/helm/templates/hpa.yaml template creates an autoscaling/v2 HorizontalPodAutoscaler resource that targets the AxonHub Deployment. It supports both CPU and memory utilization metrics, scaling the pod count between the configured minimum and maximum replicas.
# deploy/helm/templates/hpa.yaml
{{- if .Values.autoscaling.enabled }}
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: {{ include "axonhub.fullname" . }}
spec:
minReplicas: {{ .Values.autoscaling.minReplicas }}
maxReplicas: {{ .Values.autoscaling.maxReplicas }}
metrics:
{{- if .Values.autoscaling.targetCPUUtilizationPercentage }}
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: {{ .Values.autoscaling.targetCPUUtilizationPercentage }}
{{- end }}
{{- if .Values.autoscaling.targetMemoryUtilizationPercentage }}
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: {{ .Values.autoscaling.targetMemoryUtilizationPercentage }}
{{- end }}
{{- end }}
Service Discovery and Load Balancing
When scaling horizontally, Kubernetes handles traffic distribution automatically. The Helm chart creates a ClusterIP Service defined in deploy/helm/templates/service.yaml that exposes AxonHub on port 8090. This Service fronts all pod replicas, using Kubernetes' built-in load balancing to distribute requests across available AxonHub instances. If you enable Ingress via ingress.enabled: true, external traffic routes through the Ingress controller to this Service, maintaining balanced distribution across all replicas.
Resource Requirements for Effective Scaling
To ensure the Horizontal Pod Autoscaler calculates utilization percentages correctly, you must define resource requests and limits in values.yaml. The HPA compares current usage against these defined values to determine when to scale. Without proper resource specifications, the autoscaler cannot make informed scaling decisions.
Example configuration:
axonhub:
resources:
limits:
cpu: "2000m"
memory: "2Gi"
requests:
cpu: "1000m"
memory: "1Gi"
Step-by-Step Deployment Workflow
Follow these steps to deploy and scale AxonHub in your Kubernetes cluster.
-
Install with default settings – Deploy AxonHub with a single replica for initial testing:
helm install axonhub ./deploy/helm -
Configure autoscaling – Create a custom
scaling-values.yamlfile enabling HPA with your desired thresholds:autoscaling: enabled: true minReplicas: 2 maxReplicas: 10 targetCPUUtilizationPercentage: 70 -
Apply the configuration – Upgrade the Helm release to apply scaling settings:
helm upgrade axonhub ./deploy/helm -f scaling-values.yaml -
Verify the deployment – Confirm the HPA is active and pods are distributed:
kubectl get hpa kubectl get pods -l app.kubernetes.io/component=axonhub
Summary
- AxonHub supports horizontal scaling through the official Helm chart in
deploy/helmdue to its stateless architecture and external PostgreSQL persistence. - Manual scaling uses
axonhub.replicaCountinvalues.yaml, rendered indeploy/helm/templates/deployment.yamlwhenautoscaling.enabledis false. - Automated scaling deploys a Horizontal Pod Autoscaler via
deploy/helm/templates/hpa.yaml, configured throughautoscaling.minReplicas,maxReplicas, and target utilization percentages. - Kubernetes Services automatically load-balance traffic across AxonHub pods on port 8090.
- Proper resource requests and limits are required for the HPA to calculate utilization metrics correctly.
Frequently Asked Questions
What makes AxonHub suitable for horizontal scaling?
AxonHub is stateless by design, storing all persistent data in PostgreSQL rather than local storage. This means each pod replica operates independently without requiring session affinity or shared state, allowing Kubernetes to distribute traffic across any number of AxonHub instances safely.
How do I disable autoscaling and set a fixed number of replicas?
Set autoscaling.enabled: false in your values.yaml file and specify the desired count using axonhub.replicaCount. The Helm template in deploy/helm/templates/deployment.yaml will render the static replicas field only when autoscaling is disabled, preventing conflicts with HPA controllers.
Why is the Horizontal Pod Autoscaler not scaling my AxonHub deployment?
The HPA requires properly defined resource requests and limits to calculate utilization percentages. If you haven't specified axonhub.resources.requests and limits in values.yaml, the autoscaler cannot determine when CPU or memory thresholds are exceeded. Additionally, ensure autoscaling.enabled is set to true and that metrics-server is running in your cluster.
Can I expose multiple AxonHub replicas externally using an Ingress?
Yes, the Helm chart supports Ingress configuration via ingress.enabled: true. When enabled, the Ingress routes external traffic to the AxonHub Service, which automatically load-balances across all pod replicas. This works seamlessly whether you have 2 replicas or 20, with the Service distributing requests on port 8090.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →