Argo CD Health Checks: How Resource Health Is Determined and Aggregated
Argo CD health checks evaluate the operational state of every Kubernetes resource to determine whether an Application is Healthy, Progressing, Degraded, or Unknown, using built-in logic for standard resources and customizable Lua scripts for CRDs.
Argo CD, the declarative GitOps continuous delivery tool for Kubernetes, leverages a sophisticated health assessment system defined in the gitops-engine package to report the real-time status of deployed Applications. According to the argoproj/argo-cd source code, the system evaluates individual resources against a deterministic hierarchy of states and aggregates them to reflect the overall Application health.
Health Status Hierarchy and Precedence
Argo CD defines six distinct health statuses, ordered from best to worst. This precedence is encoded in the healthOrder slice inside gitops-engine/pkg/health/health.go and is used by the IsWorse function to determine the most severe status when aggregating multiple resources.
- Healthy – The resource is fully operational and ready.
- Suspended – The resource is paused, such as a CronJob with
spec.suspend: true. - Progressing – The resource is not yet healthy but is making progress, such as a Deployment during a rolling update.
- Missing – The resource is absent from the cluster.
- Degraded – The resource has reported a failure condition.
- Unknown – Health could not be determined, typically due to an error in a custom health script.
When assessing Application health, the controller selects the worst status among all immediate child resources. This means a single Degraded resource will mark the entire Application as Degraded, even if all other resources are Healthy.
Core Health Assessment Architecture
The health evaluation flow follows a deterministic path from resource detection to final aggregation.
Entry Point: GetResourceHealth
All health assessments originate from the GetResourceHealth function in gitops-engine/pkg/health/health.go. This function accepts an *unstructured.Unstructured object and an optional HealthOverride:
func GetResourceHealth(obj *unstructured.Unstructured, healthOverride HealthOverride) (*HealthStatus, error)
The function first checks for a deletion timestamp, returning Progressing if the resource is being deleted. If a custom health override is provided, it executes the override logic. Otherwise, it delegates to GetHealthCheckFunc to locate the appropriate built-in evaluator.
Built-in Health Lookup
GetHealthCheckFunc switches on the object's GroupVersionKind to return a type-specific health function, such as getDeploymentHealth, getIngressHealth, or getJobHealth. These mappings are maintained in gitops-engine/pkg/health/health.go and route the evaluation to dedicated implementation files.
Resource-Specific Health Implementations
Built-in health logic resides in dedicated files following the health_<resource>.go naming convention:
- Deployments (
health_deployment.go) – Verifies thatstatus.observedGenerationmatchesmetadata.generationand thatupdatedReplicasequals the desired replica count. - Services (
health_service.go) – For LoadBalancer types, checks thatstatus.loadBalancer.ingressis non-empty. - Ingress (
health_ingress.go) – Ensuresstatus.loadBalancer.ingresscontains at least one entry. - Jobs – Respects
spec.suspendto report Suspended, and evaluates completion and failure conditions. - CronJobs – Inspects the latest Job to determine if the status is Degraded (failed) or Progressing (running).
- PersistentVolumeClaims – Requires
status.phase == "Bound". - Pods – Evaluates
phase,readystatus, and container restart counts.
Aggregation to Application Health
After individual resources are evaluated, the Application controller computes the worst health among its immediate children using the IsWorse function. The resulting status is stored in the Application CR's status.health field and displayed in the UI via color-coded indicators (green for Healthy, orange for Progressing, red for Degraded, etc.).
Built-in Health Check Implementation: Deployments
The Deployment health check in gitops-engine/pkg/health/health_deployment.go demonstrates the standard evaluation pattern. It extracts nested fields from the unstructured object to compare desired and actual states:
func getDeploymentHealth(obj *unstructured.Unstructured) (*HealthStatus, error) {
status, _, _ := unstructured.NestedMap(obj.Object, "status")
desired, _, _ := unstructured.NestedInt64(status, "replicas")
updated, _, _ := unstructured.NestedInt64(status, "updatedReplicas")
if desired == updated {
return &HealthStatus{Status: HealthStatusHealthy, Message: "All replicas up‑to‑date"}, nil
}
return &HealthStatus{Status: HealthStatusProgressing, Message: "Rolling update in progress"}, nil
}
This implementation ensures that a Deployment is only considered Healthy when all replicas are updated and ready, preventing false positives during rollouts.
Custom Health Checks with Lua
For CRDs or to override built-in logic, Argo CD supports custom health checks written in Lua. These scripts are defined in the resource_customizations directory or via the argocd-cm ConfigMap under the key resource.customizations.health.<group>_<kind>.
The Lua script receives the full resource as the obj variable and must return a table with status and optional message fields. For example, a custom check for cert-manager.io/Certificate might inspect the Ready condition:
resource.customizations.health.cert-manager.io_Certificate: |
hs = {}
if obj.status ~= nil then
if obj.status.conditions ~= nil then
for i, condition in ipairs(obj.status.conditions) do
if condition.type == "Ready" and condition.status == "False" then
hs.status = "Degraded"
hs.message = condition.message
return hs
end
if condition.type == "Ready" and condition.status == "True" then
hs.status = "Healthy"
hs.message = condition.message
return hs
end
end
end
end
hs.status = "Progressing"
hs.message = "Waiting for certificate"
return hs
Custom scripts must include a corresponding health_test.yaml file and testdata/ resources for validation. Run go test -v ./util/lua/ to verify custom logic.
Disabling Health Checks for Specific Resources
You can exclude specific resources from health assessment by applying the argocd.argoproj.io/ignore-healthcheck annotation:
apiVersion: apps/v1
kind: Deployment
metadata:
name: my‑app
annotations:
argocd.argoproj.io/ignore-healthcheck: "true"
spec:
# …
When this annotation is present, Argo CD skips health evaluation for the resource, effectively treating it as Healthy for aggregation purposes.
Summary
- Argo CD health checks use a deterministic hierarchy (Healthy → Unknown) defined in
gitops-engine/pkg/health/health.goto evaluate Kubernetes resources. - The
GetResourceHealthfunction serves as the entry point, routing to built-in implementations likegetDeploymentHealthor custom Lua overrides. - Application health is determined by aggregating the worst status among all child resources using the
IsWorsecomparison function. - Custom health checks for CRDs are implemented via Lua scripts in
resource_customizationsor theargocd-cmConfigMap. - Resources can be excluded from health checks using the
argocd.argoproj.io/ignore-healthcheckannotation.
Frequently Asked Questions
How does Argo CD determine the overall Application health from individual resources?
The Application controller aggregates health statuses by selecting the worst state among all immediate child resources according to the precedence order defined in healthOrder. The IsWorse function compares the current aggregated status against each resource's status, ensuring that a single Degraded or Missing resource will mark the entire Application as unhealthy. This result is persisted in the Application CR's status.health field.
Can I override built-in health checks for standard Kubernetes resources?
Yes, you can override built-in health checks by defining a Lua script in the argocd-cm ConfigMap under resource.customizations.health.<group>_<kind> or by placing a health.lua file in the resource_customizations/<group>/<Kind>/ directory. When GetResourceHealth detects a custom health override, it executes the Lua script instead of the built-in function, allowing you to modify health evaluation logic even for native types like Deployments or Services.
How do I write a custom health check for a CRD?
Create a Lua script that receives the resource object as obj and returns a table containing status (e.g., Healthy, Progressing, Degraded) and an optional message field. Store the script in resource_customizations/<group>/<Kind>/health.lua and include a health_test.yaml file with test cases in the testdata/ directory. Run go test -v ./util/lua/ to validate the script before deployment.
What happens if a health check script returns an error?
If a custom Lua script fails to execute or returns an invalid response, GetResourceHealth catches the error and returns the Unknown status, which is the lowest priority in the health hierarchy. This ensures that health evaluation errors do not falsely report applications as healthy while still surfacing that the assessment could not be completed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →