How to Get GKE Cluster Autoscaler Visibility Logs for Debugging

GKE Cluster Autoscaler visibility logs are stored in Cloud Logging under the Log ID container.googleapis.com/cluster-autoscaler-visibility and can be retrieved using gcloud logging read or the log-autoscaler-events.sh script from the google/skills repository.

GKE Cluster Autoscaler visibility logs provide detailed telemetry explaining why scale-up or scale-down actions succeed or fail. These logs are essential for debugging scenarios where pods remain pending or nodes fail to terminate. According to the google/skills source code, you can access these logs through direct Cloud Logging queries or automated scripts that filter for specific error codes like scale.up.error.out.of.resources.

What Are GKE Cluster Autoscaler Visibility Logs?

GKE Cluster Autoscaler visibility logs are structured log entries emitted by the Cluster Autoscaler component that explain the decision-making process behind node scaling operations. Unlike standard GKE logs, these entries contain specific messageId values that map to granular error conditions such as stock-outs, quota exceeded limits, and configuration blockers.

As documented in skills/cloud/gke-cluster-autoscaler/SKILL.md, these logs capture both scale-up failures (e.g., GCE resource unavailability) and scale-down blockers (e.g., pods marked with safe-to-evict: "false"). The logs are essential for distinguishing between application-level issues and infrastructure constraints preventing cluster scaling.

Querying Visibility Logs with gcloud

The most direct method to retrieve GKE cluster autoscaler visibility logs uses the gcloud logging read command with a filter on the specific Log ID. This approach requires appropriate IAM permissions to read Cloud Logging entries.

To fetch the most recent 50 visibility events from the last hour:

gcloud logging read \
  'log_id("container.googleapis.com/cluster-autoscaler-visibility")' \
  --freshness=1h \
  --limit=50

For programmatic parsing, pipe the JSON output to jq to extract specific fields like messageId and message:

gcloud logging read \
  'log_id("container.googleapis.com/cluster-autoscaler-visibility")' \
  --freshness=2h \
  --limit=100 \
  --format="json" | jq '.jsonPayload.messageId, .jsonPayload.message'

Adjust the --freshness parameter (e.g., 30m, 24h) and --limit values based on your investigation window. The --format flag supports json, yaml, or table output depending on your analysis needs.

Live-Tailing Events with the Helper Script

The google/skills repository provides a dedicated shell script at skills/cloud/gke-cluster-autoscaler/assets/log-autoscaler-events.sh that automates the filtering and formatting of visibility logs. This script continuously tails events for a specific cluster, inserting the correct log filter and formatting output for readability.

To stream live visibility events for your cluster:

./assets/log-autoscaler-events.sh <CLUSTER_NAME>

Replace <CLUSTER_NAME> with the actual name of your GKE cluster. The script automatically constructs the proper log_id filter and applies formatting that highlights critical messageId values, making it easier to spot errors like scale.up.error.quota.exceeded or scale.up.no.scale.up in real-time.

Interpreting Common Error Codes

Visibility logs use standardized messageId values to categorize autoscaler behavior. According to skills/cloud/gke-cluster-autoscaler/references/ca-debug.md and skills/cloud/gke-compute-classes/references/compute-class-debug.md, these identifiers reveal the root cause of scaling issues:

  • scale.up.error.out.of.resources — Indicates a GCE stockout in the target zone. The autoscaler cannot provision nodes because the specified machine family is unavailable. Remediation involves adding fallback zones or alternative machine types.
  • scale.up.error.quota.exceeded — Signals that the project has hit CPU, IP address, or disk quota limits. Requires quota increase or resource cleanup.
  • scale.up.no.scale.up — Occurs when no node group matches the pending pod's resource requests or constraints. Verify pod resource specifications and node pool configuration.
  • Scale-down blockers — Entries referencing local storage or PDB violations explain why specific nodes are not being removed despite low utilization.

Cross-reference these IDs with the message payload details to determine whether issues stem from transient capacity limitations, configuration errors, or policy constraints.

Summary

  • GKE Cluster Autoscaler visibility logs are accessible via the Log ID container.googleapis.com/cluster-autoscaler-visibility in Cloud Logging.
  • Use gcloud logging read with --freshness and --limit parameters for targeted historical queries.
  • Leverage the log-autoscaler-events.sh script from skills/cloud/gke-cluster-autoscaler/assets/ for continuous live-tailing of cluster events.
  • Analyze the messageId field to identify specific error conditions like stock-outs (scale.up.error.out.of.resources) or quota limits (scale.up.error.quota.exceeded).
  • Reference skills/cloud/gke-cluster-autoscaler/SKILL.md and the debug reference files for complete error code documentation.

Frequently Asked Questions

What Log ID stores GKE Cluster Autoscaler visibility events?

GKE Cluster Autoscaler visibility events are stored under the Log ID container.googleapis.com/cluster-autoscaler-visibility in Cloud Logging. This specific ID filters the structured logs that contain detailed autoscaler decision telemetry, distinct from standard container logs or GKE audit logs.

How do I filter visibility logs for a specific time range?

Use the --freshness flag with gcloud logging read to define the time window (e.g., --freshness=30m for the last 30 minutes or --freshness=24h for the last day). Combine this with --limit to control the number of returned entries and avoid overwhelming output during peak scaling activity.

What does the scale.up.error.out.of.resources error indicate?

The scale.up.error.out.of.resources messageId indicates that the Cluster Autoscaler cannot provision new nodes due to a GCE stockout in the specified zone. This means the requested machine family or accelerator is temporarily unavailable in that location. The recommended remediation is to configure additional node pools in alternative zones or specify fallback machine families.

Where can I find the log-autoscaler-events.sh script?

The log-autoscaler-events.sh script is located at skills/cloud/gke-cluster-autoscaler/assets/log-autoscaler-events.sh in the google/skills repository. This helper utility wraps the gcloud logging read command with appropriate filters for the visibility Log ID and formats output for easier debugging of specific GKE clusters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →