Setting Up Google Cloud Monitoring and Observability: A Complete Guide

You can deploy a production-grade observability stack on Google Cloud by enabling Cloud Logging, Cloud Monitoring, Managed Prometheus, and Distributed Tracing through the reusable SKILL modules in the google/skills repository.

Setting up Google Cloud monitoring and observability involves configuring a unified pipeline that spans log ingestion, metrics collection, and distributed tracing. The google/skills repository codifies these configurations into modular SKILL files—located in paths like skills/cloud/gke-observability/SKILL.md—that provide copy-paste-ready commands and Terraform resources for GKE clusters, VPC networking, and SLO-based alerting.

Core Architecture and Components

Google Cloud’s observability stack consists of six integrated layers that feed into a single pipeline controllable via variables like enable_monitoring.

Cloud Logging (logging.googleapis.com) serves as the centralized ingestion point for system component logs, workload logs, and audit logs, offering a queryable Logging Query Language (LQL) interface. Cloud Monitoring (monitoring.googleapis.com) collects system-component, control-plane, and custom metrics, exposing them to Metrics Explorer, Grafana, and Managed Prometheus for PromQL queries without self-managed infrastructure.

Cloud Trace and Cloud Profiler provide OpenTelemetry-enabled agents for distributed tracing and continuous CPU/memory profiling. For network visibility, VPC Flow Logs, Firewall Logs, and Cloud NAT Logs capture packet-level telemetry via BigQuery and Cloud Monitoring MCPs. Finally, SLO Alerting uses PromQL-driven alert policies in Cloud Monitoring to automate burn-rate-aware breach detection, as configured in skills/cloud/google-cloud-slo-alert-configuration/SKILL.md.

Step-by-Step Implementation

Enable Required APIs

Before configuring specific services, activate the foundational APIs using gcloud:

gcloud services enable \
  logging.googleapis.com \
  monitoring.googleapis.com \
  container.googleapis.com \
  file.googleapis.com \
  --quiet

Configure GKE Cluster Monitoring

Adopt the "golden-path" defaults to expose the full suite of metrics, logs, and tracing. Update your GKE cluster with comprehensive monitoring components and enable Managed Prometheus:


# Enable full suite of monitoring components

gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --monitoring=SYSTEM,API_SERVER,SCHEDULER,CONTROLLER_MANAGER,STORAGE,POD,DEPLOYMENT,STATEFULSET,DAEMONSET,HPA,CADVISOR,KUBELET,DCGM \
  --quiet

# Enable Managed Prometheus

gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --enable-managed-prometheus \
  --quiet

# Enable Dataplane V2 flow metrics

gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --enable-dataplane-v2-flow-observability \
  --quiet

Refer to skills/cloud/gke-observability/SKILL.md for the complete command reference and golden-path defaults.

Query Logs with Logging Query Language

Once Cloud Logging is active, use LQL to extract specific events. To fetch OOM-Killed events from Kubernetes:

resource.type="k8s_event" AND jsonPayload.reason="OOMKilling"

Execute queries via gcloud logging read to filter by namespace:

gcloud logging read 'resource.type="k8s_container" AND resource.labels.namespace_name="<NAMESPACE>"' \
  --project <PROJECT_ID> \
  --limit 100 \
  --quiet

Optimize Costs with System-Only Monitoring

For non-production environments, reduce costs by limiting metric collection to system-only telemetry:

gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --monitoring=SYSTEM \
  --quiet

Cost considerations and toggle strategies are documented in the GKE Observability SKILL.

Implement SLO-Based Alerting with Terraform

Create burn-rate-aware alerts using PromQL queries. The following Terraform snippet, consistent with skills/cloud/google-cloud-slo-alert-configuration/SKILL.md, creates an API-server latency alert:

resource "google_monitoring_alert_policy" "api_latency" {
  display_name = "API Server Latency P99 > 5s"
  combiner     = "OR"

  conditions {
    display_name = "apiserver_request_duration_seconds P99"
    condition_prometheus_query_language {
      duration = "60s"
      evaluation_missing_data = "DEFAULT"
      query = "apiserver_request_duration_seconds{quantile=\"0.99\"} > 5"
    }
  }

  alert_strategy {
    auto_close = "86400s"
  }
  # Notification channels omitted; configure as needed

}

Enable Distributed Tracing

Integrate OpenTelemetry exporters into your application code to send traces to Cloud Trace. For Go applications:

import (
    "go.opentelemetry.io/otel"
    "go.opentelemetry.io/otel/exporters/trace/googlecloud"
)

func initTracer() {
    exporter, _ := googlecloud.NewExporter()
    tp := otel.NewTracerProvider(otel.WithSyncer(exporter))
    otel.SetTracerProvider(tp)
}

Deployment patterns for other languages are available in the GKE Observability SKILL under the distributed tracing section.

Practical Implementation Examples

Full Monitoring Enablement: Use the gcloud container clusters update commands detailed in the implementation steps above to activate the complete metrics suite.

VPC Flow Logs Cost Estimation: Estimate log ingestion costs programmatically via the Monitoring API endpoint referenced in skills/cloud/google-cloud-networking-observability/references/vpc-flow-logs-cost-estimation.md:

curl "https://monitoring.googleapis.com/v3/projects/${PROJECT_ID}/timeSeries?filter=metric.type%3D%22networking.googleapis.com/vpc_flow/predicted_max_vpc_flow_logs_count%22&interval.startTime=${START}&interval.endTime=${END}"

Managed Grafana Dashboard Deployment: Deploy dashboards using JSON configurations:

gcloud monitoring dashboards create \
  --project <PROJECT_ID> \
  --config-from-file=grafana-dashboard.json

Key Source Files and References

Summary

Setting up Google Cloud monitoring and observability using the google/skills repository provides:

  • One-command activation of production-grade monitoring via gcloud cluster updates.
  • Native PromQL and LQL support for querying metrics and logs with documented examples.
  • Cost-control mechanisms to toggle between full and system-only metric collection.
  • Terraform-based SLO alerting aligned with Google SRE best practices.
  • Integrated distributed tracing for microservice performance debugging.

These modules deliver a repeatable, secure, and documented foundation for observing GKE workloads and Google Cloud networking environments.

Frequently Asked Questions

What APIs must be enabled before setting up Google Cloud monitoring and observability?

You must enable logging.googleapis.com, monitoring.googleapis.com, and container.googleapis.com using gcloud services enable. The google/skills repository typically requires file.googleapis.com as well for comprehensive GKE observability.

How do I reduce monitoring costs for non-production GKE clusters?

Limit metric collection to system-only telemetry by updating your cluster with --monitoring=SYSTEM instead of the full component list. This eliminates workload-specific metrics while preserving node health visibility, significantly reducing ingestion costs.

What is the difference between Cloud Monitoring and Managed Prometheus?

Cloud Monitoring is Google Cloud’s native metrics service that collects system and custom metrics through the Cloud Monitoring API. Managed Prometheus is a Google-hosted Prometheus-compatible service that accepts native PromQL queries and stores time-series data without requiring you to operate Prometheus servers, as configured via --enable-managed-prometheus.

Where can I find the canonical PromQL queries for SLO alerts?

Reference skills/cloud/google-cloud-slo-alert-configuration/references/service_metrics.md within the repository. This file lists the standard GKE and Cloud metrics—such as apiserver_request_duration_seconds—used in burn-rate-aware alert policies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →