Setting Up Google Cloud Monitoring and Observability: A Complete Guide
You can deploy a production-grade observability stack on Google Cloud by enabling Cloud Logging, Cloud Monitoring, Managed Prometheus, and Distributed Tracing through the reusable SKILL modules in the google/skills repository.
Setting up Google Cloud monitoring and observability involves configuring a unified pipeline that spans log ingestion, metrics collection, and distributed tracing. The google/skills repository codifies these configurations into modular SKILL files—located in paths like skills/cloud/gke-observability/SKILL.md—that provide copy-paste-ready commands and Terraform resources for GKE clusters, VPC networking, and SLO-based alerting.
Core Architecture and Components
Google Cloud’s observability stack consists of six integrated layers that feed into a single pipeline controllable via variables like enable_monitoring.
Cloud Logging (logging.googleapis.com) serves as the centralized ingestion point for system component logs, workload logs, and audit logs, offering a queryable Logging Query Language (LQL) interface. Cloud Monitoring (monitoring.googleapis.com) collects system-component, control-plane, and custom metrics, exposing them to Metrics Explorer, Grafana, and Managed Prometheus for PromQL queries without self-managed infrastructure.
Cloud Trace and Cloud Profiler provide OpenTelemetry-enabled agents for distributed tracing and continuous CPU/memory profiling. For network visibility, VPC Flow Logs, Firewall Logs, and Cloud NAT Logs capture packet-level telemetry via BigQuery and Cloud Monitoring MCPs. Finally, SLO Alerting uses PromQL-driven alert policies in Cloud Monitoring to automate burn-rate-aware breach detection, as configured in skills/cloud/google-cloud-slo-alert-configuration/SKILL.md.
Step-by-Step Implementation
Enable Required APIs
Before configuring specific services, activate the foundational APIs using gcloud:
gcloud services enable \
logging.googleapis.com \
monitoring.googleapis.com \
container.googleapis.com \
file.googleapis.com \
--quiet
Configure GKE Cluster Monitoring
Adopt the "golden-path" defaults to expose the full suite of metrics, logs, and tracing. Update your GKE cluster with comprehensive monitoring components and enable Managed Prometheus:
# Enable full suite of monitoring components
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--monitoring=SYSTEM,API_SERVER,SCHEDULER,CONTROLLER_MANAGER,STORAGE,POD,DEPLOYMENT,STATEFULSET,DAEMONSET,HPA,CADVISOR,KUBELET,DCGM \
--quiet
# Enable Managed Prometheus
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--enable-managed-prometheus \
--quiet
# Enable Dataplane V2 flow metrics
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--enable-dataplane-v2-flow-observability \
--quiet
Refer to skills/cloud/gke-observability/SKILL.md for the complete command reference and golden-path defaults.
Query Logs with Logging Query Language
Once Cloud Logging is active, use LQL to extract specific events. To fetch OOM-Killed events from Kubernetes:
resource.type="k8s_event" AND jsonPayload.reason="OOMKilling"
Execute queries via gcloud logging read to filter by namespace:
gcloud logging read 'resource.type="k8s_container" AND resource.labels.namespace_name="<NAMESPACE>"' \
--project <PROJECT_ID> \
--limit 100 \
--quiet
Optimize Costs with System-Only Monitoring
For non-production environments, reduce costs by limiting metric collection to system-only telemetry:
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--monitoring=SYSTEM \
--quiet
Cost considerations and toggle strategies are documented in the GKE Observability SKILL.
Implement SLO-Based Alerting with Terraform
Create burn-rate-aware alerts using PromQL queries. The following Terraform snippet, consistent with skills/cloud/google-cloud-slo-alert-configuration/SKILL.md, creates an API-server latency alert:
resource "google_monitoring_alert_policy" "api_latency" {
display_name = "API Server Latency P99 > 5s"
combiner = "OR"
conditions {
display_name = "apiserver_request_duration_seconds P99"
condition_prometheus_query_language {
duration = "60s"
evaluation_missing_data = "DEFAULT"
query = "apiserver_request_duration_seconds{quantile=\"0.99\"} > 5"
}
}
alert_strategy {
auto_close = "86400s"
}
# Notification channels omitted; configure as needed
}
Enable Distributed Tracing
Integrate OpenTelemetry exporters into your application code to send traces to Cloud Trace. For Go applications:
import (
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/exporters/trace/googlecloud"
)
func initTracer() {
exporter, _ := googlecloud.NewExporter()
tp := otel.NewTracerProvider(otel.WithSyncer(exporter))
otel.SetTracerProvider(tp)
}
Deployment patterns for other languages are available in the GKE Observability SKILL under the distributed tracing section.
Practical Implementation Examples
Full Monitoring Enablement: Use the gcloud container clusters update commands detailed in the implementation steps above to activate the complete metrics suite.
VPC Flow Logs Cost Estimation: Estimate log ingestion costs programmatically via the Monitoring API endpoint referenced in skills/cloud/google-cloud-networking-observability/references/vpc-flow-logs-cost-estimation.md:
curl "https://monitoring.googleapis.com/v3/projects/${PROJECT_ID}/timeSeries?filter=metric.type%3D%22networking.googleapis.com/vpc_flow/predicted_max_vpc_flow_logs_count%22&interval.startTime=${START}&interval.endTime=${END}"
Managed Grafana Dashboard Deployment: Deploy dashboards using JSON configurations:
gcloud monitoring dashboards create \
--project <PROJECT_ID> \
--config-from-file=grafana-dashboard.json
Key Source Files and References
skills/cloud/gke-observability/SKILL.md: Blueprint for logging, monitoring, Prometheus, and tracing configuration.skills/cloud/google-cloud-networking-observability/SKILL.md: Guides VPC Flow Logs, firewall logs, NAT logs, and network cost estimation.skills/cloud/google-cloud-slo-alert-configuration/SKILL.md: Wizard for PromQL-based SLO alert policies.skills/cloud/google-cloud-recipe-foundation-builder/SKILL.md: Sets up centralized logging and monitoring scopes across organizations.skills/cloud/google-cloud-slo-alert-configuration/references/service_metrics.md: Canonical list of GKE metrics for PromQL alerts.
Summary
Setting up Google Cloud monitoring and observability using the google/skills repository provides:
- One-command activation of production-grade monitoring via
gcloudcluster updates. - Native PromQL and LQL support for querying metrics and logs with documented examples.
- Cost-control mechanisms to toggle between full and system-only metric collection.
- Terraform-based SLO alerting aligned with Google SRE best practices.
- Integrated distributed tracing for microservice performance debugging.
These modules deliver a repeatable, secure, and documented foundation for observing GKE workloads and Google Cloud networking environments.
Frequently Asked Questions
What APIs must be enabled before setting up Google Cloud monitoring and observability?
You must enable logging.googleapis.com, monitoring.googleapis.com, and container.googleapis.com using gcloud services enable. The google/skills repository typically requires file.googleapis.com as well for comprehensive GKE observability.
How do I reduce monitoring costs for non-production GKE clusters?
Limit metric collection to system-only telemetry by updating your cluster with --monitoring=SYSTEM instead of the full component list. This eliminates workload-specific metrics while preserving node health visibility, significantly reducing ingestion costs.
What is the difference between Cloud Monitoring and Managed Prometheus?
Cloud Monitoring is Google Cloud’s native metrics service that collects system and custom metrics through the Cloud Monitoring API. Managed Prometheus is a Google-hosted Prometheus-compatible service that accepts native PromQL queries and stores time-series data without requiring you to operate Prometheus servers, as configured via --enable-managed-prometheus.
Where can I find the canonical PromQL queries for SLO alerts?
Reference skills/cloud/google-cloud-slo-alert-configuration/references/service_metrics.md within the repository. This file lists the standard GKE and Cloud metrics—such as apiserver_request_duration_seconds—used in burn-rate-aware alert policies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →