# Setting Up Google Cloud Monitoring and Observability: A Complete Guide

> Set up Google Cloud monitoring and observability using Cloud Logging, Cloud Monitoring, Managed Prometheus, and Distributed Tracing with SKILL modules. Deploy a production-ready stack.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-13

---

**You can deploy a production-grade observability stack on Google Cloud by enabling Cloud Logging, Cloud Monitoring, Managed Prometheus, and Distributed Tracing through the reusable SKILL modules in the `google/skills` repository.**

Setting up Google Cloud monitoring and observability involves configuring a unified pipeline that spans log ingestion, metrics collection, and distributed tracing. The `google/skills` repository codifies these configurations into modular SKILL files—located in paths like [`skills/cloud/gke-observability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-observability/SKILL.md)—that provide copy-paste-ready commands and Terraform resources for GKE clusters, VPC networking, and SLO-based alerting.

## Core Architecture and Components

Google Cloud’s observability stack consists of six integrated layers that feed into a single pipeline controllable via variables like `enable_monitoring`.

**Cloud Logging** (`logging.googleapis.com`) serves as the centralized ingestion point for system component logs, workload logs, and audit logs, offering a queryable **Logging Query Language (LQL)** interface. **Cloud Monitoring** (`monitoring.googleapis.com`) collects system-component, control-plane, and custom metrics, exposing them to Metrics Explorer, Grafana, and **Managed Prometheus** for PromQL queries without self-managed infrastructure.

**Cloud Trace and Cloud Profiler** provide OpenTelemetry-enabled agents for distributed tracing and continuous CPU/memory profiling. For network visibility, **VPC Flow Logs**, Firewall Logs, and Cloud NAT Logs capture packet-level telemetry via BigQuery and Cloud Monitoring MCPs. Finally, **SLO Alerting** uses PromQL-driven alert policies in Cloud Monitoring to automate burn-rate-aware breach detection, as configured in [`skills/cloud/google-cloud-slo-alert-configuration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/SKILL.md).

## Step-by-Step Implementation

### Enable Required APIs

Before configuring specific services, activate the foundational APIs using `gcloud`:

```bash
gcloud services enable \
  logging.googleapis.com \
  monitoring.googleapis.com \
  container.googleapis.com \
  file.googleapis.com \
  --quiet

```

### Configure GKE Cluster Monitoring

Adopt the "golden-path" defaults to expose the full suite of metrics, logs, and tracing. Update your GKE cluster with comprehensive monitoring components and enable Managed Prometheus:

```bash

# Enable full suite of monitoring components

gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --monitoring=SYSTEM,API_SERVER,SCHEDULER,CONTROLLER_MANAGER,STORAGE,POD,DEPLOYMENT,STATEFULSET,DAEMONSET,HPA,CADVISOR,KUBELET,DCGM \
  --quiet

# Enable Managed Prometheus

gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --enable-managed-prometheus \
  --quiet

# Enable Dataplane V2 flow metrics

gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --enable-dataplane-v2-flow-observability \
  --quiet

```

Refer to [`skills/cloud/gke-observability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-observability/SKILL.md) for the complete command reference and golden-path defaults.

### Query Logs with Logging Query Language

Once Cloud Logging is active, use LQL to extract specific events. To fetch OOM-Killed events from Kubernetes:

```text
resource.type="k8s_event" AND jsonPayload.reason="OOMKilling"

```

Execute queries via `gcloud logging read` to filter by namespace:

```bash
gcloud logging read 'resource.type="k8s_container" AND resource.labels.namespace_name="<NAMESPACE>"' \
  --project <PROJECT_ID> \
  --limit 100 \
  --quiet

```

### Optimize Costs with System-Only Monitoring

For non-production environments, reduce costs by limiting metric collection to system-only telemetry:

```bash
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
  --monitoring=SYSTEM \
  --quiet

```

Cost considerations and toggle strategies are documented in the GKE Observability SKILL.

### Implement SLO-Based Alerting with Terraform

Create burn-rate-aware alerts using PromQL queries. The following Terraform snippet, consistent with [`skills/cloud/google-cloud-slo-alert-configuration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/SKILL.md), creates an API-server latency alert:

```hcl
resource "google_monitoring_alert_policy" "api_latency" {
  display_name = "API Server Latency P99 > 5s"
  combiner     = "OR"

  conditions {
    display_name = "apiserver_request_duration_seconds P99"
    condition_prometheus_query_language {
      duration = "60s"
      evaluation_missing_data = "DEFAULT"
      query = "apiserver_request_duration_seconds{quantile=\"0.99\"} > 5"
    }
  }

  alert_strategy {
    auto_close = "86400s"
  }
  # Notification channels omitted; configure as needed

}

```

### Enable Distributed Tracing

Integrate OpenTelemetry exporters into your application code to send traces to Cloud Trace. For Go applications:

```go
import (
    "go.opentelemetry.io/otel"
    "go.opentelemetry.io/otel/exporters/trace/googlecloud"
)

func initTracer() {
    exporter, _ := googlecloud.NewExporter()
    tp := otel.NewTracerProvider(otel.WithSyncer(exporter))
    otel.SetTracerProvider(tp)
}

```

Deployment patterns for other languages are available in the GKE Observability SKILL under the distributed tracing section.

## Practical Implementation Examples

**Full Monitoring Enablement**: Use the `gcloud container clusters update` commands detailed in the implementation steps above to activate the complete metrics suite.

**VPC Flow Logs Cost Estimation**: Estimate log ingestion costs programmatically via the Monitoring API endpoint referenced in [`skills/cloud/google-cloud-networking-observability/references/vpc-flow-logs-cost-estimation.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-networking-observability/references/vpc-flow-logs-cost-estimation.md):

```bash
curl "https://monitoring.googleapis.com/v3/projects/${PROJECT_ID}/timeSeries?filter=metric.type%3D%22networking.googleapis.com/vpc_flow/predicted_max_vpc_flow_logs_count%22&interval.startTime=${START}&interval.endTime=${END}"

```

**Managed Grafana Dashboard Deployment**: Deploy dashboards using JSON configurations:

```bash
gcloud monitoring dashboards create \
  --project <PROJECT_ID> \
  --config-from-file=grafana-dashboard.json

```

## Key Source Files and References

- **[`skills/cloud/gke-observability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-observability/SKILL.md)**: Blueprint for logging, monitoring, Prometheus, and tracing configuration.
- **[`skills/cloud/google-cloud-networking-observability/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-networking-observability/SKILL.md)**: Guides VPC Flow Logs, firewall logs, NAT logs, and network cost estimation.
- **[`skills/cloud/google-cloud-slo-alert-configuration/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/SKILL.md)**: Wizard for PromQL-based SLO alert policies.
- **[`skills/cloud/google-cloud-recipe-foundation-builder/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-recipe-foundation-builder/SKILL.md)**: Sets up centralized logging and monitoring scopes across organizations.
- **[`skills/cloud/google-cloud-slo-alert-configuration/references/service_metrics.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/references/service_metrics.md)**: Canonical list of GKE metrics for PromQL alerts.

## Summary

Setting up Google Cloud monitoring and observability using the `google/skills` repository provides:

- **One-command activation** of production-grade monitoring via `gcloud` cluster updates.
- **Native PromQL and LQL support** for querying metrics and logs with documented examples.
- **Cost-control mechanisms** to toggle between full and system-only metric collection.
- **Terraform-based SLO alerting** aligned with Google SRE best practices.
- **Integrated distributed tracing** for microservice performance debugging.

These modules deliver a repeatable, secure, and documented foundation for observing GKE workloads and Google Cloud networking environments.

## Frequently Asked Questions

### What APIs must be enabled before setting up Google Cloud monitoring and observability?

You must enable `logging.googleapis.com`, `monitoring.googleapis.com`, and `container.googleapis.com` using `gcloud services enable`. The `google/skills` repository typically requires `file.googleapis.com` as well for comprehensive GKE observability.

### How do I reduce monitoring costs for non-production GKE clusters?

Limit metric collection to system-only telemetry by updating your cluster with `--monitoring=SYSTEM` instead of the full component list. This eliminates workload-specific metrics while preserving node health visibility, significantly reducing ingestion costs.

### What is the difference between Cloud Monitoring and Managed Prometheus?

**Cloud Monitoring** is Google Cloud’s native metrics service that collects system and custom metrics through the Cloud Monitoring API. **Managed Prometheus** is a Google-hosted Prometheus-compatible service that accepts native PromQL queries and stores time-series data without requiring you to operate Prometheus servers, as configured via `--enable-managed-prometheus`.

### Where can I find the canonical PromQL queries for SLO alerts?

Reference [`skills/cloud/google-cloud-slo-alert-configuration/references/service_metrics.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-slo-alert-configuration/references/service_metrics.md) within the repository. This file lists the standard GKE and Cloud metrics—such as `apiserver_request_duration_seconds`—used in burn-rate-aware alert policies.