# Strategies for Google Cloud Cost Optimization: A Practical Guide Based on the google/skills Repository

> Optimize Google Cloud costs with practical strategies from the google/skills repository. Discover CUDs, Spot VMs, GKE cost allocation, and rightsizing.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-08-13

---

**The `google/skills` repository provides production-ready templates and best practices for Google Cloud cost optimization, including Committed Use Discounts, Spot VM deployment, GKE cost allocation, and automated rightsizing through Vertical Pod Autoscaling.**

Google Cloud cost optimization follows the **Cost Optimization pillar** of the Google Cloud Well-Architected Framework. The open-source `google/skills` repository encodes these principles into reusable skills—structured documentation with copy-paste configurations—for immediate implementation. This guide distills the repository's architectural patterns into an actionable playbook that covers compute, storage, networking, and governance.

## Core Cost Optimization Principles

According to [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md), all cost optimization efforts rest on four foundational principles:

- **Align spending with business value** — Tie every resource to a measurable outcome before provisioning.
- **Foster a culture of cost awareness** — Embed FinOps practices into daily workflows and make cost data visible to all teams.
- **Optimize resource usage** — Right-size instances, select the most cost-effective machine types, and eliminate idle resources.
- **Continuous optimization** — Establish regular review cycles using automated recommendations and billing analysis.

These principles appear in the repository's operational instructions and validation checklists, providing a governance framework rather than ad-hoc fixes.

## Visibility and Governance Foundations

You cannot optimize what you cannot measure. The repository emphasizes establishing granular cost attribution before implementing technical optimizations.

Enable **BigQuery billing export** to analyze spend programmatically. For GKE environments, activate cost allocation to attribute cluster costs to namespaces and workloads:

```bash
gcloud container clusters update my-gke-cluster \
  --enable-cost-allocation \
  --region us-central1

```

*Source:* [`skills/cloud/gke-cost-analysis/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-analysis/SKILL.md), lines 85‑88.

Apply consistent **labels** (e.g., `env`, `team`, `app`) across all resources to enable chargeback and showback mechanisms. Use **Cloud Billing reports** and **Looker Studio** dashboards for baseline visibility, but rely on BigQuery for granular analysis as shown in [`skills/cloud/gke-cost-analysis/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-analysis/SKILL.md), lines 75‑77.

## Compute Optimization Strategies

### Committed Use Discounts (CUDs)

For steady-state workloads, commit to **Committed Use Discounts (CUDs)** to receive significant price reductions. Supplement with **Sustained Use Discounts (SUDs)** for predictable resource patterns. Use `gcloud compute commitments create` to purchase commitments programmatically, as referenced in [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md).

### Spot VMs for Fault-Tolerant Workloads

Deploy fault-tolerant batch jobs on **Spot VMs** to achieve 60–90 % cost reduction. In GKE, use node selectors to target Spot capacity:

```yaml
nodeSelector:
  cloud.google.com/gke-provisioning: Spot

```

*Source:* [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md).

### GKE Autopilot and Rightsizing

**GKE Autopilot** charges per pod request rather than provisioned node capacity. Avoid over-requesting CPU and memory by setting requests close to observed P95 usage. For existing Standard clusters, implement **Vertical Pod Autoscaling (VPA)** in recommendation mode to analyze actual utilization before applying changes:

```yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: myapp-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: myapp
  updatePolicy:
    updateMode: "Off"

```

*Source:* [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md), lines 74‑90.

### Resource Quotas for Cost Control

Prevent runaway spending by enforcing **Resource Quotas** at the namespace level. This caps total CPU and memory consumption regardless of individual pod specifications:

```yaml
apiVersion: v1
kind: ResourceQuota
metadata:
  name: compute-quota
  namespace: analytics
spec:
  hard:
    requests.cpu: "4"
    requests.memory: 16Gi
    limits.cpu: "8"
    limits.memory: 32Gi

```

*Source:* [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md), lines 53‑66.

## Storage Cost Management

Storage costs accumulate through unnecessary replication and suboptimal lifecycle management. The repository recommends three specific techniques from [`skills/cloud/google-cloud-storage-basics/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-storage-basics/SKILL.md) and [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md):

- **Lifecycle Policies** — Automatically transition objects to cheaper storage classes (Nearline, Coldline, Archive) based on age or access frequency. Review usage patterns with **Storage Insights** before applying policies (lines 102‑105).
- **Autoclass** — Enable `Autoclass` to let Cloud Storage automatically transition objects between Standard and colder classes based on actual access patterns, eliminating manual policy management.
- **Versioning Control** — Disable Object Versioning unless required for compliance to reduce storage and deletion operation costs.

## Networking and CDN Optimization

Network egress charges often surprise teams. The [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md) (lines 107‑111) recommends:

- **Location awareness** — Keep traffic within a single region to avoid inter-region egress fees.
- **Standard Network Service Tier** — Use the Standard tier for latency-tolerant traffic instead of the Premium tier.
- **Cloud CDN** — Cache content at the edge to reduce origin egress and compute load.

## Managed Services and Serverless Economics

Prefer fully managed services that bill per usage rather than provisioned capacity:

- **Cloud Run** and **Cloud Functions** scale to zero and bill only for actual request processing time.
- **GKE Autopilot** eliminates node management overhead and wasted capacity from over-provisioned node pools.
- **Cloud SQL** offers per-second billing and automatic storage increases to prevent manual over-provisioning.

*Reference:* [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md), lines 94‑100.

## Automation with Active Assist

Manual cost optimization does not scale. Implement **Active Assist** and **Recommender** APIs to automatically detect:

- Idle resources and unattached persistent disks
- Rightsizing opportunities for Compute Engine instances
- Unutilized commitments and reservations

Use the **FinOps Hub** for a consolidated view of savings opportunities across all projects, as documented in [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md), lines 82‑89.

## Continuous Monitoring and Budgeting

Establish guardrails to catch anomalies before they become invoices:

- **Budgets and alerts** — Configure billing alerts at configurable percentages (e.g., 50%, 90%) of monthly budget.
- **VPC Flow Logs sampling** — Reduce log ingestion costs by sampling at `flow_sampling = 0.1` and disabling entirely in non-production environments.

*Reference:* [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md), lines 75‑77 and [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md), lines 69‑71.

## End-to-End Implementation Workflow

Based on the skill definitions across the repository, implement cost optimization in this sequence:

1. **Enable visibility** — Activate BigQuery billing export and GKE cost allocation using the `gcloud` command above.
2. **Analyze current spend** — Execute BigQuery queries to identify top cost drivers:

```bash
bq query --nouse_legacy_sql '
SELECT
  SUM(cost) + SUM(IFNULL((SELECT SUM(c.amount) FROM UNNEST(credits) c), 0)) AS net_cost,
  labels.value AS cluster_name
FROM `myproject.billing_dataset.gcp_billing_export_resource_v1_*` AS bqe
LEFT JOIN UNNEST(bqe.labels) AS labels ON labels.key = "goog-k8s-cluster-name"
WHERE _PARTITIONTIME >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 30 DAY)
GROUP BY 2
ORDER BY net_cost DESC
LIMIT 10;
'

```

*Source:* [`skills/cloud/gke-cost-analysis/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-analysis/SKILL.md), query example lines 11‑26.

3. **Apply rightsizing** — Deploy VPA in recommendation mode, then adjust pod requests to `P95 × 1.2` safety margin.
4. **Adopt Spot instances** — Create a `ComputeClass` with Spot fallback for fault-tolerant workloads:

```yaml
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
  name: spot-with-fallback
spec:
  priorities:
  - machineFamily: n4
    spot: true
  - machineFamily: n4
    spot: false

```

*Source:* [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md), lines 24‑31.

5. **Purchase commitments** — Based on historical analysis, buy 1-year CUDs for consistently used machine families.
6. **Enforce governance** — Apply labels universally, deploy ResourceQuotas in all namespaces, and configure billing budgets.

## Summary

- The `google/skills` repository encodes Google Cloud cost optimization into actionable skills covering the Well-Architected Framework's Cost Optimization pillar.
- Foundational visibility requires BigQuery billing export, GKE cost allocation (`--enable-cost-allocation`), and consistent resource labeling.
- Compute savings come from CUDs for steady-state workloads, Spot VMs for batch jobs, and VPA/ResourceQuotas for Kubernetes rightsizing.
- Storage costs drop through lifecycle policies, Autoclass, and disabled versioning where compliance allows.
- Active Assist and FinOps Hub provide automated, continuous optimization beyond initial configuration.

## Frequently Asked Questions

### What are the most effective Google Cloud cost optimization strategies for GKE workloads?

The most effective strategies combine **GKE cost allocation** for visibility, **Vertical Pod Autoscaling** for rightsizing, **Spot VMs** for fault-tolerant workloads, and **Resource Quotas** to prevent over-provision. According to [`skills/cloud/gke-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-optimization/SKILL.md), enabling VPA in recommendation mode (lines 74‑90) before auto-applying changes prevents resource thrashing while identifying accurate request sizes.

### How do I enable cost allocation for GKE clusters?

Enable cost allocation using the gcloud CLI flag `--enable-cost-allocation` during cluster creation or updates. This populates the `goog-k8s-cluster-name` and namespace labels in BigQuery billing exports, allowing you to attribute spend to specific teams or applications. See [`skills/cloud/gke-cost-analysis/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gke-cost-analysis/SKILL.md), lines 85‑88 for the exact command syntax.

### What is the difference between Committed Use Discounts and Spot VMs?

**Committed Use Discounts (CUDs)** provide 1-year or 3-year price reductions for predictable, steady-state workloads that run continuously. **Spot VMs** offer 60–90 % discounts for fault-tolerant, interruptible workloads that can handle sudden termination. Use CUDs for production databases and Spot for batch processing or CI/CD runners, as outlined in [`skills/cloud/google-cloud-waf-cost-optimization/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-waf-cost-optimization/SKILL.md).

### How can I automate storage cost optimization in Google Cloud?

Enable **Autoclass** on Cloud Storage buckets to automatically transition objects between storage classes based on access frequency without manual lifecycle rules. For existing buckets, implement **lifecycle policies** to move data to Nearline, Coldline, or Archive classes after defined periods of inactivity. Review access patterns first using **Storage Insights** to avoid retrieval fees on archived data, per [`skills/cloud/google-cloud-storage-basics/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/google-cloud-storage-basics/SKILL.md).