Optimizing GKE Costs with Proper Resource Requests and Node Pool Configuration
Aligning pod resource requests with right-sized node pools and enabling autoscaling eliminates idle capacity and can reduce GKE spend by up to 80% while maintaining workload performance.
Google Kubernetes Engine (GKE) provides a managed environment for deploying containerized applications at scale, but inefficient resource allocation drives unnecessary cloud costs. The google/skills repository contains authoritative guidance in skills/cloud/workload-manager-basics/SKILL.md and skills/cloud/gcloud/SKILL.md for implementing cost controls through precise pod sizing and intelligent node pool architecture. By applying these patterns, you can optimize GKE costs with proper resource requests and node pool configuration without sacrificing reliability.
Right-Size Pod Resources with Accurate Requests and Limits
Setting precise CPU and memory requests ensures the Kubernetes scheduler packs workloads efficiently onto nodes, reducing the total node count required. When pods request only what they actually consume, you avoid the common anti-pattern of over-provisioning that leaves resources idle.
Define requests in your pod spec to match observed usage:
apiVersion: v1
kind: Pod
metadata:
name: web-app
spec:
containers:
- name: web
image: gcr.io/my-project/web-app:latest
resources:
requests:
cpu: "250m" # 0.25 vCPU
memory: "128Mi"
limits:
cpu: "500m"
memory: "256Mi"
This configuration requests minimal resources while setting hard limits to prevent runaway consumption.
Select Cost-Optimized Node Pool Machine Types
Matching node capacity to aggregate pod demand prevents paying for unused vCPU and memory. The repository's gcloud skill documentation demonstrates selecting e2-standard or e2-micro machine types for price-sensitive workloads.
Create a cost-optimized node pool using the patterns from skills/cloud/gcloud/SKILL.md:
gcloud container node-pools create low-cost-pool \
--cluster=my-gke-cluster \
--machine-type=e2-standard-2 \
--num-nodes=1 \
--enable-autoscaling \
--min-nodes=1 \
--max-nodes=5 \
--region=us-central1
Enable Cluster and Horizontal Pod Autoscaling
Cluster Autoscaler automatically adjusts node count based on pending pod scheduling, while Horizontal Pod Autoscaler (HPA) scales replica counts to match demand. Together, these mechanisms prevent over-provisioning during low-traffic periods.
As documented in the workload-manager basics skill, autoscaling is essential for maintaining the balance between performance and cost:
kubectl autoscale deployment web-app \
--cpu-percent=60 \
--min=2 \
--max=10
The Cluster Autoscaler works in tandem with HPA by adding nodes when pending pods cannot be scheduled and removing underutilized nodes when demand drops.
Leverage Preemptible VMs for Non-Critical Workloads
Preemptible VMs offer up to 80% cost savings compared to standard compute instances, making them ideal for batch processing and fault-tolerant applications. Configure a dedicated node pool in gcloud with the --preemptible flag:
gcloud container node-pools create batch-preemptible \
--cluster=my-gke-cluster \
--machine-type=e2-micro \
--preemptible \
--num-nodes=0 \
--enable-autoscaling \
--min-nodes=0 \
--max-nodes=4 \
--region=us-central1
Start with zero nodes and let autoscaling provision capacity only when batch jobs are submitted.
Isolate Workloads with Node Taints and Pod Tolerations
Use node taints to prevent general workloads from consuming expensive or specialized resources, and pod tolerations to allow specific jobs onto preemptible or dedicated nodes. This separation ensures critical services run on stable infrastructure while cost-sensitive workloads utilize discounted capacity.
Apply a taint to the preemptible pool:
gcloud container node-pools update batch-preemptible \
--cluster=my-gke-cluster \
--node-taints=preemptible=true:NoSchedule
Then add the corresponding toleration to batch job pod specs:
tolerations:
- key: "preemptible"
operator: "Equal"
value: "true"
effect: "NoSchedule"
Continuously Monitor and Adjust Resource Allocation
Cost optimization requires ongoing observation of actual utilization versus provisioned capacity. The repository's guidance in skills/cloud/workload-manager-basics/SKILL.md stresses that static configurations become inefficient as workloads evolve.
Use Cloud Monitoring dashboards to track CPU and memory utilization across namespaces. When pods consistently use less than requested, reduce the request values and let Cluster Autoscaler downscale the node pool. Conversely, increase limits before performance degrades.
Summary
- Right-size requests: Set CPU and memory requests to match actual consumption to maximize node packing efficiency.
- Choose appropriate machine types: Select e2-standard or smaller instance types for cost-sensitive pools.
- Enable autoscaling: Combine Cluster Autoscaler for nodes and Horizontal Pod Autoscaler for replicas to handle demand fluctuations.
- Use preemptible VMs: Configure node pools with
--preemptiblefor batch workloads to reduce compute costs by up to 80%. - Implement taints and tolerations: Direct workloads to appropriate node pools based on criticality and cost constraints.
- Monitor continuously: Adjust requests and limits based on observed utilization patterns from Cloud Monitoring.
Frequently Asked Questions
How do I determine the right CPU and memory requests for my GKE pods?
Analyze historical usage metrics from Cloud Monitoring or kubectl top pods to identify baseline consumption. Set requests slightly above the average usage (e.g., 25th percentile) and limits at the 95th percentile to handle spikes without over-provisioning. The workload-manager basics skill in google/skills recommends starting with conservative estimates and refining based on observed behavior.
What is the difference between Cluster Autoscaler and Horizontal Pod Autoscaler?
Cluster Autoscaler adjusts the number of nodes in your pool based on pending pods that cannot be scheduled, while Horizontal Pod Autoscaler changes the replica count of a deployment based on CPU, memory, or custom metrics. Cluster Autoscaler optimizes infrastructure costs by removing idle nodes, whereas HPA optimizes application costs by scaling down replicas during low demand.
When should I use preemptible VMs in GKE?
Use preemptible VMs for fault-tolerant, stateless batch jobs, CI/CD pipelines, or development environments that can tolerate interruption. According to the gcloud skill documentation, these instances provide up to 80% cost savings but may be terminated with 24 hours notice, making them unsuitable for long-running critical services or databases.
How do node taints help reduce GKE costs?
Node taints prevent pods from scheduling onto specific nodes unless they have matching tolerations. This allows you to isolate expensive workloads to dedicated node pools while keeping non-critical jobs on cheaper preemptible instances. By ensuring high-priority services do not consume discount capacity unnecessarily, you maximize the efficiency of your node pool architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →