Deploying Vertex AI and ML Infrastructure on GKE: Migration Patterns and Implementation Guide
Google Kubernetes Engine (GKE) serves as the preferred platform for production-grade, self-hosted AI/ML inference when you require tighter control over latency, cost, or custom hardware configurations compared to managed Vertex AI endpoints.
The google/skills repository provides comprehensive, battle-tested patterns for deploying Vertex AI and ML infrastructure on GKE, enabling you to either migrate existing workloads or provision fresh inference services. These guides address the complete lifecycle from initial discovery through production cut-over, with specific optimizations for GPU/TPU scheduling, model storage, and LLM-aware autoscaling.
When to Deploy Vertex AI Workloads on GKE
GKE becomes the optimal choice when your use case demands hardware flexibility, cost optimization at scale, or specific compliance requirements that managed endpoints cannot satisfy. The repository outlines two primary scenarios: migrating existing Vertex AI services to gain infrastructure control, and deploying net-new AI inference using modern GKE-native APIs.
Two Deployment Patterns for GKE-Based AI/ML
The google/skills repository defines distinct pathways depending on your starting state.
Pattern 1: Migrating Existing Vertex AI Workloads
Use the google-cloud-solution-guided-gke-ai-migration Skill when relocating models currently served on Vertex AI Endpoints or the Gemini Enterprise Agent Platform. This four-phase workflow, documented in skills/cloud/google-cloud-solution-guided-gke-ai-migration/SKILL.md, covers discovery, solution design, manifest generation, and traffic cutover.
The migration process begins with auditing your current Vertex AI configuration using gcloud ai endpoints list to capture endpoint specifications and model artifact locations. You then design a GKE-centric architecture selecting appropriate GPU tiers (NVIDIA L4, A100, H100) and storage backends before generating deployment manifests from the repository's template files.
Pattern 2: Fresh AI Inference Deployment
For new projects without existing Vertex AI endpoints, the gke-inference Skill in skills/cloud/gke-inference/SKILL.md provides a streamlined path using the AI Profiles API. This approach leverages gcloud container ai profiles to discover compatible accelerator configurations and generate ready-to-apply Kubernetes manifests without manual YAML templating.
Core Architectural Components
Custom Compute Classes for Accelerator Selection
GKE Autopilot utilizes Custom Compute Classes (CCC) to schedule pods onto specific GPU hardware. According to the source code in skills/cloud/gke-inference/SKILL.md, you define a ComputeClass to request exact accelerator types:
apiVersion: cloud.google.com/v1
kind: ComputeClass
metadata:
name: l4-inference
spec:
priorities:
- machineFamily: g2
gpu:
type: nvidia-l4
count: 1
minCores: 4
minMemoryGb: 16
This ensures the scheduler places inference workloads on nodes with the specific NVIDIA L4, A100, or H100 GPUs your model requires.
Model Storage and Cold Start Optimization
The repository recommends Cloud Storage FUSE CSI driver for mounting model buckets directly into pods, eliminating repeated downloads during pod initialization. For ultra-low latency requirements or PiB-scale workloads, Managed Lustre provides high-performance parallel filesystem access, as detailed in skills/cloud/google-cloud-solution-guided-gke-ai-migration/SKILL.md.
Ingress and Traffic Management
Expose LLM endpoints via the GKE Gateway API using the internal gke-l7-rilb class by default. For advanced traffic shaping, the optional GKE Inference Gateway adds LLM-aware routing through InferencePool resources, enabling sophisticated request distribution across model replicas.
Observability and GPU Monitoring
Enable Google Cloud Managed Service for Prometheus with DCGM metrics to track GPU utilization, memory consumption, and custom LLM metrics including queue size and request latency. This monitoring stack, referenced in the migration Skill, provides the visibility necessary to tune autoscaling behaviors.
Step-by-Step Migration Workflow
Phase 1: Discovery and Audit
Extract your existing Vertex AI endpoint configuration to understand current resource allocations and model sources:
# List existing Vertex AI endpoints
gcloud ai endpoints list --project=$PROJECT_ID
# Describe specific endpoint details
gcloud ai endpoints describe <ENDPOINT_ID> --project=$PROJECT_ID
Document the model artifact location, current secret bindings, and traffic patterns before proceeding.
Phase 2: Solution Design
Calculate required VRAM using the deterministic formula provided in the migration guide:
VRAM_total = (Params × 2 / Quantization + KV_Cache) × 1.2
Map this calculation to an accelerator tier, decide between FUSE versus Lustre storage, and determine if custom autoscaling metrics are required based on queue depth rather than raw GPU utilization.
Phase 3: Implementation and Manifest Generation
Generate concrete YAML files from the repository's assets/ directory templates. Key files include vllm-deployment.yaml.tmpl for the model server, ccc-profile.yaml.tmpl for GPU selection, and gke-inference-gateway.yaml.tmpl for ingress configuration.
For gated Hugging Face models, create secrets before applying manifests:
kubectl create secret generic hf-secret \
--namespace=my-namespace \
--from-literal=hf_api_token=<YOUR_HF_TOKEN>
Apply the generated configurations:
kubectl apply -f vllm-deployment.yaml
kubectl apply -f ccc-profile.yaml
Phase 4: Validation and Traffic Cut-Over
Verify pod health and run inference tests using port-forwarding before migrating production traffic:
kubectl port-forward svc/my-model-vllm-svc 8000:8000 &
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"my-model","messages":[{"role":"user","content":"Hello"}]}'
Once validated, update your application's endpoint configuration to target the GKE service rather than the Vertex AI endpoint.
AI Profiles API for New Deployments
For fresh deployments, the AI Profiles API automates hardware compatibility checking and manifest generation:
# Discover available profiles for a specific model
gcloud container ai profiles list \
--model=gemma-2-9b-it \
--quiet
# Generate optimized manifest
gcloud container ai profiles manifests create \
--model=gemma-2-9b-it \
--model-server=vllm \
--accelerator-type=nvidia-l4 \
--target-ntpot-milliseconds=50 \
--quiet > inference.yaml
# Deploy to cluster
kubectl apply -f inference.yaml
Security and Operational Best Practices
Workload Identity: Bind GKE service accounts to Google service accounts with roles/storage.objectUser for secure bucket access without static credentials, as implemented in the migration security section.
Secret Management: Never embed token values in Git-tracked manifests. Always use Kubernetes Secrets created via kubectl create secret generic before applying deployment jobs.
Autoscaling Awareness: LLM workloads often appear fully utilized because vLLM pre-allocates VRAM for KV-cache. According to skills/cloud/gke-inference/SKILL.md, prefer queue-size or custom GPU metrics over raw GPU utilization to prevent premature or unnecessary scaling events.
Summary
- The
google/skillsrepository provides two validated patterns for deploying Vertex AI and ML infrastructure on GKE: migration of existing endpoints and fresh deployment via AI Profiles. - Custom Compute Classes enable precise GPU selection (L4, A100, H100) and scheduling on GKE Autopilot clusters.
- Cloud Storage FUSE and Managed Lustre offer tiered storage solutions for model artifacts, optimizing for either simplicity or high-throughput access.
- The four-phase migration workflow (Discovery, Design, Implementation, Validation) ensures systematic relocation of Vertex AI workloads with minimal downtime.
- GKE Gateway API and optional Inference Gateway provide sophisticated traffic management and LLM-aware routing capabilities.
- Queue-based autoscaling metrics prevent over-scaling of LLM inference pods that naturally maintain high VRAM allocation.
Frequently Asked Questions
When should I migrate from Vertex AI to GKE instead of staying on the managed service?
Migrate when you require specific GPU types not available in Vertex AI, need to co-locate inference with preprocessing pipelines in the same cluster, or want to optimize costs through custom autoscaling policies and sustained-use discounts. The migration Skill provides a checklist to validate these requirements before committing to the move.
How do I calculate the correct GPU memory requirements for my model?
Use the deterministic VRAM sizing formula provided in skills/cloud/google-cloud-solution-guided-gke-ai-migration/SKILL.md: VRAM_total = (Params × 2 / Quantization + KV_Cache) × 1.2. This accounts for model parameters at your chosen quantization level (e.g., FP16, INT8) plus the KV-cache overhead required for context windows, with a 20% safety buffer for system overhead.
What is the difference between Cloud Storage FUSE and Managed Lustre for model storage?
Cloud Storage FUSE mounts GCS buckets as filesystems in your pods, offering simplicity and cost-effectiveness for standard inference workloads. Managed Lustre provides a high-performance parallel filesystem designed for ultra-low latency or petabyte-scale model repositories where FUSE latency would bottleneck throughput. Choose Lustre only when FUSE performance proves insufficient for your latency requirements.
How does autoscaling work differently for LLM inference compared to traditional web services?
Traditional CPU-based autoscaling triggers on utilization percentages, but LLM inference servers like vLLM pre-allocate VRAM for KV-cache, making GPU utilization appear constantly high. According to the gke-inference Skill, you should configure HorizontalPodAutoscaler to scale on request queue depth or custom metrics like time-to-first-token (TTFT) rather than GPU utilization percentages to ensure scaling reflects actual demand rather than memory allocation patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →