Cloud SQL High Availability Configuration: A Complete Terraform Implementation Guide
Use the enable_ha variable in google/skills to toggle Regional High Availability, which sets availability_type = "REGIONAL" and enables point-in-time recovery for automatic failover across zones.
The google/skills repository demonstrates production-ready Cloud SQL high availability (HA) through declarative Terraform configuration. This implementation combines regional replication, private connectivity, and IAM-based authentication to deliver enterprise-grade database resilience without public internet exposure.
Understanding Cloud SQL High Availability in the Skills Repository
High availability in this architecture means synchronous replication across multiple zones within a single region. When enable_ha = true, the Terraform module automatically provisions a standby instance that maintains real-time consistency with the primary.
The failover mechanism achieves < 60 seconds recovery time with zero data loss for committed transactions. This satisfies most production SLAs without the complexity of cross-region read replicas.
The enable_ha Variable and Core Configuration
Variable Declaration
The HA toggle lives in assets/main.tf as a simple boolean:
variable "enable_ha" {
type = bool
description = "Whether to configure Regional High Availability (availability_type = REGIONAL) and Point-in-Time Recovery for Cloud SQL."
default = false
}
This single variable drives three critical behaviors: availability type selection, backup configuration, and point-in-time recovery enablement.
Instance Resource Definition
The google_sql_database_instance resource implements conditional logic based on enable_ha:
resource "google_sql_database_instance" "private_db" {
name = "private-postgres-db"
database_version = "POSTGRES_18"
region = var.region
deletion_protection = true
settings {
edition = var.db_edition
availability_type = var.enable_ha ? "REGIONAL" : "ZONAL"
backup_configuration {
enabled = true
point_in_time_recovery_enabled = var.enable_ha
}
ip_configuration {
ipv4_enabled = false
psc_config {
psc_enabled = true
allowed_consumer_projects = [var.project_id]
}
}
database_flags {
name = "cloudsql.iam_authentication"
value = "on"
}
}
}
Key implementation details from assets/main.tf lines 31-39:
availability_typeswitches between"ZONAL"(single zone) and"REGIONAL"(multi-zone)point_in_time_recovery_enabledmirrors the HA setting since continuous transaction logs are essential for regional failover
HA Architecture Components
Regional Availability Type
REGIONAL availability creates a primary instance in one zone with a synchronous standby in another zone of the same region. The Cloud SQL control plane manages:
- Automatic health checks every few seconds
- Synchronous replication of write-ahead logs
- Automatic promotion of standby on primary failure
- DNS endpoint updates to redirect application traffic
This differs from ZONAL deployment, which operates a single instance with no automatic recovery capability.
Point-in-Time Recovery (PITR)
When HA is enabled, point_in_time_recovery_enabled = true activates continuous transaction log archiving. This provides:
- Recovery to any second within the 7-day retention window
- Protection against application-level data corruption
- Foundation for cross-region disaster recovery strategies
Private Service Connect Integration
The repository enforces Private Service Connect (PSC) for all connectivity, regardless of HA status. From assets/main.tf lines 41-48:
ip_configuration {
ipv4_enabled = false
psc_config {
psc_enabled = true
allowed_consumer_projects = [var.project_id]
}
}
This configuration ensures database traffic never traverses the public internet, maintaining security boundaries during failover events.
IAM Database Authentication
The cloudsql.iam_authentication = "on" database flag enables passwordless authentication through Cloud SQL Auth Proxy. From assets/main.tf lines 50-53:
database_flags {
name = "cloudsql.iam_authentication"
value = "on"
}
The Auth Proxy sidecar automatically handles:
- Short-lived certificate rotation (every 10 minutes)
- Secure tunnel establishment over PSC
- Failover awareness without application reconfiguration
Deploying High Availability
Enable HA via Terraform Variables
Create or modify a terraform.tfvars file:
enable_ha = true
region = "us-central1"
Or pass directly during apply:
terraform init
terraform apply -var="enable_ha=true" -var="region=us-central1"
Verify HA Configuration Post-Deployment
# Check availability type
gcloud sql instances describe private-postgres-db \
--format='value(settings.availabilityType)'
# Expected output: REGIONAL
# Verify PITR status
gcloud sql instances describe private-postgres-db \
--format='value(settings.backupConfiguration.pointInTimeRecoveryEnabled)'
# Expected output: True
Cloud Run Integration with HA-Aware Proxy
The repository's recommended deployment pattern uses the built-in Cloud SQL Auth Proxy sidecar. From SKILL.md lines 55-60, the connection path uses Unix sockets:
containers:
- name: backend
image: gcr.io/$PROJECT_ID/backend-image
volumeMounts:
- name: cloudsql-instance
mountPath: /cloudsql
env:
- name: DB_SOCKET_PATH
value: "/cloudsql/PROJECT_ID:REGION:INSTANCE_NAME"
volumes:
- name: cloudsql-instance
cloudSqlInstance:
instance: private-postgres-db
The sidecar transparently handles failovers—applications continue using the same socket path while the proxy reconnects to the new primary.
Production Best Practices
Monitoring and Alerting
Enable Cloud Monitoring integration to track HA events:
variable "enable_monitoring" {
type = bool
default = true
}
# In Cloud Run service configuration
env:
- name: ENABLE_QUERY_INSIGHTS
value: "true"
Critical metrics to alert on:
cloudsql.googleapis.com/database/failover_countcloudsql.googleapis.com/database/replication_lagcloudsql.googleapis.com/database/availability
Cost Considerations
Regional HA doubles compute costs (primary + standby) and increases storage costs by ~100% due to synchronous replication. The repository uses Cloud SQL Enterprise edition to access HA features; Standard edition does not support REGIONAL availability.
Disaster Recovery Extension
For cross-region resilience, combine HA with cross-region read replicas:
resource "google_sql_database_instance" "replica" {
name = "private-postgres-db-replica"
database_version = "POSTGRES_18"
region = "us-west1" # Different region
master_instance_name = google_sql_database_instance.private_db.name
settings {
tier = var.db_tier
availability_type = "ZONAL" # Replicas use ZONAL
}
}
Source Files Reference
| File Path | Purpose |
|---|---|
skills/cloud/google-cloud-solution-n-tier-serverless-web-app/assets/main.tf |
Complete Terraform definition including HA variable, instance configuration, PSC setup, and IAM flags |
skills/cloud/google-cloud-solution-n-tier-serverless-web-app/SKILL.md |
Architectural documentation and deployment workflow |
skills/cloud/cloud-sql-basics/references/dr-backups.md |
Backup and disaster recovery concepts |
skills/cloud/google-cloud-solution-architecture/references/best-practices-guides.md (lines 95-102) |
Cloud SQL operational best practices |
Summary
- Single variable control:
enable_hatoggles all HA-related settings consistently - Regional replication:
availability_type = "REGIONAL"provides automatic cross-zone failover - Integrated security: PSC and IAM authentication work identically in HA and non-HA modes
- Zero application changes: Cloud SQL Auth Proxy handles failover transparency
- Production requirement: Enterprise edition required; monitoring strongly recommended
Frequently Asked Questions
What happens during a Cloud SQL failover with HA enabled?
The Cloud SQL control plane detects primary instance failure through health checks, promotes the standby to primary, and updates the connection endpoint—all within 60 seconds. Applications using the Cloud SQL Auth Proxy automatically reconnect without configuration changes. No committed data is lost due to synchronous replication.
Does enabling HA affect database performance?
Regional HA introduces slight write latency overhead (typically < 5ms) because writes must commit to both primary and standby before acknowledgment. Read performance remains unchanged. The repository mitigates this through PSC, which provides lower-latency connectivity than public IP.
Can I enable HA on an existing ZONAL instance?
Yes, but it requires instance recreation. Modify enable_ha = true and run terraform apply; Terraform will destroy and recreate the instance with REGIONAL availability. Plan for maintenance window migration or use export/import to minimize downtime for production databases.
Is point-in-time recovery available without HA?
No in this implementation. The repository couples point_in_time_recovery_enabled to enable_ha because continuous transaction logging is required for both features and significantly impacts storage costs. For PITR without HA, modify the backup_configuration block to decouple these settings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →