How Kubernetes Rolling Updates and Rollbacks Work: A Complete Guide
Kubernetes rolling updates allow you to update container images and pod templates without downtime by gradually replacing old pods with new ones, while rollbacks restore previous ReplicaSet revisions when issues occur.
The litu54/DevOps-Interview-Guide repository documents how Kubernetes handles application updates through declarative Deployment controllers. This guide explains the mechanics behind zero-downtime deployments, the specific parameters that govern update behavior, and the revision history system that enables rapid recovery from failed releases.
Understanding Kubernetes Rolling Updates
A rolling update is a Deployment strategy that replaces old pods with new ones incrementally. When you modify a Deployment's pod template—such as changing a container image or environment variables—the Deployment controller creates a new ReplicaSet rather than modifying the existing one.
The Deployment Controller and ReplicaSets
The controller manages the transition by orchestrating two ReplicaSets simultaneously:
- New ReplicaSet: Created from the updated pod specification
- Old ReplicaSet: Contains the previous pod specification being replaced
As the new pods become Ready, the controller terminates pods from the old ReplicaSet. This process continues until all replicas run the new version.
Key Parameters: maxSurge and maxUnavailable
Two critical parameters in the Deployment strategy control the update behavior:
-
maxSurge: The maximum number of additional pods that can be created above the desired replica count during the update. Setting
maxSurge: 1allows one extra pod to exist temporarily, ensuring capacity during the transition. -
maxUnavailable: The maximum number of pods that may be unavailable during the update. Setting
maxUnavailable: 0(combined with an appropriate maxSurge) ensures zero downtime by preventing any reduction in available capacity.
These parameters are documented in the interview guide within SquareOps/DevOps_Engineer.md, which specifically questions candidates on how to implement rolling deployments using these settings.
Step-by-Step Rolling Update Process
The Deployment controller follows this sequence according to the Kubernetes implementation referenced in HCL/DevOps_Engineer_2.md:
- Create a new ReplicaSet from the updated Deployment specification
- Scale up the new ReplicaSet while respecting the
maxSurgeconstraint - Wait for new pods to pass readiness probes and report Ready status
- Scale down the old ReplicaSet while respecting the
maxUnavailablelimit - Repeat the scale-up/scale-down cycle until all old pods are replaced
If you need to intervene during the process, you can pause the rollout to inspect the new pods before committing to the full migration.
How Kubernetes Rollbacks Work
Kubernetes maintains a revision history of all previous ReplicaSets. When a rolling update fails or produces unexpected behavior, you can revert to any previous revision stored in this history.
Revision History Management
The system stores old ReplicaSets as revisions rather than deleting them immediately. You can view the revision history using:
kubectl rollout history deployment/web-app
Each revision represents a complete snapshot of the previous pod template, including container images, resource limits, and configuration.
Executing Rollbacks
Rollback re-creates the previous ReplicaSet and scales it up while terminating the current version. According to TCS/SRE_1.md, you can rollback to the immediate previous revision or a specific revision:
# Rollback to previous revision
kubectl rollout undo deployment/web-app
# Rollback to specific revision (e.g., revision 2)
kubectl rollout undo deployment/web-app --to-revision=2
The same file also addresses failure scenarios in the entry "If a rollback fails, how will you handle it?", emphasizing that manual intervention may be required if the rollback mechanism itself encounters issues.
Practical kubectl Commands for Updates and Rollbacks
Trigger and manage rolling updates using these commands from the repository examples:
# Trigger a rolling update by changing the container image
kubectl set image deployment/web-app web-app=repo/web-app:v2
# Monitor rollout progress in real-time
kubectl rollout status deployment/web-app
# Pause a rollout to troubleshoot or inspect new pods
kubectl rollout pause deployment/web-app
# Resume a paused rollout
kubectl rollout resume deployment/web-app
# View complete revision history
kubectl rollout history deployment/web-app
# Rollback to previous revision
kubectl rollout undo deployment/web-app
# Rollback to specific revision number
kubectl rollout undo deployment/web-app --to-revision=2
Interview Perspectives from the Repository
The litu54/DevOps-Interview-Guide repository contains specific technical questions that validate understanding of these concepts:
TCS/SRE_1.md: Documents the specific command for rolling back to a revision and asks how to handle failed rollbacksSquareOps/DevOps_Engineer.md: Tests knowledge ofmaxSurgeandmaxUnavailableconfigurationHCL/DevOps_Engineer_2.md: Explores the definition of rolling updates and asks about deployment strategies when both Deployments and StatefulSets are involvedVirtusa/Tech_Lead.md: References rolling updates alongside canary deployment strategies in YAML configurations
Summary
- Rolling updates work by creating new ReplicaSets and gradually scaling them up while scaling down old ones, governed by
maxSurgeandmaxUnavailableparameters. - Revision history preserves old ReplicaSets, enabling rollbacks to any previous state using
kubectl rollout undo. - Zero-downtime deployments require configuring
maxUnavailable: 0to ensure pods remain available throughout the update process. - StatefulSets behave differently from Deployments regarding rolling updates, as noted in the HCL interview materials.
- Pause and resume capabilities allow you to inspect new pods before committing to a full rollout.
Frequently Asked Questions
What is the difference between a rolling update and a recreate strategy?
A rolling update gradually replaces old pods with new ones to maintain availability, while a recreate strategy terminates all existing pods before creating new ones, causing downtime. The rolling update strategy is the default for Deployments and is essential for production workloads that require continuous availability.
How do maxSurge and maxUnavailable interact during an update?
These parameters work together to control capacity and availability. maxSurge allows temporary over-provisioning (e.g., running 4 pods when 3 are desired), while maxUnavailable sets a floor on how many pods can be down (e.g., 0 means all 3 must remain available). You can set maxSurge: 25% and maxUnavailable: 25% to allow flexibility in either direction without violating capacity constraints.
What happens to a failed rollout if I don't manually intervene?
If new pods fail readiness checks or crash during a rolling update, the Deployment controller will halt the rollout but keep the old ReplicaSet running. The system does not automatically rollback; you must manually execute kubectl rollout undo or fix the underlying issue and resume. As documented in TCS/SRE_1.md, interviewers specifically ask how to handle situations where rollbacks themselves fail, requiring manual pod deletion or revision management.
Can I rollback a StatefulSet the same way as a Deployment?
No, StatefulSets manage pods with stable identities and persistent storage, making rollbacks more complex. While Deployments automatically maintain revision history and support kubectl rollout undo, StatefulSets require manual intervention or the use of kubectl rollout undo with caution, as it may not handle persistent volume claims correctly. The HCL/DevOps_Engineer_2.md file specifically highlights this distinction when asking about strategies involving both resource types.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →