Kubernetes troubleshooting
Helm Upgrade: Atomic, Dry Run and Rollback
helm upgrade changes a release in place. Four flags decide how safe that is: --dry-run to preview, --wait to wait for ready pods, --atomic to roll back on failure, and --force to replace resources that cannot be patched. This guide explains each one and how to recover from a failed upgrade.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- Preview with helm upgrade --dry-run=server. It renders against the live cluster, so it catches many errors that a client-side dry run misses.
- --atomic implies --wait and rolls the release back to the last good revision when the upgrade fails or times out.
- --force replaces resources by delete and recreate when a patch is rejected. That can cause downtime, so use it on purpose.
- If an upgrade was killed midway, the release can be left pending-upgrade. Roll back to the last deployed revision before you retry.
The upgrade command and the flags that matter
Use --install so the same command creates the release on the first run and upgrades it afterwards. A safe deploy command for a pipeline looks like this:
helm upgrade --install my-api ./charts/my-api \
--namespace prod --create-namespace \
--values values/prod.yaml \
--set image.tag="$GIT_SHA" \
--atomic --timeout 5m- --dry-run=server sends the rendered manifests to the API server without applying them. Use it before the first production upgrade of a new chart version.
- --diff needs the helm-diff plugin. It shows the changed lines of the rendered manifests against the live release.
- --wait waits until Deployments, StatefulSets and Jobs are ready. --atomic includes --wait.
- By default an upgrade does not reuse the values of the last release. A value set with --set last time must be passed again. --reuse-values keeps the stored values, and --reset-values discards them. Pick one on purpose.
- --cleanup-on-fail removes resources created by a failed install or upgrade. It does not roll back an existing release.
Preview the upgrade with --dry-run=server
A client-side dry run only renders the chart. It does not check the API server, so a bad field or a missing CRD appears only at apply time. A server dry run asks the API server to validate each object.
helm upgrade my-api ./charts/my-api -n prod \
-f values/prod.yaml --set image.tag="$GIT_SHA" \
--dry-run=server
# Show what would change, with the helm-diff plugin installed
helm diff upgrade my-api ./charts/my-api -n prod \
-f values/prod.yaml --set image.tag="$GIT_SHA"Let --atomic roll back a failed upgrade
A plain helm upgrade that fails leaves the release at the new revision, with some pods updated and some not. --atomic prevents that. When the upgrade fails, or the wait times out, Helm restores the previous revision and reports the failure. Check the outcome from the release history:
helm history my-api -n prod
# Roll back to a specific revision by number
helm rollback my-api 7 -n prod --wait
# Confirm the release is deployed
helm status my-api -n prod--atomic cannot help if Helm itself is killed while it waits. A CI job that times out or a pod that is evicted can stop Helm before it rolls back. That is how a release gets stuck, and it is covered below.
Use --force only when a patch is rejected
Some fields cannot be changed in place. The most common one is the label selector of a Deployment:
Error: UPGRADE FAILED: cannot patch "my-api" with kind Deployment:
Deployment.apps "my-api" is invalid: spec.selector: Invalid value: ...: field is immutable--force tells Helm to delete and recreate the resource instead of patching it. The pods go away during the replacement, so the service has a gap. For a selector change, a planned maintenance window or a new release name is often safer than --force. Use it only after you have read the diff.
Recover from a release stuck in pending-upgrade
When an upgrade is killed midway, the release status stays pending-upgrade. Every later upgrade fails with another operation is in progress:
helm status my-api -n prod | grep STATUS
# STATUS: pending-upgrade
helm history my-api -n prod
# Find the last revision with STATUS deployed, then roll back to it
helm rollback my-api 6 -n prod --waitAvoid deleting the Helm release secrets by hand unless you have read the history and backed them up. Helm stores each revision as a secret named sh.helm.release.v1.<release>.v<n>, and removing the wrong one can make the release history inconsistent. The full recovery steps for the in-progress error are in the post linked below.
Timeouts that are not a stuck release
UPGRADE FAILED: timed out waiting for the condition means the new pods did not become ready. Check the new pods, their probes and their events before you change the timeout. A longer timeout only helps when the pods are slow to start, not when they crash.
kubectl get pods -n prod -l app.kubernetes.io/instance=my-api
kubectl describe pod <new-pod> -n prod | sed -n '/Events:/,$p'How SeaGit handles this
On SeaGit you do not run helm for your apps. SeaGit runs the install and upgrade for you, in your cluster. Its cluster add-on installs use the same pattern as this guide: helm upgrade --install with --atomic and --wait and a timeout, so a failed add-on install rolls back instead of leaving a half-applied release.
SeaGit's app deployer also handles the stuck case. If a previous run was killed and left a release in pending-install, the deployer clears that state before it reinstalls, so a retry does not fail on another operation in progress. The check and the clean-up run inside the deploy, so you do not need to run helm rollback yourself for that case.
SeaGit does not change how your chart renders or which values it uses. Chart and value changes still need the preview steps above.