Troubleshooting
Helm: another operation is in progress
The error means Helm found a release record that never finished. The earlier command was killed, timed out or cancelled, and the release is still marked pending. This guide shows how to confirm that, how to recover without losing the running application, and how to keep the lock from forming again.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- The message is
Error: UPGRADE FAILED: another operation (install/upgrade/rollback) is in progress. Helm is refusing to act on a release whose status is still pending-install, pending-upgrade or pending-rollback. - Run helm history first. A revision with STATUS deployed is the safe rollback target. If there is none, the release never installed and uninstall is the right step.
- Fix in this order: confirm nothing is running, roll back to the last deployed revision, uninstall only a failed first install, and delete the pending Secret only as a last resort.
- Prevent it with --atomic on Helm 3 or --rollback-on-failure on Helm 4, a bounded --timeout, and one writer per release.
What the error means
The full message from Helm is Error: UPGRADE FAILED: another operation (install/upgrade/rollback) is in progress. The word in parentheses changes with the command you ran. Helm prints the same error for install, upgrade and rollback. Running helm status on the affected release usually shows a STATUS of pending-install, pending-upgrade or pending-rollback.
Helm lists those states in its reference for the status command. The other states are unknown, deployed, uninstalled, superseded, failed and uninstalling. A release in a pending state is one that Helm has started changing and has not finished. The error is not about the cluster as a whole. It is about one release record.
Why Helm locks a release
Helm 3 and later keep no state outside the cluster. Each revision of a release is stored as a Secret in the release namespace, and the Secret carries the labels owner=helm and the release name. The Helm storage page states that release information is stored in Secrets by default, and gives the command kubectl get secret --all-namespaces -l "owner=helm" to list them.
When an upgrade starts, Helm writes a new revision with a pending status before it applies anything. When the work completes, Helm writes the revision again as deployed and marks the previous one superseded. That pending record is the lock. Helm checks it before it starts a second operation, and refuses to start while it is there. This is what stops two writers from changing the same release at once, and it is also why a dead writer blocks everyone after it.
The lock is only the status field in the record. The record does not say which process wrote it or whether that process is still alive. If the writer is gone, the status stays pending until a person or a script changes it.
How a release gets stuck
Four events leave the pending record behind. Each is common in CI and in Kubernetes clusters, and each shows up differently in the logs.
- The process was killed. The CI runner, a Kubernetes Job pod or a laptop was stopped during the upgrade. A pod eviction under memory pressure does the same thing when the Helm command runs inside a pod.
- Helm timed out in --wait. With
--wait, Helm keeps polling until resources are ready or the timeout passes. The default timeout is 5 minutes. If the timeout passes without--atomic, Helm stops, and the status can remain pending on some paths. Scheduling problems such as Pending pods are a common trigger. - Someone cancelled the job. A CI cancel, a Ctrl-C in a terminal, or a timeout set on the pipeline stops Helm mid-operation.
- Two writers ran at once. Two pipelines, or a pipeline and a person, ran an upgrade of the same release. The second one fails on the lock. If the first one then dies, nothing clears the record.
Two of these four are in your control: the timeout and the concurrency. The other two happen to any process, which is why the fix has to be a recovery step and not only prevention.
Diagnose the stuck release
Three read-only commands tell you what state the release is in. Run them before you change anything. The --max flag on helm history limits the list; the default is 256 revisions.
# Current state of the release: STATUS line and last revision
helm status my-release --namespace my-namespace
# Every revision, with its STATUS column
helm history my-release --namespace my-namespace --max 10In the history table, find the newest row whose STATUS is deployed. That revision is the one the running application matches. Note its number. If the newest row is pending and the one before it is deployed, the release can be rolled back. If no row is deployed, the first install never completed.
To see the Secrets behind the rows, list them by label. Each Secret holds one revision, so this confirms what Helm sees directly, without relying on the CLI output.
# Helm keeps one Secret per revision, labelled owner=helm.
# The name label is set to the release name by Helm.
kubectl get secrets --namespace my-namespace -l owner=helm,name=my-releaseTwo things to check before you act. First, confirm that no helm process is still running for this release: look at the CI job list, and at any Kubernetes Job or pod named after the release. Second, check whether the application is healthy right now with kubectl get pods in the release namespace. If the pods are healthy, the rest of this guide is about the release record, not the application.
Fix it in order of safety
Work through the steps in order and stop at the first one that works. The later steps are more destructive, and the earlier steps are often enough.
1. Wait for a live operation
If a helm command is still running, its status will change when it finishes. A pipeline that has been cancelled may still have a process on a runner. Wait for it to end, then look at the status again. Do not start a second command to speed this up.
2. Roll back to the last deployed revision
This is the normal fix. Rolling back creates a new revision that copies the deployed one, and it finishes the pending state. The rollback command takes a revision number. Omitting the revision, or passing 0, rolls back to the previous release. Pass the number you picked from helm history so you know what you are returning to.
# Roll back to the last revision whose STATUS is "deployed".
# Pick the number from helm history. Omitting it rolls back to the previous release.
helm rollback my-release 4 --namespace my-namespace --wait --timeout 5mConfirm with helm status my-release that the STATUS line now reads deployed. Then run your normal upgrade again.
3. Uninstall a first install that never deployed
If no revision in the history is deployed, rollback has nothing to return to. This happens when the very first install of a release was interrupted. Uninstall removes the resources the release owns and then the record. Use it only in that case, because on a release with a deployed revision it removes the running application.
# Only when no revision was ever "deployed" (a first install that failed).
helm uninstall my-release --namespace my-namespace --wait --timeout 5mIf you want the history kept, add --keep-history. That flag marks the release deleted and keeps its revisions, so the name can be reused without losing the record of what happened. Without it, uninstall removes the history as well.
4. Delete the pending Secret (last resort)
Deleting the Secret of the pending revision removes the lock directly. It is a last resort because it skips Helm. Helm does not clean up the resources the interrupted operation created, so you may have to remove those by hand, and the revision history will no longer describe what ran in the cluster. Take a copy of the Secret before you delete it.
# LAST RESORT. Find the Secret of the pending revision, then delete that one.
# Take a copy first, so the record can be restored if the diagnosis was wrong.
kubectl get secret <pending-secret-name> --namespace my-namespace -o yaml > pending-revision-backup.yaml
kubectl delete secret <pending-secret-name> --namespace my-namespaceAfter a manual delete, run helm status again, then check the cluster for resources the interrupted run may have left behind, such as a half-created Deployment or a Job that never finished. Helm will not reconcile those on its own.
Prevent it
Prevention has three parts: make a failed upgrade roll itself back, bound the time it can run, and allow one writer at a time. The flags depend on the Helm major version, so check which one your runners use with helm version.
Roll back on failure
On Helm 3, --atomic rolls back changes if the upgrade fails. It sets --wait automatically. On Helm 4.3, the upgrade reference lists --rollback-on-failure with the same intent, and says it defaults --wait to watcher. The Helm 4.3 upgrade reference does not list --atomic, so use the newer name there.
# Helm 3: roll back automatically on failure, wait for readiness, bound the time.
helm upgrade --install my-release ./my-chart \
--namespace my-namespace \
--atomic --wait --timeout 5m \
--history-max 10# Helm 4: the same behaviour is spelled --rollback-on-failure.
# --wait is set automatically by --rollback-on-failure.
helm upgrade --install my-release ./my-chart \
--namespace my-namespace \
--rollback-on-failure --timeout 5m \
--history-max 10Rollback-on-failure only helps when Helm is still running to do the rollback. If the process is killed, nothing rolls back. That is the case the rest of this section is for.
Bound the time
The --timeout flag sets the time Helm waits for each Kubernetes operation. Its default is 5 minutes. Set it deliberately, and set it lower than the time your CI job allows, so Helm finishes and writes its status before the runner kills it. A job timeout that is shorter than the Helm timeout is the usual way a pending record is born.
One writer per release
Run deploys for one release name through a single queue. Most CI systems have a concurrency setting that does this per key, such as the release name or the environment. Do not let a new pipeline start an upgrade while an earlier one is still running against the same release. Do not cancel a pipeline in the middle of a Helm step unless the next run includes the recovery steps above.
If you run Helm from a Kubernetes Job, give the Job enough memory and a priority that keeps it from being evicted during the install. An evicted Job is a killed process from Helm’s point of view.
How SeaGit handles this
SeaGit installs some cluster add-ons with Helm on clusters it provisions in your AWS account. The installer runs the same checks described above, and the limits are stated below.
- Before each add-on install. The installer reads the release state. A release in pending-install is uninstalled and then reinstalled. A release in pending-upgrade, pending-rollback or failed is rolled back to its last deployed revision, and it is uninstalled if there is no such revision.
- The install itself. It runs
helm upgrade --installwith--atomic --wait. The timeout is 5 minutes, or 6 minutes for the ingress controller, which schedules more slowly than the other add-ons. - Application deploys. The job that installs an application release runs the same pending-state check in its own script before the Helm command, so a retried job does not collide with a release its killed predecessor left pending.
- Limits. This applies to releases that SeaGit’s installer manages. A release you install with your own helm commands on your own clusters is not changed by it. The steps in this guide still apply to those.
For how SeaGit clusters and deployments are organised, see the clusters documentation, the deployments documentation, and the applications documentation. The troubleshooting guide lists the other common deploy failures.