Kubernetes troubleshooting
Issuing Certificate as Secret Does Not Exist
The message "Issuing certificate as Secret does not exist" is normal for a few seconds while the first certificate is issued. If it stays for minutes, the certificate is blocked at one step in a chain of objects that cert-manager creates. This guide shows how to find that step.
By Muhammad Soliman, Founder at SeaGit
·
Key takeaways
- The Certificate is only the summary. The real work happens in a CertificateRequest, then an ACME Order, then one or more Challenges.
- Read the Challenge first. A pending Challenge names the reason: the HTTP-01 path is not reachable, or the DNS-01 record was not created or did not propagate.
- Check the ClusterIssuer is Ready. A broken ACME account or a wrong email blocks every certificate that uses it.
- Re-issuing after repeated failures counts against Let's Encrypt limits. Test with the staging issuer first.
The chain of objects
When an Ingress or a Certificate asks for TLS, cert-manager creates a CertificateRequest. For an ACME issuer such as Let's Encrypt, it then creates an Order, and an Order creates one Challenge for each name. The Secret is written only after every Challenge passes.
kubectl get certificate,certificaterequest,order,challenge -n prod
kubectl describe certificate app-tls -n prod | sed -n '/Status:/,$p'Step 1: find the blocked object
Work down the chain. Stop at the first object that is not in a good state. Each describe command shows the events and the message.
# The request and its condition
kubectl describe certificaterequest <name> -n prod
# The ACME order, if one was created
kubectl describe order <name> -n prod
# The challenge: this is where most failures are reported
kubectl describe challenge <name> -n prodIf there is no Order at all, the CertificateRequest was not accepted by the issuer. Check the issuer, then the cert-manager logs:
kubectl describe clusterissuer letsencrypt-prod
kubectl logs -n cert-manager deploy/cert-manager --tail=100 | grep -i 'app-tls\|error'Cause 1: the HTTP-01 challenge cannot be reached
With HTTP-01, cert-manager creates a temporary Ingress that serves a token under /.well-known/acme-challenge/. Let's Encrypt must be able to fetch that URL from the internet. The challenge stays pending when the DNS name does not point at the right load balancer, when the ingress class is wrong, or when a redirect or firewall blocks port 80.
# The temporary solver ingress, and the class it uses
kubectl get ingress -n prod
kubectl get ingress -n prod -o jsonpath='{range .items[*]}{.metadata.name}{" class="}{.spec.ingressClassName}{"\n"}{end}'
# The public name must resolve to the controller's load balancer
dig +short app.example.com
# The token URL must answer over plain HTTP
curl -sS http://app.example.com/.well-known/acme-challenge/test | head -3Cause 2: the DNS-01 record is missing or not visible
With DNS-01, cert-manager writes a TXT record named _acme-challenge under the domain. The challenge fails if the DNS provider rejects the write, if the credentials lack permission, or if the record has not propagated when Let's Encrypt checks it.
# The TXT record should appear while the challenge is pending
dig +short TXT _acme-challenge.app.example.com
# On AWS, the issuer's credentials need these actions on the hosted zone
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::111122223333:role/cert-manager \
--action-names route53:ChangeResourceRecordSets route53:GetChange route53:ListHostedZonesA record that is created and then removed quickly is normal. A challenge that stays pending with no TXT record usually means the DNS-01 solver is not using the zone you expect. Check the zone name in the issuer against the domain.
Cause 3: the issuer or ACME account is broken
A ClusterIssuer whose ACME account was never registered, or whose private key secret was deleted, blocks every certificate that uses it. Its Ready condition shows the reason.
- Use the staging endpoint while you test: https://acme-staging-v02.api.letsencrypt.org/directory. Staging certificates are not trusted by browsers, but the limits are much higher.
- Set a real contact email in the issuer. A missing or malformed email is rejected by the ACME server.
- Do not delete the account key secret. A new account key means a new registration, which the server may rate limit.
Retry without burning Let's Encrypt limits
Failed validations count against Let's Encrypt limits. The one that matters most here is five failed validations per account, per hostname, per hour. A tight retry loop on a broken challenge uses that budget quickly. Fix the cause first, then retry once.
# Delete the failed request. cert-manager creates a new one from the Certificate
kubectl delete certificaterequest <name> -n prod
# Watch the Certificate until Ready=True
kubectl get certificate app-tls -n prod -wFor the case where a certificate hits the limit for an exact set of names, see the Let's Encrypt rate limit guide.
Confirm the Secret and the Ingress match
Once the Certificate is Ready, the Secret it names must exist, and the Ingress must reference that same Secret name:
kubectl get secret app-tls -n prod
kubectl get certificate app-tls -n prod -o jsonpath='{.spec.secretName}{"\n"}'
kubectl get ingress -n prod -o jsonpath='{range .items[*].spec.tls[*]}{.secretName}{"\n"}{end}'How SeaGit handles this
For domains that SeaGit manages in your AWS account, SeaGit sets up the cert-manager side for you. It creates a DNS-01 ClusterIssuer for the domain and a wildcard certificate, and it adds the matching cert-manager.io/cluster-issuer annotation to the app ingress. For those certificates the HTTP-01 reachability step above does not apply, and the DNS-01 section is the one to check.
The chain of objects is still there, so the commands in this guide apply to a SeaGit cluster as well. SeaGit does not hide a failed issuance: check the Certificate and Challenge in the same way. For the setup of domains and certificates, see the certificates and TLS docs.