How I saved my Kubernetes cluster from corrupted etcd secrets

Published: 2026-09-17

Somehow, when upgrading from Kubernetes from v1.36 to v.1.37, all my secrets got corrupted. The Kubernetes API was stuck in a Not ready state and trying to schedule a new pod would return with CreateContainerConfigError. Pretty scary! But this was fixable, and these are the steps I took to do just that.

Examining the logs and diagnosing the problem

Reading the logs by running kubectl logs -n kube-system kube-api-server showed me this error output:

cacher (secrets): unexpected ListAndWatch error: failed to list : failed to read one or more secrets from the storage: StorageError: corrupt object, Code: 7, Key: /registry/secrets/cert-manager/cert-manager-webhook-ca, ResourceVersion: 0, AdditionalErrorMsg: data from the storage is not transformable revision=0: no matching key was found for the provided Secretbox transformer (and 89 more); reinitializing…

Running kubectl get secrets -A also provided this error:

Error from server (StorageReadError): failed to read one or more secrets from the storage: StorageError: corrupt object, Code: 7, Key: /registry/secrets/cert-manager/cert-manager-webhook-ca, ResourceVersion: 0, AdditionalErrorMsg: data from the storage is not transformable revision=0: no matching key was found for the provided Secretbox transformer (and 89 more)

All signs pointed to the secrets within etcd (the underlying storage for the Kubernetes cluster state) being corrupted. My hypothesis was that, while upgrading Kubernetes, the decryption key that was handling all the etcd secrets (which were encrypted at rest) got rotated during the upgrade process and the original one was lost. This may be because I updated all my nodes all at once, resulting in an improperly handled transition. Either way, I had 90 corrupt secrets that needed to be handled.

How did I fix this?

Trying to delete the corrupted secrets with the API just returned the same errors as above. This meant I couldn’t rely on the Kubernetes API to remove the offending secrets. I had to bypass the API and manually remove these corrupt secrets from etcd directly and hope that would get the API back to a Ready state.

Here’s an overview of the steps I took to fix the issue:

  1. Generated etcdctl client key for my machine using openssl and my cluster configuration
  2. Fetched all the secret keys in /registry/secrets/ and cross examined the output with the logs
  3. Deleted all the corrupt secrets with a script
  4. Rebooted the nodes one by one

Connecting directly to etcd

I needed to connect with the cluster’s etcd directly using etcdctl. Because this isn’t the intended way to interface with a Kubernetes cluster, I didn’t have the necessary credentials to do that and I needed to mint client credentials for my machine using the information I had available.

Thankfully, I use Talos to manage my cluster nodes and was able to use the Talos API to retrieve my cluster’s etcd secrets from the machine configuration. I was then able to process and save necessary keys into the PEM formatted keys etcdctl expects.

Note: this command operates on a now deprecated but still supported v1alpha1 schema of the Talos machine configuration. My machine configuration was a multi doc YAML file which, Talos has now spit into multiple files since v1.14. Refer to the Talos Linux documentation for updated machine configuration retrieval.

# extract both the .cluster.etcd.ca.crt and .cluster.etcd.ca.key fields
# from the machine configuration and save as PEM formatted files.
cat machineconfig.yaml | \
yq -r 'select(documentIndex == 0) | .cluster.etcd.ca.crt' | \
base64 -d > etcd.crt

cat machineconfig.yaml | \
yq -r 'select(documentIndex == 0) | .cluster.etcd.ca.key' | \
base64 -d > etcd.key

With these files, now you can mint etcd credentials for your machine.

# genrate a generic client key for your machine
openssl genrsa -out client.key 2048
# generate a new signing request for that key
oppenssl req -new -key client.key -out client.csr -subj "/CN=talos-etcdctl"
# generate the client certificate using the request and the etcd secrets
openssl x509 -req -in client.csr \
-CA etcd.crt -CAkey etcd.key \
-Cacreateserial -out client.crt -days 1

Test the connection by fetching all the secret keys from an etcd node.

etcdctl \
--endpoints=https://<ETCD_NODE_IP>:2379 \
--cacert=etcd.crt \
--cert=client.crt \
--key=client.key \
get /registry/secrets/ --prefix --keys-only

If this works, congratulations, you’ve got direct access to your cluster’s etcd! I wrote the output to a file named broken.txt and inspected its contents.

Confirmation and secret deletion

Inspecting broken.txt, I ensured that each secret key was on its own line separated from the next key using only a \n to delimit them. The output produced 90 keys, the exact amount returned in the error logs!

I needed to delete all of them. And the best way to do so was to read and loop over the document’s contents.

wile IFS= read -r key; do
  echo "deleteing: $key";
  etcdctl \
  --endpoints=https://<ETCD_NODE_IP>:2379 \
  --cacert=etcd.crt \
  --cert=client.crt \
  --key=client.key \
  del "$key";
done < broken.txt

Ideally, the etcd cluster itself was still healthy and in sync, so this command only needed to be performed on one node. I was able to confirm this by running talosctl etcd status.

Rebooting the nodes

All that’s left was to reboot the nodes. Secrets were now completely gone, which still resulted in CreateContainerConfigErrors as pods got moved to a yet-to-be rebooted node; this was to to be expected. I just waited for a rebooted node to become ready and healthy before restarting another node. Many of the secrets were regenerated when pods got re-rescheduled onto a healthy node. Using a GitOps operator, like FluxCD or ArgoCD, helped a lot to automatically regenerate secrets. For secrets defined on my end, I relied on the the reconciliation processes of External Secrets Operator (more on that topic here) to regenerate those secrets! If you’re going through this process yourself and none of these operators are present, your Helm charts might need to be manually reapplied and any user-defined secrets imperatively redefined.

I gave it some time and the cluster naturally healed itself.

Reflections and how I could avoid this issue in the future

Before running the upgrade, I did a test-run on one of my development clusters and encountered no issues. So why this issue occurred in the first place probably boiled down to some real world environment variables (i.e. how long the production cluster had been running vs how long the development cluster had been running).

The issue might’ve been mitigated if I had run the upgrade on one node at a time, however! Running my tests had inflated my confidence, and proper procedure still needed to be respected. I’m glad that disaster was fixed by messing with etcd entries manually, but that was a last resort solution. Had only one node gone down, I’d still have had access to the secrets via Kubernetes’ API and a lot of downtime could’ve been avoided.

Do you have a cool idea? I love helping bring complex visions to life.

Say hi or follow me here: