Etcd restore
Hi,
This question is related to etcd backup and restore.
I've set up a kubernetes cluster using stacked etc topology using kubeadm. Dual control-plane nodes and dual worker-nodes. I end up with the static pod manifests in /etc/kubernetes/manifests and the control-plane services are running as pods. Only kubelet is running as a systemd service.
I've created a snapshot of my etcd databases on both control-plane nodes.
I simulate a data failure:
1. Stop API servers by removing the manifest files.
2. Delete /var/lib/etcd/ on both control-plane nodes.
Now I want to do a complete restore of etcd database.
I'm doing it using etcdctl on both control-plane nodes.
It seems to work and I get the output:
2022-07-01 09:05:14.228623 I | mvcc: restore compact to 42333 2022-07-01 09:05:14.233684 I | etcdserver/membership: added member a874c87fd42044f [https://127.0.0.1:2380] to cluster c9be114fc2da2776
kubectl get nodes is showing all nodes ready. However, when I do a single change, such as scaling a deployment. Something gets really wrong. Suddenly the API servers disagree on node health, and the replication change is not performed.
Am I doing this wrong?
Best Answer
-
@oleksazhel Thank you for your willingness to help.
I think my error was to do separate etcd snapshot. One for each node. Then I restored each node from its own snapshot.
I read in etcd documentation that "all that is needed is a single snapshot “db” file" and "all members should restore using the same snapshot."
https://etcd.io/docs/v3.5/op-guide/recovery/Doing like this made it work:
- Make one snapshot from a control-plane node.
- Stop etcd, api-server, kube-scheduler, kube-controller on both nodes.
- Delete etcd data-dir on both nodes to simulate data loss.
- Restore etcd data-dir on first node using
etcdctl snapshot restore - Copy snapshot to second node and restore using
etcdctl snapshot restore - Start etcd and verify
etcdctl endpoint health,etcdctl endpoint status,etcdctl member list - Start api-server, kube-controller and kube-scheduler
- Restart kubelet
1
Answers
-
@pnts Could you provide output of
ETCDCTL_API=3 etcdctl -w table \ --cacert /etc/kubernetes/pki/etcd/ca.crt \ --cert /etc/kubernetes/pki/etcd/server.crt \ --key /etc/kubernetes/pki/etcd/server.key \ member list
and
ETCDCTL_API=3 etcdctl -w table \ --endpoints <CP1_IP_ADDRESS>:2379,<CP2_IP_ADDRESS>:2379 \ --cacert /etc/kubernetes/pki/etcd/ca.crt \ --cert /etc/kubernetes/pki/etcd/server.crt \ --key /etc/kubernetes/pki/etcd/server.key \ endpoint status
0
Categories
- All Categories
- 177 LFX Mentorship
- 177 LFX Mentorship: Linux Kernel
- 750 Linux Foundation IT Professional Programs
- 373 Cloud Engineer IT Professional Program
- 169 Advanced Cloud Engineer IT Professional Program
- 74 DevOps IT Professional Program - Discontinued
- 4 DevOps & GitOps IT Professional Program
- 99 Cloud Native Developer IT Professional Program
- 7.6K Training Courses & Learning Paths
- 1 AI & ML Training
- 1 Blockchain & Decentralized Identity Training
- 5 Cloud & Containers Training
- 1 Cybersecurity Training
- 2 DevOps & Site-Reliability Training
- 1 Linux Kernel Development Training
- 1 Networking Training
- 2 Open Source Best Practice Training
- 1 System Administration Training
- 1 System Engineering Training
- 1 Web & Application Development Training
- 792 Hardware
- 202 Drivers
- 68 I/O Devices
- 37 Monitors
- 95 Multimedia
- 173 Networking
- 91 Printers & Scanners
- 87 Storage
- 769 Linux Distributions
- 81 Debian
- 68 Fedora
- 22 Linux Mint
- 13 Mageia
- 24 openSUSE
- 150 Red Hat Enterprise
- 31 Slackware
- 13 SUSE Enterprise
- 356 Ubuntu
- 465 Linux System Administration
- 31 Cloud Computing
- 73 Command Line/Scripting
- Github systems admin projects
- 98 Linux Security
- 78 Network Management
- 101 System Management
- 46 Web Management
- 106 Mobile Computing
- 18 Android
- 73 Development
- 1.2K New to Linux
- 1K Getting Started with Linux
- 392 Off Topic
- 121 Introductions
- 181 Small Talk
- 29 Study Material
- 956 Programming and Development
- 310 Kernel Development
- 628 Software Development
- 984 Software
- 376 Applications
- 182 Command Line
- 5 Compiling/Installing
- 68 Games
- 317 Installation
- Archived
- 2 LFD140 Class Forum
- 1.4K LFS258 Class Forum
Upcoming Training
-
August 20, 2018
Kubernetes Administration (LFS458)
-
August 20, 2018
Linux System Administration (LFS301)
-
August 27, 2018
Open Source Virtualization (LFS462)
-
August 27, 2018
Linux Kernel Debugging and Security (LFD440)