Recommendation on 20GB+ per node (VM) disk space to avoid "Evicted pod death-spiral"
Hopefully this will help others who have taken their time going through the LFS258 material over a longer period of time, and ended up with dozens of pods in "Evicted" state around Lab 11.1, 11.2, or after there.
Any long-running pods from earlier in the lab builds will gradually increase their local disk usage over time, until ultimately the nodes will "Evict" any newly-scheduled pods. This results in a "death-spiral" because the only way to reclaim the space completely is to undeploy ~everything. Since the labs use the CP node as an additional Worker node, this impacts ~everything else.
One solution is to simply not leave your environment running overnight after linkerd & linkerd-viz are deployed, since it's the first Lab that really consumes a lot of space over time.
Alternatively, increase the disk space on your VMs up to at least 20GB, and that should give you enough buffer to continue through the end of all labs. (Search "growpart" and "lvextend" online, and "qemu-img resize" if you're on KVM.) In my case, adding the ingress controller pushed things into the endless spiral of ephemeral disk space being claimed then evicted, and thus it was far faster to resize the VMs and "throw space at the problem" rather than reconfigure external volumes, edit logging, etc.
Hopefully that helps someone else too.
Categories
- All Categories
- 178 LFX Mentorship
- 178 LFX Mentorship: Linux Kernel
- 775 Linux Foundation IT Professional Programs
- 384 Cloud Engineer IT Professional Program
- 175 Advanced Cloud Engineer IT Professional Program
- 75 DevOps IT Professional Program - Discontinued
- 7 DevOps & GitOps IT Professional Program
- 103 Cloud Native Developer IT Professional Program
- 7.6K Training Courses & Learning Paths
- 12 AI & ML Training
- 1 Blockchain & Decentralized Identity Training
- 32 Cloud & Containers Training
- 3 Cybersecurity Training
- 3 DevOps & Site-Reliability Training
- 1 Linux Kernel Development Training
- 2 Networking Training
- 2 Open Source Best Practice Training
- 5 System Administration Training
- 1 System Engineering Training
- 6 Web & Application Development Training
- 798 Hardware
- 202 Drivers
- 68 I/O Devices
- 37 Monitors
- 96 Multimedia
- 173 Networking
- 91 Printers & Scanners
- 92 Storage
- 772 Linux Distributions
- 81 Debian
- 68 Fedora
- 24 Linux Mint
- 13 Mageia
- 24 openSUSE
- 151 Red Hat Enterprise
- 31 Slackware
- 13 SUSE Enterprise
- 356 Ubuntu
- 469 Linux System Administration
- 31 Cloud Computing
- 73 Command Line/Scripting
- Github systems admin projects
- 101 Linux Security
- 79 Network Management
- 101 System Management
- 46 Web Management
- 166 Mobile Computing
- 32 Android
- 119 Development
- 1.2K New to Linux
- 1K Getting Started with Linux
- 407 Off Topic
- 128 Introductions
- 35 Study Material
- 1K Programming and Development
- 310 Kernel Development
- 720 Software Development
- 1K Software
- 418 Applications
- 182 Command Line
- 5 Compiling/Installing
- 71 Games
- 320 Installation
- Archived
- 183 Small Talk
- 2 LFD140 Class Forum
- 1.4K LFS258 Class Forum
Upcoming Training
-
August 20, 2018
Kubernetes Administration (LFS458)
-
August 20, 2018
Linux System Administration (LFS301)
-
August 27, 2018
Open Source Virtualization (LFS462)
-
August 27, 2018
Linux Kernel Debugging and Security (LFD440)