KubeRusted
Spinning up your cluster…

k9s Lab — hands-on Kubernetes troubleshooting

A simulated Kubernetes cluster in a k9s-style terminal with 18 hands-on missions based on the official docs: CrashLoopBackOff, OOMKilled, Pending pods, readiness probes, Service selectors, HPA autoscaling, CronJobs, rollbacks and node drains.

Look around

Open the pods view, then widen it to every namespace.

Everything that runs in Kubernetes runs in a Pod — even the control plane itself, over in kube-system. Namespaces are just folders for names: 0 shows all of them at once, 1/2/3 jump to your favourites.

Based on the Kubernetes docs: Pods

Meet the control plane

Find the etcd pod in kube-system and describe it.

Look at "Controlled By: Node/control-plane" and the config.mirror annotation: etcd, the API server, scheduler and controller-manager are static pods. The kubelet runs them straight from files in /etc/kubernetes/manifests — try deleting one and watch it reappear instantly, because the API object is only a mirror.

Based on the Kubernetes docs: Create static Pods

Self-healing

Delete one of the web pods in default and watch a replacement come up.

The ReplicaSet controller (inside kube-controller-manager) saw 2 of 3 desired pods, created a new one, the scheduler bound it to a node, and that node's kubelet started it. Nobody "restarted" your pod — a brand-new one replaced it, with a new name and a new IP. That's why you never talk to pods directly.

Based on the Kubernetes docs: ReplicaSet

Pets vs. cattle

Now delete the debug-shell pod in default. Does it come back?

It's gone for good. debug-shell was created with a plain kubectl run — no ReplicaSet, no Deployment, "Controlled By: <none>". Nothing is watching it, so nothing recreates it. In production, always wrap pods in a controller.

Based on the Kubernetes docs: Pods — working with Pods

Scale out

Scale the web Deployment to 5 replicas and wait for all of them to be Running.

You only changed one number (spec.replicas). The Deployment updated its ReplicaSet, the ReplicaSet created pods, and the scheduler spread them across the least-busy workers — never onto the control-plane node, because of its NoSchedule taint. Check :events to see each step.

Based on the Kubernetes docs: Deployments — scaling a Deployment

Why is it crashing?

The payments pod in shop is in CrashLoopBackOff. Open its logs and find out why.

CrashLoopBackOff means the container starts and then exits, so the kubelet keeps restarting it with a growing delay (watch RESTARTS climb). The logs tell you why: v2.3.0 needs a DATABASE_URL variable nobody set. In the logs view, p shows the previous container's logs — kubectl logs --previous. Two fixes: roll back (next mission), or e on the deployment and add a DATABASE_URL entry under the container's env:.

Based on the Kubernetes docs: Debug Pods

Roll it back

Roll payments back to the last good image, shop/payments:2.2.1.

Changing the image changed the pod template, so the Deployment rolled out a new revision — and because 2.2.1's ReplicaSet was still around (scaled to 0), it simply scaled that one back up. Look at :rs: old ReplicaSets are how kubectl rollout undo works.

Based on the Kubernetes docs: Deployments — rolling back

Typo in production

The frontend pods in shop are stuck in ImagePullBackOff. Describe one, spot the problem in its Events, and fix the Deployment's image.

ImagePullBackOff is the kubelet telling you the container runtime could not pull the image — here a typo'd tag (alpnie). The Events at the bottom of describe are almost always the fastest route to the cause. Fixing the Deployment rolled out new pods; the broken ones were removed straight away since they were never Ready.

Based on the Kubernetes docs: Images

Service discovery

Open a shell in any web pod and curl catalog.shop:8080.

catalog.shop was resolved by CoreDNS (see cat /etc/resolv.conf — nameserver 10.96.0.10 is the kube-dns Service) to the Service's stable ClusterIP, and kube-proxy's iptables rules forwarded it to one of the Running catalog pods. Describe the catalog Service to see those endpoints — try curl payments.shop:8080 while payments is crashing, too.

Based on the Kubernetes docs: DNS for Services and Pods

Node maintenance

Drain worker-2 for maintenance, watch its pods move, then uncordon it.

Drain = cordon (mark unschedulable) + evict. Controller-owned pods were recreated on the other workers; the kube-proxy DaemonSet pod stayed because it tolerates the cordon. Uncordoning does not move anything back — the scheduler only places new pods.

Based on the Kubernetes docs: Safely Drain a Node

The Service that routes nowhere

curl orders.shop:8080 is refused although both orders pods are Running. Follow the docs' "Debug Services" checklist: describe the Service, compare its selector with the pods' labels, and fix the selector in its YAML, just like kubectl edit.

A Service finds its pods purely by label selector — app=order matched nothing, so it had no endpoints and kube-proxy had nowhere to send traffic. "Does the Service have any Endpoints?" is the first question in the docs' Debug Services guide for exactly this reason. Now describe it again: the endpoints are the pod IPs, kept up to date by the EndpointSlice controller.

Based on the Kubernetes docs: Debug Services

Running but not Ready

The checkout pods are Running yet show READY 0/1, so the checkout Service gets no traffic. Find out what the readiness probe is checking (describe + logs) and point it at the path the app really serves.

A readiness probe decides whether a pod receives Service traffic — it never restarts anything (that's the liveness probe's job). The probe hit /healthz, the app only answers on /ready, so every check returned 404 and the pods were kept out of the endpoints. Changing the probe changed the pod template, so the Deployment rolled out new pods, and the Service picked them up as soon as they passed.

Based on the Kubernetes docs: Configure Liveness, Readiness and Startup Probes

OOMKilled

The reports pod in shop keeps restarting. Describe it — what does Last State say? — and give the container enough memory.

Last State: Terminated, Reason: OOMKilled, Exit Code: 137 (128 + SIGKILL). A container that goes over its memory limit is killed by the kernel — Kubernetes doesn't throttle memory the way it throttles CPU. The app peaks around 200Mi, the limit was 128Mi. Raising the limit rolled out a new ReplicaSet whose pod finally stays up.

Based on the Kubernetes docs: Assign Memory Resources to Containers and Pods

Too big to schedule

The ml-train pod in default has been Pending since it was created. Read the scheduler's reason in its Events, then fix the Deployment so it fits on a node.

The scheduler places pods by their *requests*, not their actual usage: 0/4 nodes are available: … 3 Insufficient cpu — the pod asked for 6000m (6 CPUs) and every worker only has 4. A request bigger than any node can offer stays Pending forever; lower it (or add bigger nodes) and it is scheduled at once. Describe a node to see its "Allocated resources".

Based on the Kubernetes docs: Assign CPU Resources to Containers and Pods

Assign pods to nodes

The ssd-cache pod only runs on nodes labelled disktype=ssd — and none are. worker-3 has the SSD: label it so the scheduler can place the pod.

A nodeSelector is the simplest node-placement constraint: the pod only fits nodes carrying every listed label. This is the exact kubectl label nodes <node> disktype=ssd walk-through from the docs. Note the scheduler did it on its own the moment a node matched — you never told it where to go. For richer rules (preferences, "not on"), use node affinity.

Based on the Kubernetes docs: Assign Pods to Nodes

Autoscale under load

The catalog HPA keeps CPU at 50% of requests. Open a shell in a web pod, start the docs' load generator against catalog, and watch :hpa scale it to 4+ ready replicas.

The HPA controller read CPU from the metrics pipeline, compared it with the target and applied desiredReplicas = ceil(currentReplicas × currentUtilization / target), capped at maxReplicas. Utilization is a percentage of the CPU *request*, which is why an HPA can't work on pods with no requests. When the load stops, it waits out a stabilization window before scaling back down — watch it in :hpa over the next minute or two.

Based on the Kubernetes docs: HorizontalPodAutoscaler Walkthrough

Run a CronJob now

Don't wait until 03:00: trigger the nightly-report CronJob by hand, then follow the Job it creates until its pod is Completed.

A CronJob is just a Job factory on a schedule — t did kubectl create job --from=cronjob/nightly-report. The Job controller created a pod with restartPolicy OnFailure, and once it exited 0 the Job counted 1 Succeeded. The pod stays around as Completed so you can still read its logs; history limits (and ttlSecondsAfterFinished) clean them up eventually.

Based on the Kubernetes docs: Running Automated Tasks with a CronJob

Take a pod out of rotation

One web pod in default is acting up and you want to debug it without it serving traffic or being killed. Change its app label (e.g. to web-debug) so it leaves both its ReplicaSet and the web Service — while web keeps 3 ready replicas.

Controllers and Services find pods only by label. With app changed, the ReplicaSet no longer counts the pod as one of its own — it releases it (no owner any more) and creates a replacement to get back to 3 — and the Service drops it from its endpoints. The pod keeps running untouched, so you can exec into it and read its logs at leisure; the docs call this isolating Pods from a ReplicaSet. Delete it yourself when you're done — nothing else will.

Based on the Kubernetes docs: ReplicaSet — Isolating Pods from a ReplicaSet