A senior-engineer guide to the hidden mechanics, fast diagnosis techniques, and deep debugging flows you won't find in most tutorials.
# Health conditions in a glance
kubectl get nodes
# Full condition detail (Ready, MemoryPressure, DiskPressure, PIDPressure)
kubectl describe node NODE_NAME | sed -n '/Conditions:/,/Addresses:/p'
# kubelet health from the API perspective
kubectl get node NODE_NAME -o jsonpath='{range .status.conditions[*]}{.type}={.status}{"\n"}{end}'
# Top nodes by CPU / memory
kubectl top nodes --sort-by=cpu
kubectl top nodes --sort-by=memory
On the node itself, verify kubelet via systemd (independent of the API server):
systemctl status kubelet --no-pager
journalctl -u kubelet -n 50 --no-pager
The 3 most important fields in kubectl describe pod:
# Logs from a previous (crashed) container instance
kubectl logs POD_NAME --previous
kubectl logs POD_NAME -c CONTAINER_NAME -p # -p is shorthand
# Compact event view, newest last
kubectl get events --sort-by=.lastTimestamp -o custom-columns=TIME:.lastTimestamp,TYPE:.type,REASON:.reason,MSG:.message
# Ephemeral container for immutable/distroless images
kubectl debug -it POD_NAME --image=busybox --target=CONTAINER_NAME
# Throwaway pod for connectivity tests (no curl/wget in target)
kubectl run nettest --rm -it --image=busybox --restart=Never -- /bin/sh
# inside: nc -vz SERVICE_NAME 80 ; wget -qO- http://SERVICE_NAME
# DNS resolution & the ndots:5 gotcha
kubectl exec POD_NAME -- cat /etc/resolv.conf
# Service endpoints vs actual pods
kubectl get endpoints SERVICE_NAME
kubectl get pods -o wide --show-labels | grep APP_LABEL
# Network policies that might be blocking
kubectl get networkpolicy -A
ndots:5 problem: With ndots:5, a lookup of api.internal.example.com (4 dots) first tries the fully-qualified name plus 5 search domains before falling back. This multiplies DNS queries and can time out external calls. Check /etc/resolv.conf and reduce ndots or use fully-qualified names with a trailing dot.# CPU/memory usage and throttling hints
kubectl top pods -A --sort-by=memory
# Find pods with no requests/limits (common cause of throttling/OOM)
kubectl get pods -A -o json | jq -r '.items[] | select(.spec.containers[].resources.requests == null) | .metadata.namespace + "/" + .metadata.name'
# Inode usage on a node
kubectl debug node/NODE_NAME -it --image=busybox -- sh -c 'df -i'
For Cgroups v2, CPU throttling appears in cpu.stat as nr_throttled and throttled_usec. On the node:
cat /sys/fs/cgroup/kubepods.slice/.../cpu.stat # check nr_throttled
API server latency is best read from its own metrics (via Prometheus or kubectl get --raw):
kubectl get --raw /metrics | grep apiserver_request_duration_seconds_sum
kubectl describe pod POD_NAME | sed -n '/Events:/,$p' — look for FailedScheduling, Insufficient cpu/memory, didn't match pod anti-affinity.kubectl describe node NODE_NAME | grep -A 12 "Allocated resources"kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taintskubectl get pvc -A — a Pending PVC with WaitForFirstConsumer blocks scheduling.kubectl get pods POD_NAME -o json | jq '.spec.containers[] | {name, resources}'
kubectl get nodes -o json | jq '.items[] | {name:.metadata.name, allocatable:.status.allocatable}'
kubectl describe pod POD_NAME | grep -A 5 "State" then kubectl logs POD_NAME --previous.137 = SIGKILL (often OOM); 1 or 2 = app error; 0 + restart = wrong container contract (e.g., a one-shot command).kubectl get pod POD_NAME -o json | jq '.spec.containers[0].livenessProbe'.kubectl debug -it POD_NAME --image=busybox --target=CONTAINER -- sh (works even if the original is distroless).kubectl get events | grep -i oom and kubectl describe pod POD_NAME | grep -i "OOMKilled".kubectl get endpoints SERVICE_NAME — if ENDPOINTS is empty, the selector matches no pods.kubectl get svc SERVICE_NAME -o jsonpath='{.spec.selector}' vs kubectl get pods --show-labels.iptables-save | grep SERVICE_NAME (iptables) or ipvsadm -ln (IPVS).kubectl get networkpolicy -A and check ingress rules for the target pods.kubectl get pods -n kube-system -l k8s-app=kube-dns — ensure they're Running.kubectl run dnstest --rm -it --image=busybox -- nslookup kubernetes.default.svc.cluster.local./etc/resolv.conf: kubectl exec POD_NAME -- cat /etc/resolv.conf — check nameserver IP and ndots.kubectl get endpoints -n kube-system kube-dns.kubectl logs -n kube-system -l k8s-app=kube-dns and kubectl get cm -n kube-system coredns -o yaml.kubectl describe pvc PVC_NAME then kubectl get storageclass.kubectl get pods -n kube-system | grep csi — the controller pod must be Running.WaitForFirstConsumer PVC with a zone-specific topology that no node satisfies will wait forever.CreateVolume.kubectl get pv | grep -i failed.Top 10 kube-state-metrics to track:
kube_pod_status_phasekube_pod_container_status_restarts_totalkube_deployment_status_replicaskube_node_status_conditionkube_pod_container_resource_requestskube_pod_container_resource_limitskube_persistentvolume_status_phasekube_persistentvolumeclaim_resource_requests_storage_byteskube_job_status_failedkube_hpa_status_conditionkubelet_volume_stats_used_bytes against the PVC capacity to catch PVs filling up before the pod's write starts failing. A 85% threshold with 15m persistence is a good starting point.Detect CrashLoopBackOff early with increase(kube_pod_container_status_restarts_total[10m]) > 0.
kubelet --max-pods: raising it allows more pods per node, but each pod consumes a PID, IP, and memory overhead. It can exhaust node IPs and the conntrack table — the hidden trade-off.kube-api-server --max-requests-inflight: limits concurrent non-mutating requests (default 400). Too high and the API server can OOM under a thundering herd; too low and clients see 429s.memory.available<100Mi, nodefs.available<10%. Tune these with care — too aggressive and you evict healthy pods.# Snapshot etcd
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%F).db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
# Verify the snapshot
ETCDCTL_API=3 etcdctl snapshot status /backup/etcd-DATE.db
Restoring a single resource from a snapshot is not a direct operation — you restore the snapshot to a scratch etcd and extract the object, or use a tool like etcd-snapshot-restore to a temp cluster. In practice, back up at the object level too (e.g. Velero) for single-resource restore.
Terminating/Unknown. Force-delete only after confirming the node is truly gone: kubectl delete pod POD_NAME --force --grace-period=0.alias k=kubectl
alias kg='kubectl get'
alias kgp='kubectl get pods'
alias kd='kubectl describe'
alias kl='kubectl logs'
alias kex='kubectl exec -it'
alias kns='kubectl config set-context --current --namespace'
# Delete pods stuck in Terminating
kubectl delete pod POD_NAME --force --grace-period=0
# All pods sorted by restart count (descending)
kubectl get pods -A --sort-by=.status.containerStatuses[0].restartCount | sort -k4 -n -r
# Pods without resource requests/limits
kubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.containers[]; (.resources.requests == null) or (.resources.limits == null))) | .metadata.namespace + "/" + .metadata.name'
# PVs by status with age
kubectl get pv --sort-by=.metadata.creationTimestamp -o custom-columns=NAME:.metadata.name,STATUS:.status.phase,AGE:.metadata.creationTimestamp
# Node conditions as a table
kubectl get nodes -o custom-columns=NAME:.metadata.name,READY:.status.conditions[?(@.type==\"Ready\")].status,MEM:.status.conditions[?(@.type==\"MemoryPressure\")].status,DISK:.status.conditions[?(@.type==\"DiskPressure\")].status,PID:.status.conditions[?(@.type==\"PIDPressure\")].status
A stateful service routinely lost writes on every rollout. Logs showed the app never flushed. The pod looked healthy (Running, ready), so the team blamed the DB. The real issue: the container's entrypoint was sh -c "java -jar app.jar". The shell was PID 1 and never forwarded SIGTERM, so kubelet waited the full 30s then SIGKILL'd — losing buffered data. Aha moment: kubectl logs --previous showed no graceful-shutdown line at all. Fix: exec java -jar app.jar.
External API calls intermittently timed out, but only from certain pods. Nothing in the app code pointed to DNS. The "aha" came from cat /etc/resolv.conf: ndots:5 meant every external hostname first tried 5 search suffixes before resolving, multiplying lookups 6x and occasionally blowing past the 5s resolver timeout. Fix: use fully-qualified names (trailing dot) or lower ndots.
A PVC with a working StorageClass sat in Pending. The provisioner was Running, IAM was fine, and the events were silent. The "aha": the StorageClass had volumeBindingMode: WaitForFirstConsumer, but the pod that used the PVC had a node affinity to a zone with no available capacity. The scheduler never placed the pod, so the volume was never provisioned — a clean but invisible deadlock. Fix: relax the affinity or add capacity to that zone.
Deploy this manifest in minikube or k3d. It contains 3 deliberate issues. Debug and fix each one.
# kubernetes-lab.yaml
# Issue 1: Pending — requests more CPU than any single node can provide
apiVersion: v1
kind: Pod
metadata:
name: lab-pending
spec:
containers:
- name: app
image: nginx:alpine
resources:
requests:
cpu: "100" # <-- unrealistic request
memory: "64Mi"
---
# Issue 2: CrashLoopBackOff — liveness probe hits a path nginx 404s
apiVersion: v1
kind: Pod
metadata:
name: lab-crash
labels:
app: web
spec:
containers:
- name: app
image: nginx:alpine
ports:
- containerPort: 80
livenessProbe:
httpGet:
path: /healthz # <-- nginx has no /healthz
port: 80
initialDelaySeconds: 5
periodSeconds: 5
---
# Issue 3: Service with no endpoints — selector doesn't match pod labels
apiVersion: v1
kind: Service
metadata:
name: lab-web
spec:
selector:
app: web-app # <-- pods use app: web
ports:
- port: 80
targetPort: 80
kubectl apply -f kubernetes-lab.yamlkubectl get pods → lab-pending (Pending), lab-crash (CrashLoopBackOff).kubectl describe pod lab-pending shows Insufficient cpu. Change cpu: "100" to cpu: "100m" and re-apply.kubectl describe pod lab-crash + kubectl logs lab-crash --previous shows the liveness probe failing on /healthz. Change the path to / (or remove the probe) and re-apply.kubectl get endpoints lab-web shows empty endpoints. Compare the Service selector (app: web-app) with the pod labels (app: web) via kubectl get pods --show-labels. Change the selector to app: web and re-apply, then confirm kubectl get endpoints lab-web.