← Back to Home

☸️ Kubernetes Internals

A senior-engineer guide to the hidden mechanics, fast diagnosis techniques, and deep debugging flows you won't find in most tutorials.

1. Hidden & Unknown Concepts

a) Pod Lifecycle Nuances

Pending vs PodInitializing vs Running

These three states confuse even experienced engineers because they live at different layers of the object model.

💡 Pro Tip: Always cross-check phase against conditions. kubectl get pod NAME -o jsonpath='{.status.conditions}' tells you the real truth (Ready, ContainersReady, Initialized, PodScheduled).

CrashLoopBackOff and exponential backoff

When a container exits, the kubelet restarts it with an exponential backoff: 10s, 20s, 40s, 80s, 160s, capped at 300s (5 minutes). The backoff resets only after the container has run successfully for 10 minutes.

# See the restart timeline and current backoff state
kubectl describe pod NAME | grep -A 8 "State"

# The "Reason" (CrashLoopBackOff) and "Exit Code" live here
kubectl get pod NAME -o jsonpath='{.status.containerStatuses[0].state}'

terminationGracePeriodSeconds gotchas

The default grace period is 30 seconds. On deletion, the kubelet:

  1. Runs any preStop hook (this time counts toward the grace period).
  2. Sends SIGTERM to PID 1 in each container.
  3. Waits the remainder of the grace period, then sends SIGKILL.
⚠️ Common mistakes:
  • Running your app as sh -c "app" — the shell (PID 1) does not forward SIGTERM. Use exec app or a real init system.
  • Setting a huge preStop sleep without increasing terminationGracePeriodSeconds — the pod gets SIGKILL'd mid-hook.
  • Assuming SIGTERM is enough for stateful apps that need to flush buffers — you still must handle it in code.

b) Scheduling Internals

How the scheduler scores nodes

The scheduler runs two phases: filtering (remove infeasible nodes) then scoring (rank the rest 0–100). Default scoring plugins include:

PluginWhat it rewards
NodeResourcesFitNodes with the most free CPU/memory for the pod's requests (LeastAllocated) or, optionally, the tightest fit (MostAllocated)
BalancedAllocationNodes whose CPU-to-memory ratio stays balanced after placement
ImageLocalityNodes that already have the container image pulled
TaintTolerationNodes whose taints the pod tolerates (fewer tolerated taints = higher score)
InterPodAffinityNodes that satisfy pod (anti-)affinity rules
NodeAffinityNodes matching preferred node affinity terms

Weights are configurable in the scheduler profile. To inspect what the scheduler decided:

# Verbose scheduler logs (on a control-plane node, or via kube-scheduler pod)
kubectl logs -n kube-system kube-scheduler-control-plane --v=6 | grep -A 30 "Scheduled"

The hidden spec.schedulerName

Every pod implicitly uses default-scheduler. You can override it to route specific pods to a custom scheduler (batch scheduling, topology-aware, GPU-aware). The default scheduler ignores pods whose schedulerName differs from its own.

spec:
  schedulerName: my-custom-scheduler   # default-scheduler will skip this pod

PodTopologySpread — when it silently fails

With whenUnsatisfiable: ScheduleAnyway, the scheduler places the pod even if the spread can't be satisfied, so it silently spreads unevenly. With DoNotSchedule, it refuses and the pod stays Pending — the more visible failure.

Priority & preemption

Pods get a priorityClassName. When a high-priority pod can't schedule, the scheduler preempts (evicts) lower-priority pods to make room. Preemption is not graceful — evicted pods return to Pending and restart their lifecycle.

⚠️ Watch for: A "ping-pong" where a high-priority pod preempts a low-priority one, the low one reschedules elsewhere, then preempts another — cascading churn. Audit with kubectl get events --field-selector reason=Preempted.

c) etcd & Control Plane

What lives in etcd vs the API server cache

etcd is the source of truth. The API server maintains an in-memory watch cache to serve kubectl get without hitting etcd every time. This is why:

etcd compaction

etcd keeps a history of every key revision for MVCC. Without compaction, the database grows unbounded and reads slow down. Compaction removes old revisions while keeping the latest state.

# Manual compaction (revision is from: etcdctl endpoint status --write-out=json)
ETCDCTL_API=3 etcdctl compact REVISION
# Defragment to reclaim space
ETCDCTL_API=3 etcdctl defrag
💡 Pro Tip: etcd automatically compacts periodically, but on high-churn clusters (many ConfigMap updates, lease renewals) you may need to tune --etcd-compaction-interval and defrag regularly. Watch the etcd_mvcc_db_total_size_in_bytes metric.

Reading etcd directly

# Via the API server (safe, no etcd certs needed)
kubectl proxy &
curl http://localhost:8001/api/v1/namespaces/default/pods

# Directly via etcdctl (requires client certs, normally on control-plane nodes)
ETCDCTL_API=3 etcdctl \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key \
  get / --prefix --keys-only | head

d) Networking Deep Dive

Pod network namespaces — clearing the myth

A common claim is that a pod has "4 network namespaces". That's inaccurate. A pod has one network namespace, created by the pause (infra) container and shared by all containers in the pod. Containers in a pod share localhost and the same IP because they join this single netns.

What is true: a running pod involves multiple Linux namespaces of different kinds — network, PID, mount, UTS, IPC, and optionally user. Containers can share some (net, IPC, PID via shareProcessNamespace) and isolate others (mount is always isolated per-container).

kube-proxy: iptables vs IPVS

AspectiptablesIPVS
LookupLinear rule chains — O(n)Kernel hash table — O(1)
Load balancingRandom onlyrr, least-conn, source-hash, etc.
At scaleDegrades past ~1000 servicesHandles many more services well
Connection trackingconntrack requiredLighter, still uses conntrack for SNAT

Packet flow: Pod A → Pod B across nodes

  1. Pod A sends to Pod B's IP. Packet leaves Pod A's netns via its veth pair into the node's root netns.
  2. The node's bridge/CNI (e.g. cni0) routes it toward the destination.
  3. For cross-node traffic, the CNI overlay encapsulates the packet (VXLAN/Geneve) or BGP/underlay routes it directly.
  4. The destination node decapsulates, routes to its bridge, and delivers into Pod B's veth.

Why port-forward works but NodePort doesn't

kubectl port-forward proxies through the API server over SPDY directly to the pod — it bypasses Services, kube-proxy, and node firewalls entirely. If port-forward works but NodePort doesn't, suspect:

e) Storage Secrets

CSI volume provisioning and attachment

When you create a PVC, the external-provisioner calls the CSI driver's CreateVolume. Later, when a pod is scheduled to a node, the external-attacher calls ControllerPublishVolume to attach the volume to that node, then the kubelet calls NodeStageVolume/NodePublishVolume to mount it into the pod.

volumeBindingMode: Immediate vs WaitForFirstConsumer

ModeBehaviorUse case
ImmediateBinds PVC→PV and provisions as soon as the PVC is createdSimple, but volume may land in the wrong availability zone
WaitForFirstConsumerWaits until a pod using the PVC is scheduled, then provisions in the pod's zoneTopology-aware, avoids cross-zone mounts

PVC stuck Pending even when StorageClass exists

2. Quick Diagnosis Techniques (60-Second Checks)

a) Node-Level Diagnostics

# Health conditions in a glance
kubectl get nodes

# Full condition detail (Ready, MemoryPressure, DiskPressure, PIDPressure)
kubectl describe node NODE_NAME | sed -n '/Conditions:/,/Addresses:/p'

# kubelet health from the API perspective
kubectl get node NODE_NAME -o jsonpath='{range .status.conditions[*]}{.type}={.status}{"\n"}{end}'

# Top nodes by CPU / memory
kubectl top nodes --sort-by=cpu
kubectl top nodes --sort-by=memory

On the node itself, verify kubelet via systemd (independent of the API server):

systemctl status kubelet --no-pager
journalctl -u kubelet -n 50 --no-pager

b) Pod-Level Diagnostics

The 3 most important fields in kubectl describe pod:

  1. Events — the scheduler/kubelet narrative (why Pending, why evicted).
  2. State — exit code + reason (OOMKilled, Error, CrashLoopBackOff).
  3. Conditions — the Ready/ContainersReady/PodScheduled truth.
# Logs from a previous (crashed) container instance
kubectl logs POD_NAME --previous
kubectl logs POD_NAME -c CONTAINER_NAME -p   # -p is shorthand

# Compact event view, newest last
kubectl get events --sort-by=.lastTimestamp -o custom-columns=TIME:.lastTimestamp,TYPE:.type,REASON:.reason,MSG:.message

# Ephemeral container for immutable/distroless images
kubectl debug -it POD_NAME --image=busybox --target=CONTAINER_NAME

c) Network Diagnostics

# Throwaway pod for connectivity tests (no curl/wget in target)
kubectl run nettest --rm -it --image=busybox --restart=Never -- /bin/sh
  # inside: nc -vz SERVICE_NAME 80 ; wget -qO- http://SERVICE_NAME

# DNS resolution & the ndots:5 gotcha
kubectl exec POD_NAME -- cat /etc/resolv.conf

# Service endpoints vs actual pods
kubectl get endpoints SERVICE_NAME
kubectl get pods -o wide --show-labels | grep APP_LABEL

# Network policies that might be blocking
kubectl get networkpolicy -A
⚠️ The ndots:5 problem: With ndots:5, a lookup of api.internal.example.com (4 dots) first tries the fully-qualified name plus 5 search domains before falling back. This multiplies DNS queries and can time out external calls. Check /etc/resolv.conf and reduce ndots or use fully-qualified names with a trailing dot.

d) Resource & Performance

# CPU/memory usage and throttling hints
kubectl top pods -A --sort-by=memory

# Find pods with no requests/limits (common cause of throttling/OOM)
kubectl get pods -A -o json | jq -r '.items[] | select(.spec.containers[].resources.requests == null) | .metadata.namespace + "/" + .metadata.name'

# Inode usage on a node
kubectl debug node/NODE_NAME -it --image=busybox -- sh -c 'df -i'

For Cgroups v2, CPU throttling appears in cpu.stat as nr_throttled and throttled_usec. On the node:

cat /sys/fs/cgroup/kubepods.slice/.../cpu.stat   # check nr_throttled

API server latency is best read from its own metrics (via Prometheus or kubectl get --raw):

kubectl get --raw /metrics | grep apiserver_request_duration_seconds_sum

3. Deep Debugging Scenarios (Step-by-Step)

Scenario 1: Pod Stuck in Pending

  1. Check scheduler events: kubectl describe pod POD_NAME | sed -n '/Events:/,$p' — look for FailedScheduling, Insufficient cpu/memory, didn't match pod anti-affinity.
  2. Analyze allocatable vs requested: kubectl describe node NODE_NAME | grep -A 12 "Allocated resources"
  3. Check taints/tolerations & selectors: kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints
  4. Check PVC binding (if volumes): kubectl get pvc -A — a Pending PVC with WaitForFirstConsumer blocks scheduling.
  5. Compare requests vs capacity:
kubectl get pods POD_NAME -o json | jq '.spec.containers[] | {name, resources}'
kubectl get nodes -o json | jq '.items[] | {name:.metadata.name, allocatable:.status.allocatable}'

Scenario 2: Pod Stuck in CrashLoopBackOff

  1. Capture exit code & logs: kubectl describe pod POD_NAME | grep -A 5 "State" then kubectl logs POD_NAME --previous.
  2. App vs runtime error: exit code 137 = SIGKILL (often OOM); 1 or 2 = app error; 0 + restart = wrong container contract (e.g., a one-shot command).
  3. Liveness probe failing: if the app is up but restarting, check the probe's path/port: kubectl get pod POD_NAME -o json | jq '.spec.containers[0].livenessProbe'.
  4. Override the command to inspect: kubectl debug -it POD_NAME --image=busybox --target=CONTAINER -- sh (works even if the original is distroless).
  5. Check OOM kills: kubectl get events | grep -i oom and kubectl describe pod POD_NAME | grep -i "OOMKilled".

Scenario 3: Service Unreachable

  1. Check endpoints: kubectl get endpoints SERVICE_NAME — if ENDPOINTS is empty, the selector matches no pods.
  2. Verify labels match selector: kubectl get svc SERVICE_NAME -o jsonpath='{.spec.selector}' vs kubectl get pods --show-labels.
  3. Test pod-to-pod directly: from a debug pod, hit the pod IP (bypasses the Service).
  4. Check kube-proxy rules: on the node, iptables-save | grep SERVICE_NAME (iptables) or ipvsadm -ln (IPVS).
  5. Verify no NetworkPolicy blocks: kubectl get networkpolicy -A and check ingress rules for the target pods.

Scenario 4: DNS Resolution Failing

  1. Check CoreDNS pods: kubectl get pods -n kube-system -l k8s-app=kube-dns — ensure they're Running.
  2. Test DNS from a busybox pod: kubectl run dnstest --rm -it --image=busybox -- nslookup kubernetes.default.svc.cluster.local.
  3. Inspect /etc/resolv.conf: kubectl exec POD_NAME -- cat /etc/resolv.conf — check nameserver IP and ndots.
  4. Verify the DNS service endpoints: kubectl get endpoints -n kube-system kube-dns.
  5. Debug CoreDNS logs & ConfigMap: kubectl logs -n kube-system -l k8s-app=kube-dns and kubectl get cm -n kube-system coredns -o yaml.

Scenario 5: PVC Stuck in Pending

  1. Check StorageClass & provisioner: kubectl describe pvc PVC_NAME then kubectl get storageclass.
  2. Verify CSI driver status: kubectl get pods -n kube-system | grep csi — the controller pod must be Running.
  3. Check node affinity/zone constraints: a WaitForFirstConsumer PVC with a zone-specific topology that no node satisfies will wait forever.
  4. Validate cloud permissions: check the provisioner's IAM role/credentials for CreateVolume.
  5. Look for leftover volumes: orphaned PVs or a full quota can block provisioning — kubectl get pv | grep -i failed.

4. Production Essentials (Hidden Gems)

a) Monitoring & Alerting

Top 10 kube-state-metrics to track:

  1. kube_pod_status_phase
  2. kube_pod_container_status_restarts_total
  3. kube_deployment_status_replicas
  4. kube_node_status_condition
  5. kube_pod_container_resource_requests
  6. kube_pod_container_resource_limits
  7. kube_persistentvolume_status_phase
  8. kube_persistentvolumeclaim_resource_requests_storage_bytes
  9. kube_job_status_failed
  10. kube_hpa_status_condition
💡 Pro Tip: Alert on kubelet_volume_stats_used_bytes against the PVC capacity to catch PVs filling up before the pod's write starts failing. A 85% threshold with 15m persistence is a good starting point.

Detect CrashLoopBackOff early with increase(kube_pod_container_status_restarts_total[10m]) > 0.

b) Performance Tuning

c) Backup & Recovery

# Snapshot etcd
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%F).db \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

# Verify the snapshot
ETCDCTL_API=3 etcdctl snapshot status /backup/etcd-DATE.db

Restoring a single resource from a snapshot is not a direct operation — you restore the snapshot to a scratch etcd and extract the object, or use a tool like etcd-snapshot-restore to a temp cluster. In practice, back up at the object level too (e.g. Velero) for single-resource restore.

⚠️ "Unknown" pod states: After a node failure, pods can be stuck in Terminating/Unknown. Force-delete only after confirming the node is truly gone: kubectl delete pod POD_NAME --force --grace-period=0.

5. Quick Reference Commands

Aliases

alias k=kubectl
alias kg='kubectl get'
alias kgp='kubectl get pods'
alias kd='kubectl describe'
alias kl='kubectl logs'
alias kex='kubectl exec -it'
alias kns='kubectl config set-context --current --namespace'

One-liners for common problems

# Delete pods stuck in Terminating
kubectl delete pod POD_NAME --force --grace-period=0

# All pods sorted by restart count (descending)
kubectl get pods -A --sort-by=.status.containerStatuses[0].restartCount | sort -k4 -n -r

# Pods without resource requests/limits
kubectl get pods -A -o json | jq -r '.items[] | select(any(.spec.containers[]; (.resources.requests == null) or (.resources.limits == null))) | .metadata.namespace + "/" + .metadata.name'

# PVs by status with age
kubectl get pv --sort-by=.metadata.creationTimestamp -o custom-columns=NAME:.metadata.name,STATUS:.status.phase,AGE:.metadata.creationTimestamp

# Node conditions as a table
kubectl get nodes -o custom-columns=NAME:.metadata.name,READY:.status.conditions[?(@.type==\"Ready\")].status,MEM:.status.conditions[?(@.type==\"MemoryPressure\")].status,DISK:.status.conditions[?(@.type==\"DiskPressure\")].status,PID:.status.conditions[?(@.type==\"PIDPressure\")].status

6. Real-World War Stories

Story 1: The "healthy" pod that wouldn't shut down

A stateful service routinely lost writes on every rollout. Logs showed the app never flushed. The pod looked healthy (Running, ready), so the team blamed the DB. The real issue: the container's entrypoint was sh -c "java -jar app.jar". The shell was PID 1 and never forwarded SIGTERM, so kubelet waited the full 30s then SIGKILL'd — losing buffered data. Aha moment: kubectl logs --previous showed no graceful-shutdown line at all. Fix: exec java -jar app.jar.

Story 2: DNS that worked 80% of the time

External API calls intermittently timed out, but only from certain pods. Nothing in the app code pointed to DNS. The "aha" came from cat /etc/resolv.conf: ndots:5 meant every external hostname first tried 5 search suffixes before resolving, multiplying lookups 6x and occasionally blowing past the 5s resolver timeout. Fix: use fully-qualified names (trailing dot) or lower ndots.

Story 3: A PVC stuck Pending for hours

A PVC with a working StorageClass sat in Pending. The provisioner was Running, IAM was fine, and the events were silent. The "aha": the StorageClass had volumeBindingMode: WaitForFirstConsumer, but the pod that used the PVC had a node affinity to a zone with no available capacity. The scheduler never placed the pod, so the volume was never provisioned — a clean but invisible deadlock. Fix: relax the affinity or add capacity to that zone.

7. Hands-On Lab

Deploy this manifest in minikube or k3d. It contains 3 deliberate issues. Debug and fix each one.

# kubernetes-lab.yaml
# Issue 1: Pending — requests more CPU than any single node can provide
apiVersion: v1
kind: Pod
metadata:
  name: lab-pending
spec:
  containers:
  - name: app
    image: nginx:alpine
    resources:
      requests:
        cpu: "100"          # <-- unrealistic request
        memory: "64Mi"
---
# Issue 2: CrashLoopBackOff — liveness probe hits a path nginx 404s
apiVersion: v1
kind: Pod
metadata:
  name: lab-crash
  labels:
    app: web
spec:
  containers:
  - name: app
    image: nginx:alpine
    ports:
    - containerPort: 80
    livenessProbe:
      httpGet:
        path: /healthz       # <-- nginx has no /healthz
        port: 80
      initialDelaySeconds: 5
      periodSeconds: 5
---
# Issue 3: Service with no endpoints — selector doesn't match pod labels
apiVersion: v1
kind: Service
metadata:
  name: lab-web
spec:
  selector:
    app: web-app            # <-- pods use app: web
  ports:
  - port: 80
    targetPort: 80

Walkthrough

  1. Deploy: kubectl apply -f kubernetes-lab.yaml
  2. See the two broken pods: kubectl get pods → lab-pending (Pending), lab-crash (CrashLoopBackOff).
  3. Fix Issue 1: kubectl describe pod lab-pending shows Insufficient cpu. Change cpu: "100" to cpu: "100m" and re-apply.
  4. Fix Issue 2: kubectl describe pod lab-crash + kubectl logs lab-crash --previous shows the liveness probe failing on /healthz. Change the path to / (or remove the probe) and re-apply.
  5. Fix Issue 3: kubectl get endpoints lab-web shows empty endpoints. Compare the Service selector (app: web-app) with the pod labels (app: web) via kubectl get pods --show-labels. Change the selector to app: web and re-apply, then confirm kubectl get endpoints lab-web.
💡 Pro Tip: Don't fix all three at once — reproduce, diagnose, then fix one at a time. The whole point is training the instinct to read events, logs, and endpoints first.
Copied to clipboard!