Part 2 · 3 chapters · ~25 min

Kubernetes, Properly

Kubernetes as a control loop: the API server, etcd, controllers, scheduler and kubelet, and rollouts and self-healing as reconciliation; services, ingress and the Gateway API, network policies and storage; then requests and limits, probes, the three autoscalers and operators.

4

The control loop: objects, controllers, scheduler, kubelet

desired state and the things that chase it
  1. The control plane: the API server, etcd, the scheduler and the controllers.
  2. Apply stores desired state. Nothing runs yet.
  3. Controllers watch, compare and write more desired state.
  4. The scheduler filters and scores nodes, then binds each pod to one.
  5. The kubelet starts containers, runs probes and reports status.
  6. Self-healing and rollouts are the same reconcile loop.
code
# deployment.yaml: the minimum production-shaped Deployment
apiVersion: apps/v1
kind: Deployment
metadata: { name: payments, namespace: prod, labels: { app: payments } }
spec:
  replicas: 6
  strategy: { rollingUpdate: { maxSurge: 25%, maxUnavailable: 0 } }   # never below capacity
  selector: { matchLabels: { app: payments } }
  template:
    metadata: { labels: { app: payments } }
    spec:
      serviceAccountName: payments                 # → IAM role via IRSA / Pod Identity
      topologySpreadConstraints:
        - { maxSkew: 1, topologyKey: topology.kubernetes.io/zone, whenUnsatisfiable: DoNotSchedule,
            labelSelector: { matchLabels: { app: payments } } }        # spread across zones
      containers:
        - name: api
          image: 111122223333.dkr.ecr.eu-west-1.amazonaws.com/payments@sha256:9f2c…
          ports: [{ containerPort: 3000 }]
          resources: { requests: { cpu: 500m, memory: 768Mi }, limits: { memory: 768Mi } }
          readinessProbe: { httpGet: { path: /ready, port: 3000 }, periodSeconds: 5 }
          livenessProbe:  { httpGet: { path: /live,  port: 3000 }, periodSeconds: 10, failureThreshold: 6 }
          securityContext: { runAsNonRoot: true, readOnlyRootFilesystem: true, allowPrivilegeEscalation: false }
the PodDisruptionBudget
Node upgrades and Karpenter consolidation drain nodes on purpose. A PDB (minAvailable: 80%) tells the cluster how many payments pods may be voluntarily evicted at once. Without one, a node upgrade can take out every replica in a zone together.
KUBERNETES IS A CONTROL LOOP
desired state in the API server, controllers that compare and act, and a deployment rolled out by nothing more than that
swipe the figure sideways, or tap expand for full screen
1/6
control plane
The control plane: the API server (the only component that talks to etcd; everything else talks to it), etcd (the consistent key-value store of all objects), the scheduler (assigns pods to nodes), and the controller manager (the built-in controllers: Deployment, ReplicaSet, Job, Node, endpoints). On EKS and GKE the provider runs all of this.
5

Services, ingress, networking and storage

stable layers over disposable pods
  1. Pod networking: every pod gets a routable IP, with no NAT between pods.
  2. Services: a label selector, a stable virtual IP and DNS name, and load balancing across ready pods only.
  3. Ingress and the Gateway API: host and path rules programmed into a real load balancer.
  4. Network policies: default deny per namespace, then allow by label.
  5. Storage: a claim names a class, the CSI driver provisions a cloud volume, and zonal volumes pin the pod to a zone.
  6. StatefulSets: stable identities and per-pod volumes. The managed service is usually the better choice.
code
# gateway API: external HTTPS to the payments Service
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata: { name: payments, namespace: prod }
spec:
  parentRefs: [{ name: public-gateway, namespace: infra }]
  hostnames: ["api.acme.ng"]
  rules:
    - matches: [{ path: { type: PathPrefix, value: /api/pay } }]
      backendRefs: [{ name: payments, port: 80 }]
---
# network policy: only payments may call the ledger
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: { name: ledger-ingress, namespace: prod }
spec:
  podSelector: { matchLabels: { app: ledger } }
  policyTypes: [Ingress]
  ingress: [{ from: [{ podSelector: { matchLabels: { app: payments } } }], ports: [{ port: 8080 }] }]
SERVICES, INGRESS, NETWORKING AND STORAGE
how traffic finds pods that keep moving, and where data lives when pods are disposable
swipe the figure sideways, or tap expand for full screen
1/6
pod networking
Pod networking: the CNI plugin gives every pod an IP that every other pod can reach without NAT (the flat network model). On EKS the VPC CNI hands out real VPC addresses from the subnet, so pods are first-class in the VPC; on GKE alias IP ranges do the same; Cilium and Calico add network policies and eBPF data paths.
6

Resources, probes, autoscaling and operators

packing, protecting, growing, extending
  1. Requests decide how pods are packed onto nodes, and therefore cost. Set them from p95 usage.
  2. Limits: CPU limits throttle and memory limits kill. QoS class sets the eviction order.
  3. Probes: readiness gates traffic and liveness restarts the container. Never point liveness at a dependency.
  4. HPA scales pods, and KEDA adds queue-based triggers.
  5. Node autoscaling with Karpenter or Cluster Autoscaler. Autopilot hides nodes altogether.
  6. Operators are custom kinds with custom controllers running the same loop.
code
# KEDA: scale statement workers on SQS backlog, to zero when idle
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata: { name: statement-worker, namespace: prod }
spec:
  scaleTargetRef: { name: statement-worker }
  minReplicaCount: 0
  maxReplicaCount: 50
  triggers:
    - type: aws-sqs-queue
      metadata: { queueURL: https://sqs.eu-west-1.amazonaws.com/111122223333/statements,
                  queueLength: "20", awsRegion: eu-west-1 }
      authenticationRef: { name: keda-irsa }
do you need Kubernetes?
Kubernetes pays off when there are many services, many teams, and a platform team to run it. Three services and five engineers are usually better served by ECS, Cloud Run or a PaaS. Kubernetes is a platform for building platforms, and adopting it means adopting that job too.
REQUESTS, LIMITS, AUTOSCALING AND OPERATORS
what the scheduler packs by, what the kernel enforces, the three autoscalers, and controllers you write yourself
swipe the figure sideways, or tap expand for full screen
1/6
requests
Requests: the scheduler places a pod only on a node with that much unreserved CPU and memory. Requests too high waste nodes (a pod requesting 2 CPU and using 0.2 strands 1.8); too low overcommit the node and invite noisy neighbours and evictions. Set them from observed usage (p95 over a week), not guesses.