BEEP. BEEP. BEEP.
3:14 AM. That was the sound of my career redlining. I stared at the PagerDuty alert on my nightstand, the blue light searing my retinas. “CRITICAL: API Gateway Latency > 10s (99th Percentile).” By 3:20 AM, the latency didn’t matter because the gateway was gone. By 3:45 AM, the entire $50,000-a-month production cluster—running on Kubernetes v1.27—was a smoking crater.
I sat there, caffeine-deprived and shaking, watching the terminal output as the nodes dropped like flies. I had built this. I had used the “standard” configurations. I had trusted the defaults. And that trust cost us three hours of total downtime and a mid-five-figure cloud bill for a cluster that was doing nothing but rebooting in a frantic, recursive loop of death.
If you are reading this, you are my successor. You’ve inherited the wreckage. Don’t look at the documentation I wrote last month; it’s a lie of optimism. Read this instead. This is how we actually died.
Table of Contents
The Incident: The Cascading Collapse
It started with a minor traffic spike—nothing we hadn’t seen before. But because I relied on “magic” orchestration to handle the load, the cluster began to cannibalize itself. First, the memory-hungry Java microservices hit their limits. Instead of failing gracefully, they triggered a node-level memory pressure event. The Kubelet, trying to save the node, started evicting critical components.
I ran kubectl get events --sort-by='.lastTimestamp' and saw a wall of red:
LAST SEEN TYPE REASON OBJECT MESSAGE
14m Warning BackOff pod/payment-processor-7f8d9b6c5-4x2z1 Back-off restarting failed container
12m Warning Unhealthy pod/api-gateway-7f8d9b6c5-4x2z1 Liveness probe failed: HTTP probe failed with statuscode: 503
10m Warning Evicted pod/auth-service-66b8d5f8-m9q2p The node was low on resource: memory.
8m Normal NodeNotReady node/ip-10-0-64-122.ec2.internal Node ip-10-0-64-122.ec2.internal status is now: NodeNotReady
5m Warning FailedScheduling pod/payment-processor-7f8d9b6c5-99zll 0/12 nodes are available: 4 node(s) were NotReady, 8 Insufficient memory.
Then I checked the node health:
$ kubectl top nodes
NAME CPU(cores) CPU% MEMORY(bytes) MEMORY%
ip-10-0-64-122.ec2.internal 7800m 97% 31200Mi 98%
ip-10-0-64-125.ec2.internal 7950m 99% 31500Mi 99%
ip-10-0-65-10.ec2.internal 150m 2% 1200Mi 4%
The scheduler was trying to cram evicted pods onto the only “healthy” node left, which immediately caused that node to catch fire. It was a circular firing squad.
The Resource Limit Lie
My first sin was a lack of precision. I thought setting limits was enough. I was wrong. In Kubernetes, if you don’t set requests equal to your limits, you are gambling with the scheduler’s soul.
I had pods with limits: memory: 2Gi but no requests. The scheduler saw a pod with no requests and thought, “Great, this pod costs nothing!” and shoved it onto a node that was already at 80% capacity. When the pod actually tried to use that 2Gi of RAM, the node ran out of physical memory. The Linux kernel OOM (Out Of Memory) killer didn’t care about my “orchestration.” It started killing processes at random to save the kernel.
The kubernetes best practice I ignored was the “Guaranteed” Quality of Service (QoS) class. If you want a pod to stay alive, its requests must match its limits exactly.
Here is the YAML I should have used for our payment-processor:
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-processor
namespace: prod
spec:
replicas: 3
template:
spec:
containers:
- name: app
image: our-registry.io/payment:v1.4.2 # NOT LATEST
resources:
requests:
cpu: "1000m"
memory: "2Gi"
limits:
cpu: "1000m"
memory: "2Gi"
By setting them equally, the scheduler knows exactly how much room is left on a node. It won’t oversubscribe. Don’t just set 512Mi and pray. Profile your app. If it uses 1.2Gi at peak, set the request to 1.2Gi. The “Burstable” QoS class is a trap for production databases and core gateways.
Probes That Kill
I thought I was being smart with Liveness and Readiness probes. I wasn’t. I created a “Death Loop.”
Our api-gateway had a liveness probe that checked a /health endpoint. That endpoint, in its infinite wisdom, checked the connection to the backend database. When the database got slow due to the traffic spike, the liveness probe failed. Kubernetes, doing exactly what I told it to do, killed the pod and restarted it.
While the pod was restarting, it couldn’t handle traffic. This put more load on the remaining pods, which then failed their probes, which then restarted. Within minutes, the entire gateway tier was just a collection of containers in a CrashLoopBackOff, not because the code was broken, but because the probes were too aggressive and poorly conceived.
The kubernetes best way to handle this is to decouple your probes. A Liveness probe should only check if the process is deadlocked or crashed. It should never depend on an external dependency like a database. Use a startupProbe for slow-starting apps (like our v1.27 Java monoliths) so the liveness probe doesn’t kill them before they’ve even finished booting.
livenessProbe:
httpGet:
path: /live
port: 8080
initialDelaySeconds: 0
periodSeconds: 10
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready # This checks if the DB is reachable
port: 8080
periodSeconds: 5
startupProbe:
httpGet:
path: /health
port: 8080
failureThreshold: 30 # Gives it 5 minutes to start
periodSeconds: 10
The RBAC Wildcard Nightmare
I got lazy. A senior dev needed to debug a permissions issue in the staging namespace, and I gave them a ClusterRoleBinding to the admin role. “It’s just for an hour,” I said. Six months later, that dev’s credentials were leaked through a compromised CI/CD runner.
Because I used wildcards in our RBAC (Role-Based Access Control), the attacker didn’t just have access to staging. They had the keys to the kingdom. They were able to list all secrets in the kube-system namespace, grab the cloud provider integration keys, and start spinning up GPU instances for crypto mining on our dime.
Stop using * in your YAMLs. The kubernetes best approach is the principle of least privilege. You must define specific verbs for specific resources. If a service account only needs to update a configmap, don’t give it get, list, watch, create, update, patch, delete on *.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
namespace: prod
name: config-updater
rules:
- apiGroups: [""]
resources: ["configmaps"]
resourceNames: ["app-config"]
verbs: ["get", "update", "patch"]
I failed because I valued “developer velocity” over the basic security of our infrastructure. Now, the velocity is zero because the cluster is gone.
The NetworkPolicy Void
By default, Kubernetes allows every pod to talk to every other pod. I knew this. I ignored it. I thought our “internal” network was safe.
When one of our public-facing Nginx pods was exploited via a known CVE, the attacker didn’t stop there. They used curl to scan the internal network. They found the unauthenticated Prometheus metrics endpoint, then the ElasticSearch cluster, and finally the internal metadata service of the cloud provider. They moved laterally across the cluster like a virus because I hadn’t implemented a single NetworkPolicy.
To prevent this, you must implement a default-deny policy. This is the kubernetes best practice for any production-grade cluster. You shut everything down and then surgically open only the paths that are required.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: prod
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-api-to-db
namespace: prod
spec:
podSelector:
matchLabels:
app: api-gateway
ingress:
- from:
- podSelector:
matchLabels:
app: load-balancer
egress:
- to:
- podSelector:
matchLabels:
app: postgres-db
If I had spent the two hours required to map these flows, the breach would have been contained to a single, isolated container. Instead, I watched our data get exfiltrated over a weekend while I was at a BBQ.
The “Latest” Tag Suicide
I allowed image: latest in our production Helm charts. I deserve to be fired for this alone.
One Tuesday afternoon, a developer pushed a “quick fix” to the main branch. The CI pipeline built the image and tagged it as latest. At the same time, a node in our v1.29 cluster failed and was replaced by the autoscaler. When the new node came up, it pulled the latest image.
The “quick fix” had a breaking change in the database schema logic. Now, half of our pods were running the old code, and the new pods were running the broken code. The data became inconsistent. The logs were a nightmare of “Column not found” errors. Because the tag was just latest, I couldn’t easily roll back. kubectl rollout undo did nothing because the underlying image reference hadn’t changed—it was still just latest.
The kubernetes best way to manage images is to use immutable tags—ideally the Git SHA or a semantic version. Never, under any circumstances, allow a deployment to pull an image without a specific, traceable version.
spec:
containers:
- name: worker
image: our-registry.io/worker:sha-a1b2c3d4 # This is traceable
imagePullPolicy: IfNotPresent
Pod Disruption Budgets: The Node Upgrade Massacre
We decided to upgrade the cluster from v1.28 to v1.29. I figured the cloud provider’s “managed upgrade” would handle it. I clicked the button.
The cloud provider started draining nodes. Because I hadn’t configured Pod Disruption Budgets (PDBs), the drain process took down all replicas of our order-service at the same time. The service went offline. The frontend started throwing 500 errors. The upgrade process didn’t care; it just kept killing pods to clear the nodes.
If you have a service that needs at least two replicas to stay functional, you must tell Kubernetes. This is the kubernetes best way to ensure high availability during maintenance.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: order-service-pdb
namespace: prod
spec:
minAvailable: 2
selector:
matchLabels:
app: order-service
Without this, you are at the mercy of the kubectl drain command, which has no context regarding your business requirements. It only knows how to kill.
Helm Charts and the SecurityContext Neglect
I used third-party Helm charts because they were easy. I didn’t audit them. I didn’t realize that many of these charts run their containers as root by default.
One of our “utility” containers—a simple log forwarder—was running with privileged: true and as the root user. When a vulnerability was found in that log forwarder, the attacker gained root access not just to the container, but to the underlying host node. From there, they escaped the container entirely.
I should have enforced a securityContext at the pod and container level. Many Helm charts don’t even expose these fields in their values.yaml, and I was too lazy to fork them or use Kustomize to patch them.
The kubernetes best practice is to explicitly forbid root execution and restrict volume access.
securityContext:
runAsNonRoot: true
runAsUser: 1000
fsGroup: 2000
containers:
- name: app
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
If a container needs to write something, give it an emptyDir volume. Don’t let it write to the node’s root filesystem. I let them have it all, and they took it.
The Orphaned Volume Sin
We were using PersistentVolumes (PVs) for our legacy stateful services. When I deleted a namespace to “clean things up,” I assumed the storage would be cleaned up too. It wasn’t. Because the reclaimPolicy was set to Retain (the default in many older storage classes), the cloud provider kept charging us for those 10TB disks long after the pods were gone.
I spent $4,000 on “ghost” storage in a single month because I didn’t understand how the PV lifecycle worked. Always check your storage classes. Ensure that Delete is the policy for non-critical, ephemeral state, and that you have a manual process for auditing orphaned disks.
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: fast-deletion
provisioner: kubernetes.io/aws-ebs
reclaimPolicy: Delete # This would have saved my budget
allowVolumeExpansion: true
volumeBindingMode: WaitForFirstConsumer
The HPA Flapping Disaster
I set up Horizontal Pod Autoscalers (HPA) to handle load. I set the target CPU utilization to 50%.
The problem? My app took 3 minutes to start up. When load hit, the HPA saw 80% CPU and spun up 10 new pods. But those pods took 3 minutes to become “Ready.” During those 3 minutes, the existing pods were still overwhelmed, so the HPA spun up another 10 pods.
Once all 20 pods finally became ready, the average CPU dropped to 10%. The HPA then immediately killed 15 pods. Then the load hit the remaining 5 pods, and the cycle started again. This “flapping” caused more instability than the actual traffic spike.
The kubernetes best way to fix this is to use behavior settings in your HPA to slow down the scale-down process (stabilization window).
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-gateway
minReplicas: 5
maxReplicas: 50
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # Wait 5 minutes before scaling down
policies:
- type: Percent
value: 10
periodSeconds: 60
I didn’t do this. I let the cluster flap until the API server itself started timing out from the sheer volume of pod creation and deletion events.
Final Words of a Broken Architect
I’m leaving my laptop on the desk. The password is the name of the dog I haven’t seen in three days because I’ve been living in this terminal.
Don’t trust the “magic.” Kubernetes is not a self-healing system unless you tell it exactly how to heal. It is a complex, distributed machine that will happily grind your budget and your sanity into dust if you give it “default” instructions.
Use the kubernetes best practices I’ve outlined here. Not because they are in a manual, but because they are written in the blood of my failed production environment. Every line of YAML I’ve corrected here represents a specific hour of the night I spent watching a progress bar that never finished.
Fix the resource requests. Lock down the RBAC. Implement the NetworkPolicies. And for the love of everything holy, delete the latest tag from your registry.
Good luck. You’re going to need it. I’m going to go sleep for a week.
Related Articles
Explore more insights and best practices: