{"id":4902,"date":"2026-10-09T02:08:18","date_gmt":"2026-10-08T20:38:18","guid":{"rendered":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/"},"modified":"2026-10-09T02:08:18","modified_gmt":"2026-10-08T20:38:18","slug":"10-essential-kubernetes-best-practices-for-scalable-apps","status":"publish","type":"post","link":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/","title":{"rendered":"10 Essential Kubernetes Best Practices for Scalable Apps"},"content":{"rendered":"<p><em>BEEP. BEEP. BEEP.<\/em><\/p>\n<p>3:14 AM. That was the sound of my career redlining. I stared at the PagerDuty alert on my nightstand, the blue light searing my retinas. &#8220;CRITICAL: API Gateway Latency &gt; 10s (99th Percentile).&#8221; By 3:20 AM, the latency didn&#8217;t matter because the gateway was gone. By 3:45 AM, the entire $50,000-a-month production cluster\u2014running on Kubernetes v1.27\u2014was a smoking crater. <\/p>\n<p>I sat there, caffeine-deprived and shaking, watching the terminal output as the nodes dropped like flies. I had built this. I had used the &#8220;standard&#8221; configurations. I had trusted the defaults. And that trust cost us three hours of total downtime and a mid-five-figure cloud bill for a cluster that was doing nothing but rebooting in a frantic, recursive loop of death.<\/p>\n<p>If you are reading this, you are my successor. You\u2019ve inherited the wreckage. Don\u2019t look at the documentation I wrote last month; it\u2019s a lie of optimism. Read this instead. This is how we actually died.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_80 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6ac852e6af1c9\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6ac852e6af1c9\"  aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#The_Incident_The_Cascading_Collapse\" >The Incident: The Cascading Collapse<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#The_Resource_Limit_Lie\" >The Resource Limit Lie<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#Probes_That_Kill\" >Probes That Kill<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#The_RBAC_Wildcard_Nightmare\" >The RBAC Wildcard Nightmare<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#The_NetworkPolicy_Void\" >The NetworkPolicy Void<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#The_%E2%80%9CLatest%E2%80%9D_Tag_Suicide\" >The &#8220;Latest&#8221; Tag Suicide<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#Pod_Disruption_Budgets_The_Node_Upgrade_Massacre\" >Pod Disruption Budgets: The Node Upgrade Massacre<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#Helm_Charts_and_the_SecurityContext_Neglect\" >Helm Charts and the SecurityContext Neglect<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#The_Orphaned_Volume_Sin\" >The Orphaned Volume Sin<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#The_HPA_Flapping_Disaster\" >The HPA Flapping Disaster<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#Final_Words_of_a_Broken_Architect\" >Final Words of a Broken Architect<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#Related_Articles\" >Related Articles<\/a><\/li><\/ul><\/nav><\/div>\n<h2><span class=\"ez-toc-section\" id=\"The_Incident_The_Cascading_Collapse\"><\/span>The Incident: The Cascading Collapse<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>It started with a minor traffic spike\u2014nothing we hadn&#8217;t seen before. But because I relied on &#8220;magic&#8221; orchestration to handle the load, the cluster began to cannibalize itself. First, the memory-hungry Java microservices hit their limits. Instead of failing gracefully, they triggered a node-level memory pressure event. The Kubelet, trying to save the node, started evicting critical components.<\/p>\n<p>I ran <code>kubectl get events --sort-by='.lastTimestamp'<\/code> and saw a wall of red:<\/p>\n<pre class=\"codehilite\"><code class=\"language-text\">LAST SEEN   TYPE      REASON             OBJECT                                  MESSAGE\n14m         Warning   BackOff            pod\/payment-processor-7f8d9b6c5-4x2z1   Back-off restarting failed container\n12m         Warning   Unhealthy          pod\/api-gateway-7f8d9b6c5-4x2z1         Liveness probe failed: HTTP probe failed with statuscode: 503\n10m         Warning   Evicted            pod\/auth-service-66b8d5f8-m9q2p         The node was low on resource: memory.\n8m          Normal    NodeNotReady       node\/ip-10-0-64-122.ec2.internal        Node ip-10-0-64-122.ec2.internal status is now: NodeNotReady\n5m          Warning   FailedScheduling   pod\/payment-processor-7f8d9b6c5-99zll   0\/12 nodes are available: 4 node(s) were NotReady, 8 Insufficient memory.\n<\/code><\/pre>\n<p>Then I checked the node health:<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\">$ kubectl top nodes\nNAME                          CPU(cores)   CPU%   MEMORY(bytes)   MEMORY%   \nip-10-0-64-122.ec2.internal   7800m        97%    31200Mi         98%       \nip-10-0-64-125.ec2.internal   7950m        99%    31500Mi         99%       \nip-10-0-65-10.ec2.internal    150m         2%     1200Mi          4%        \n<\/code><\/pre>\n<p>The scheduler was trying to cram evicted pods onto the only &#8220;healthy&#8221; node left, which immediately caused that node to catch fire. It was a circular firing squad.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_Resource_Limit_Lie\"><\/span>The Resource Limit Lie<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>My first sin was a lack of precision. I thought setting <code>limits<\/code> was enough. I was wrong. In Kubernetes, if you don&#8217;t set <code>requests<\/code> equal to your <code>limits<\/code>, you are gambling with the scheduler\u2019s soul. <\/p>\n<p>I had pods with <code>limits: memory: 2Gi<\/code> but no <code>requests<\/code>. The scheduler saw a pod with no requests and thought, &#8220;Great, this pod costs nothing!&#8221; and shoved it onto a node that was already at 80% capacity. When the pod actually tried to use that 2Gi of RAM, the node ran out of physical memory. The Linux kernel OOM (Out Of Memory) killer didn&#8217;t care about my &#8220;orchestration.&#8221; It started killing processes at random to save the kernel.<\/p>\n<p>The <strong>kubernetes best<\/strong> practice I ignored was the &#8220;Guaranteed&#8221; Quality of Service (QoS) class. If you want a pod to stay alive, its requests must match its limits exactly.<\/p>\n<p>Here is the YAML I should have used for our <code>payment-processor<\/code>:<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: apps\/v1\nkind: Deployment\nmetadata:\n  name: payment-processor\n  namespace: prod\nspec:\n  replicas: 3\n  template:\n    spec:\n      containers:\n      - name: app\n        image: our-registry.io\/payment:v1.4.2 # NOT LATEST\n        resources:\n          requests:\n            cpu: &quot;1000m&quot;\n            memory: &quot;2Gi&quot;\n          limits:\n            cpu: &quot;1000m&quot;\n            memory: &quot;2Gi&quot;\n<\/code><\/pre>\n<p>By setting them equally, the scheduler knows exactly how much room is left on a node. It won&#8217;t oversubscribe. Don&#8217;t just set 512Mi and pray. Profile your app. If it uses 1.2Gi at peak, set the request to 1.2Gi. The &#8220;Burstable&#8221; QoS class is a trap for production databases and core gateways.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Probes_That_Kill\"><\/span>Probes That Kill<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I thought I was being smart with Liveness and Readiness probes. I wasn&#8217;t. I created a &#8220;Death Loop.&#8221; <\/p>\n<p>Our <code>api-gateway<\/code> had a liveness probe that checked a <code>\/health<\/code> endpoint. That endpoint, in its infinite wisdom, checked the connection to the backend database. When the database got slow due to the traffic spike, the liveness probe failed. Kubernetes, doing exactly what I told it to do, killed the pod and restarted it.<\/p>\n<p>While the pod was restarting, it couldn&#8217;t handle traffic. This put more load on the remaining pods, which then failed their probes, which then restarted. Within minutes, the entire gateway tier was just a collection of containers in a <code>CrashLoopBackOff<\/code>, not because the code was broken, but because the probes were too aggressive and poorly conceived.<\/p>\n<p>The <strong>kubernetes best<\/strong> way to handle this is to decouple your probes. A Liveness probe should only check if the process is deadlocked or crashed. It should <em>never<\/em> depend on an external dependency like a database. Use a <code>startupProbe<\/code> for slow-starting apps (like our v1.27 Java monoliths) so the liveness probe doesn&#8217;t kill them before they&#8217;ve even finished booting.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">        livenessProbe:\n          httpGet:\n            path: \/live\n            port: 8080\n          initialDelaySeconds: 0\n          periodSeconds: 10\n          failureThreshold: 3\n        readinessProbe:\n          httpGet:\n            path: \/ready # This checks if the DB is reachable\n            port: 8080\n          periodSeconds: 5\n        startupProbe:\n          httpGet:\n            path: \/health\n            port: 8080\n          failureThreshold: 30 # Gives it 5 minutes to start\n          periodSeconds: 10\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"The_RBAC_Wildcard_Nightmare\"><\/span>The RBAC Wildcard Nightmare<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I got lazy. A senior dev needed to debug a permissions issue in the <code>staging<\/code> namespace, and I gave them a ClusterRoleBinding to the <code>admin<\/code> role. &#8220;It&#8217;s just for an hour,&#8221; I said. Six months later, that dev\u2019s credentials were leaked through a compromised CI\/CD runner.<\/p>\n<p>Because I used wildcards in our RBAC (Role-Based Access Control), the attacker didn&#8217;t just have access to <code>staging<\/code>. They had the keys to the kingdom. They were able to list all secrets in the <code>kube-system<\/code> namespace, grab the cloud provider integration keys, and start spinning up GPU instances for crypto mining on our dime.<\/p>\n<p>Stop using <code>*<\/code> in your YAMLs. The <strong>kubernetes best<\/strong> approach is the principle of least privilege. You must define specific verbs for specific resources. If a service account only needs to update a configmap, don&#8217;t give it <code>get, list, watch, create, update, patch, delete<\/code> on <code>*<\/code>.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: rbac.authorization.k8s.io\/v1\nkind: Role\nmetadata:\n  namespace: prod\n  name: config-updater\nrules:\n- apiGroups: [&quot;&quot;]\n  resources: [&quot;configmaps&quot;]\n  resourceNames: [&quot;app-config&quot;]\n  verbs: [&quot;get&quot;, &quot;update&quot;, &quot;patch&quot;]\n<\/code><\/pre>\n<p>I failed because I valued &#8220;developer velocity&#8221; over the basic security of our infrastructure. Now, the velocity is zero because the cluster is gone.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_NetworkPolicy_Void\"><\/span>The NetworkPolicy Void<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>By default, Kubernetes allows every pod to talk to every other pod. I knew this. I ignored it. I thought our &#8220;internal&#8221; network was safe.<\/p>\n<p>When one of our public-facing Nginx pods was exploited via a known CVE, the attacker didn&#8217;t stop there. They used <code>curl<\/code> to scan the internal network. They found the unauthenticated Prometheus metrics endpoint, then the ElasticSearch cluster, and finally the internal metadata service of the cloud provider. They moved laterally across the cluster like a virus because I hadn&#8217;t implemented a single <code>NetworkPolicy<\/code>.<\/p>\n<p>To prevent this, you must implement a default-deny policy. This is the <strong>kubernetes best<\/strong> practice for any production-grade cluster. You shut everything down and then surgically open only the paths that are required.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: networking.k8s.io\/v1\nkind: NetworkPolicy\nmetadata:\n  name: default-deny-all\n  namespace: prod\nspec:\n  podSelector: {}\n  policyTypes:\n  - Ingress\n  - Egress\n---\napiVersion: networking.k8s.io\/v1\nkind: NetworkPolicy\nmetadata:\n  name: allow-api-to-db\n  namespace: prod\nspec:\n  podSelector:\n    matchLabels:\n      app: api-gateway\n  ingress:\n  - from:\n    - podSelector:\n        matchLabels:\n          app: load-balancer\n  egress:\n  - to:\n    - podSelector:\n        matchLabels:\n          app: postgres-db\n<\/code><\/pre>\n<p>If I had spent the two hours required to map these flows, the breach would have been contained to a single, isolated container. Instead, I watched our data get exfiltrated over a weekend while I was at a BBQ.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_%E2%80%9CLatest%E2%80%9D_Tag_Suicide\"><\/span>The &#8220;Latest&#8221; Tag Suicide<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I allowed <code>image: latest<\/code> in our production Helm charts. I deserve to be fired for this alone.<\/p>\n<p>One Tuesday afternoon, a developer pushed a &#8220;quick fix&#8221; to the main branch. The CI pipeline built the image and tagged it as <code>latest<\/code>. At the same time, a node in our v1.29 cluster failed and was replaced by the autoscaler. When the new node came up, it pulled the <code>latest<\/code> image. <\/p>\n<p>The &#8220;quick fix&#8221; had a breaking change in the database schema logic. Now, half of our pods were running the old code, and the new pods were running the broken code. The data became inconsistent. The logs were a nightmare of &#8220;Column not found&#8221; errors. Because the tag was just <code>latest<\/code>, I couldn&#8217;t easily roll back. <code>kubectl rollout undo<\/code> did nothing because the underlying image reference hadn&#8217;t changed\u2014it was still just <code>latest<\/code>.<\/p>\n<p>The <strong>kubernetes best<\/strong> way to manage images is to use immutable tags\u2014ideally the Git SHA or a semantic version. Never, under any circumstances, allow a deployment to pull an image without a specific, traceable version.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">spec:\n  containers:\n  - name: worker\n    image: our-registry.io\/worker:sha-a1b2c3d4 # This is traceable\n    imagePullPolicy: IfNotPresent\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"Pod_Disruption_Budgets_The_Node_Upgrade_Massacre\"><\/span>Pod Disruption Budgets: The Node Upgrade Massacre<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We decided to upgrade the cluster from v1.28 to v1.29. I figured the cloud provider&#8217;s &#8220;managed upgrade&#8221; would handle it. I clicked the button.<\/p>\n<p>The cloud provider started draining nodes. Because I hadn&#8217;t configured Pod Disruption Budgets (PDBs), the drain process took down all replicas of our <code>order-service<\/code> at the same time. The service went offline. The frontend started throwing 500 errors. The upgrade process didn&#8217;t care; it just kept killing pods to clear the nodes.<\/p>\n<p>If you have a service that needs at least two replicas to stay functional, you must tell Kubernetes. This is the <strong>kubernetes best<\/strong> way to ensure high availability during maintenance.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: policy\/v1\nkind: PodDisruptionBudget\nmetadata:\n  name: order-service-pdb\n  namespace: prod\nspec:\n  minAvailable: 2\n  selector:\n    matchLabels:\n      app: order-service\n<\/code><\/pre>\n<p>Without this, you are at the mercy of the <code>kubectl drain<\/code> command, which has no context regarding your business requirements. It only knows how to kill.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Helm_Charts_and_the_SecurityContext_Neglect\"><\/span>Helm Charts and the SecurityContext Neglect<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I used third-party Helm charts because they were easy. I didn&#8217;t audit them. I didn&#8217;t realize that many of these charts run their containers as <code>root<\/code> by default. <\/p>\n<p>One of our &#8220;utility&#8221; containers\u2014a simple log forwarder\u2014was running with <code>privileged: true<\/code> and as the root user. When a vulnerability was found in that log forwarder, the attacker gained root access not just to the container, but to the underlying host node. From there, they escaped the container entirely.<\/p>\n<p>I should have enforced a <code>securityContext<\/code> at the pod and container level. Many Helm charts don&#8217;t even expose these fields in their <code>values.yaml<\/code>, and I was too lazy to fork them or use Kustomize to patch them.<\/p>\n<p>The <strong>kubernetes best<\/strong> practice is to explicitly forbid root execution and restrict volume access.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">      securityContext:\n        runAsNonRoot: true\n        runAsUser: 1000\n        fsGroup: 2000\n      containers:\n      - name: app\n        securityContext:\n          allowPrivilegeEscalation: false\n          readOnlyRootFilesystem: true\n          capabilities:\n            drop:\n              - ALL\n<\/code><\/pre>\n<p>If a container needs to write something, give it an <code>emptyDir<\/code> volume. Don&#8217;t let it write to the node&#8217;s root filesystem. I let them have it all, and they took it.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_Orphaned_Volume_Sin\"><\/span>The Orphaned Volume Sin<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We were using <code>PersistentVolumes<\/code> (PVs) for our legacy stateful services. When I deleted a namespace to &#8220;clean things up,&#8221; I assumed the storage would be cleaned up too. It wasn&#8217;t. Because the <code>reclaimPolicy<\/code> was set to <code>Retain<\/code> (the default in many older storage classes), the cloud provider kept charging us for those 10TB disks long after the pods were gone.<\/p>\n<p>I spent $4,000 on &#8220;ghost&#8221; storage in a single month because I didn&#8217;t understand how the PV lifecycle worked. Always check your storage classes. Ensure that <code>Delete<\/code> is the policy for non-critical, ephemeral state, and that you have a manual process for auditing orphaned disks.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: storage.k8s.io\/v1\nkind: StorageClass\nmetadata:\n  name: fast-deletion\nprovisioner: kubernetes.io\/aws-ebs\nreclaimPolicy: Delete # This would have saved my budget\nallowVolumeExpansion: true\nvolumeBindingMode: WaitForFirstConsumer\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"The_HPA_Flapping_Disaster\"><\/span>The HPA Flapping Disaster<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I set up Horizontal Pod Autoscalers (HPA) to handle load. I set the target CPU utilization to 50%. <\/p>\n<p>The problem? My app took 3 minutes to start up. When load hit, the HPA saw 80% CPU and spun up 10 new pods. But those pods took 3 minutes to become &#8220;Ready.&#8221; During those 3 minutes, the existing pods were still overwhelmed, so the HPA spun up <em>another<\/em> 10 pods. <\/p>\n<p>Once all 20 pods finally became ready, the average CPU dropped to 10%. The HPA then immediately killed 15 pods. Then the load hit the remaining 5 pods, and the cycle started again. This &#8220;flapping&#8221; caused more instability than the actual traffic spike.<\/p>\n<p>The <strong>kubernetes best<\/strong> way to fix this is to use <code>behavior<\/code> settings in your HPA to slow down the scale-down process (stabilization window).<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: autoscaling\/v2\nkind: HorizontalPodAutoscaler\nmetadata:\n  name: api-hpa\nspec:\n  scaleTargetRef:\n    apiVersion: apps\/v1\n    kind: Deployment\n    name: api-gateway\n  minReplicas: 5\n  maxReplicas: 50\n  metrics:\n  - type: Resource\n    resource:\n      name: cpu\n      target:\n        type: Utilization\n        averageUtilization: 70\n  behavior:\n    scaleDown:\n      stabilizationWindowSeconds: 300 # Wait 5 minutes before scaling down\n      policies:\n      - type: Percent\n        value: 10\n        periodSeconds: 60\n<\/code><\/pre>\n<p>I didn&#8217;t do this. I let the cluster flap until the API server itself started timing out from the sheer volume of pod creation and deletion events.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Final_Words_of_a_Broken_Architect\"><\/span>Final Words of a Broken Architect<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I&#8217;m leaving my laptop on the desk. The password is the name of the dog I haven&#8217;t seen in three days because I&#8217;ve been living in this terminal. <\/p>\n<p>Don&#8217;t trust the &#8220;magic.&#8221; Kubernetes is not a self-healing system unless you tell it exactly how to heal. It is a complex, distributed machine that will happily grind your budget and your sanity into dust if you give it &#8220;default&#8221; instructions. <\/p>\n<p>Use the <strong>kubernetes best<\/strong> practices I&#8217;ve outlined here. Not because they are in a manual, but because they are written in the blood of my failed production environment. Every line of YAML I&#8217;ve corrected here represents a specific hour of the night I spent watching a progress bar that never finished.<\/p>\n<p>Fix the resource requests. Lock down the RBAC. Implement the NetworkPolicies. And for the love of everything holy, delete the <code>latest<\/code> tag from your registry.<\/p>\n<p>Good luck. You\u2019re going to need it. I\u2019m going to go sleep for a week.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Related_Articles\"><\/span>Related Articles<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Explore more insights and best practices:<\/p>\n<ul>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/what-is-artificial-intelligence-a-comprehensive-guide\/\">What Is Artificial Intelligence A Comprehensive Guide<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/understanding-machine-learning-models-a-complete-guide-2\/\">Understanding Machine Learning Models A Complete Guide 2<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/react-native-guide-build-powerful-cross-platform-apps\/\">React Native Guide Build Powerful Cross Platform Apps<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>BEEP. BEEP. BEEP. 3:14 AM. That was the sound of my career redlining. I stared at the PagerDuty alert on my nightstand, the blue light searing my retinas. &#8220;CRITICAL: API Gateway Latency &gt; 10s (99th Percentile).&#8221; By 3:20 AM, the latency didn&#8217;t matter because the gateway was gone. By 3:45 AM, the entire $50,000-a-month production &#8230; <a title=\"10 Essential Kubernetes Best Practices for Scalable Apps\" class=\"read-more\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\" aria-label=\"Read more  on 10 Essential Kubernetes Best Practices for Scalable Apps\">Read more<\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-4902","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.0 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>10 Essential Kubernetes Best Practices for Scalable Apps - ITSupportWale<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"10 Essential Kubernetes Best Practices for Scalable Apps - ITSupportWale\" \/>\n<meta property=\"og:description\" content=\"BEEP. BEEP. BEEP. 3:14 AM. That was the sound of my career redlining. I stared at the PagerDuty alert on my nightstand, the blue light searing my retinas. &#8220;CRITICAL: API Gateway Latency &gt; 10s (99th Percentile).&#8221; By 3:20 AM, the latency didn&#8217;t matter because the gateway was gone. By 3:45 AM, the entire $50,000-a-month production ... Read more\" \/>\n<meta property=\"og:url\" content=\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\" \/>\n<meta property=\"og:site_name\" content=\"ITSupportWale\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\" \/>\n<meta property=\"article:published_time\" content=\"2026-10-08T20:38:18+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png\" \/>\n\t<meta property=\"og:image:width\" content=\"512\" \/>\n\t<meta property=\"og:image:height\" content=\"512\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Techie\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Techie\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\"},\"author\":{\"name\":\"Techie\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\"},\"headline\":\"10 Essential Kubernetes Best Practices for Scalable Apps\",\"datePublished\":\"2026-10-08T20:38:18+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\"},\"wordCount\":1943,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\",\"name\":\"10 Essential Kubernetes Best Practices for Scalable Apps - ITSupportWale\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\"},\"datePublished\":\"2026-10-08T20:38:18+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/itsupportwale.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"10 Essential Kubernetes Best Practices for Scalable Apps\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"name\":\"ITSupportWale\",\"description\":\"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides\",\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\",\"name\":\"itsupportwale\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"contentUrl\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"width\":1119,\"height\":144,\"caption\":\"itsupportwale\"},\"image\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\",\"name\":\"Techie\",\"sameAs\":[\"https:\/\/itsupportwale.com\",\"iswblogadmin\"],\"url\":\"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"10 Essential Kubernetes Best Practices for Scalable Apps - ITSupportWale","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/","og_locale":"en_US","og_type":"article","og_title":"10 Essential Kubernetes Best Practices for Scalable Apps - ITSupportWale","og_description":"BEEP. BEEP. BEEP. 3:14 AM. That was the sound of my career redlining. I stared at the PagerDuty alert on my nightstand, the blue light searing my retinas. &#8220;CRITICAL: API Gateway Latency &gt; 10s (99th Percentile).&#8221; By 3:20 AM, the latency didn&#8217;t matter because the gateway was gone. By 3:45 AM, the entire $50,000-a-month production ... Read more","og_url":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/","og_site_name":"ITSupportWale","article_publisher":"https:\/\/www.facebook.com\/Itsupportwale-298547177495978","article_published_time":"2026-10-08T20:38:18+00:00","og_image":[{"width":512,"height":512,"url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png","type":"image\/png"}],"author":"Techie","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Techie","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#article","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/"},"author":{"name":"Techie","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d"},"headline":"10 Essential Kubernetes Best Practices for Scalable Apps","datePublished":"2026-10-08T20:38:18+00:00","mainEntityOfPage":{"@id":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/"},"wordCount":1943,"commentCount":0,"publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/","url":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/","name":"10 Essential Kubernetes Best Practices for Scalable Apps - ITSupportWale","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/#website"},"datePublished":"2026-10-08T20:38:18+00:00","breadcrumb":{"@id":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/itsupportwale.com\/blog\/10-essential-kubernetes-best-practices-for-scalable-apps\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/itsupportwale.com\/blog\/"},{"@type":"ListItem","position":2,"name":"10 Essential Kubernetes Best Practices for Scalable Apps"}]},{"@type":"WebSite","@id":"https:\/\/itsupportwale.com\/blog\/#website","url":"https:\/\/itsupportwale.com\/blog\/","name":"ITSupportWale","description":"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides","publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/itsupportwale.com\/blog\/#organization","name":"itsupportwale","url":"https:\/\/itsupportwale.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","contentUrl":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","width":1119,"height":144,"caption":"itsupportwale"},"image":{"@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/Itsupportwale-298547177495978"]},{"@type":"Person","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d","name":"Techie","sameAs":["https:\/\/itsupportwale.com","iswblogadmin"],"url":"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/"}]}},"_links":{"self":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4902","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/comments?post=4902"}],"version-history":[{"count":0,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4902\/revisions"}],"wp:attachment":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/media?parent=4902"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/categories?post=4902"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/tags?post=4902"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}