{"id":4866,"date":"2026-08-24T21:16:48","date_gmt":"2026-08-24T15:46:48","guid":{"rendered":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/"},"modified":"2026-08-24T21:16:48","modified_gmt":"2026-08-24T15:46:48","slug":"10-kubernetes-best-practices-for-production-success-3","status":"publish","type":"post","link":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/","title":{"rendered":"10 Kubernetes Best Practices for Production Success"},"content":{"rendered":"<p>It is 4:42 AM. The sun isn\u2019t up, but my blood pressure is. I\u2019ve spent the last 48 hours staring at a Grafana dashboard that looked like a heart monitor of a patient in active cardiac arrest. My eyes feel like they\u2019ve been rubbed with sandpaper, and the smell of stale, burnt coffee\u2014the kind that\u2019s been sitting in the pot since Tuesday\u2014is the only thing keeping me tethered to this mortal plane.<\/p>\n<p>You &#8220;clever&#8221; developers finally did it. You didn\u2019t just break the app; you managed to turn a high-availability, multi-zone Kubernetes cluster into a very expensive heater for a data center in Northern Virginia. <\/p>\n<p>Here is the post-mortem. Read it. Print it out. Eat it for all I care. Just stop doing this to me.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_80 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6a8ca38ad330b\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6a8ca38ad330b\"  aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#The_Incident_Death_by_a_Thousand_API_Calls\" >The Incident: Death by a Thousand API Calls<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#1_The_Resource_Limit_Lie_Why_%E2%80%98Best_Effort_is_a_Suicide_Pact\" >1. The Resource Limit Lie: Why &#8216;Best Effort&#8217; is a Suicide Pact<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#2_Probes_that_Kill_How_Your_Liveness_Check_Created_a_Cascading_Failure\" >2. Probes that Kill: How Your Liveness Check Created a Cascading Failure<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#3_RBAC_is_Not_Optional_Stop_Giving_%E2%80%98Cluster-Admin_to_Your_CICD_Bot\" >3. RBAC is Not Optional: Stop Giving &#8216;Cluster-Admin&#8217; to Your CI\/CD Bot<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#4_The_Hidden_Cost_of_%E2%80%9Ckubernetes_best%E2%80%9D_Intentions\" >4. The Hidden Cost of &#8220;kubernetes best&#8221; Intentions<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#5_Networking_Sprawl_When_Your_CNI_Decides_to_Stop_Routing_Traffic\" >5. Networking Sprawl: When Your CNI Decides to Stop Routing Traffic<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#6_Storage_Classes_and_the_Myth_of_Statelessness\" >6. Storage Classes and the Myth of Statelessness<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#The_%E2%80%9Cimage_latest%E2%80%9D_Fireable_Offense\" >The &#8220;image: latest&#8221; Fireable Offense<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#Why_Im_Still_Here\" >Why I&#8217;m Still Here<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#Related_Articles\" >Related Articles<\/a><\/li><\/ul><\/nav><\/div>\n<h3><span class=\"ez-toc-section\" id=\"The_Incident_Death_by_a_Thousand_API_Calls\"><\/span>The Incident: Death by a Thousand API Calls<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>At 03:14 AM on Saturday, the <code>kube-apiserver<\/code> on our production control plane decided it had seen enough of this world. It didn&#8217;t just crash; it OOM-killed itself with such violence that the underlying etcd nodes started desyncing. <\/p>\n<p>Why? Because one of you geniuses decided to deploy a &#8220;lightweight&#8221; microservice that didn&#8217;t have resource limits and, for reasons known only to God and your poorly written Go code, decided to list every single Secret in the namespace every 500 milliseconds.<\/p>\n<p>When the API server died, the Kubelets panicked. When the Kubelets panicked, they started killing pods. When the pods died, the ReplicaSets tried to recreate them. But since the API server was in a recursive loop of death, the scheduler couldn&#8217;t bind the pods to nodes. <\/p>\n<p>This is what my terminal looked like for six hours:<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\">$ kubectl get pods -A\nNAMESPACE     NAME                                      READY   STATUS             RESTARTS         AGE\ndefault       awesome-new-app-6f9d5f7d57-abc12          0\/1     CrashLoopBackOff   142 (4m ago)     2h\ndefault       awesome-new-app-6f9d5f7d57-def34          0\/1     Terminating        0                2h\nkube-system   kube-apiserver-ip-10-0-1-10.ec2.internal  0\/1     Error              15               48h\nkube-system   kube-controller-manager-ip-10-0-1-10      0\/1     CrashLoopBackOff   12               48h\nmonitoring    prometheus-k8s-0                          0\/2     ImagePullBackOff   0                5h\n<\/code><\/pre>\n<p>The <code>ImagePullBackOff<\/code> was a nice touch\u2014turns out that when the control plane is melting, the NAT gateway also decided to hit its connection limit because of the millions of retries. <\/p>\n<p>This isn&#8217;t v1.18 anymore. We are on v1.29. If you are still using <code>policy\/v1beta1<\/code> for your PodDisruptionBudgets or ignoring the fact that <code>selfLink<\/code> is gone, you aren&#8217;t just behind the times; you are a liability.<\/p>\n<hr \/>\n<h2><span class=\"ez-toc-section\" id=\"1_The_Resource_Limit_Lie_Why_%E2%80%98Best_Effort_is_a_Suicide_Pact\"><\/span>1. The Resource Limit Lie: Why &#8216;Best Effort&#8217; is a Suicide Pact<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>You love to talk about &#8220;elasticity.&#8221; You think Kubernetes is a magical infinite bucket of RAM. It isn&#8217;t. When you omit <code>requests<\/code> and <code>limits<\/code> from your YAML, you are telling the scheduler, &#8220;I don&#8217;t know what I&#8217;m doing, just figure it out.&#8221;<\/p>\n<p>Kubernetes assigns a Quality of Service (QoS) class to every pod. If you don&#8217;t define limits, you get <code>BestEffort<\/code>. Do you know what happens to <code>BestEffort<\/code> pods when a node gets tight on memory? They are the first ones lined up against the wall and shot by the OOM Killer.<\/p>\n<p><strong>The &#8220;Clever&#8221; Developer YAML (The Suicide Pact):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: apps\/v1\nkind: Deployment\nmetadata:\n  name: memory-hog\nspec:\n  template:\n    spec:\n      containers:\n      - name: app\n        image: our-registry.io\/app:latest # FIREABLE OFFENSE\n        # No resources defined. &quot;It'll be fine,&quot; you said.\n<\/code><\/pre>\n<p><strong>The SRE-Approved YAML (The &#8220;I Want to Sleep&#8221; Version):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: apps\/v1\nkind: Deployment\nmetadata:\n  name: stable-app\nspec:\n  template:\n    spec:\n      containers:\n      - name: app\n        image: our-registry.io\/app:v1.4.2 # Specific tag, you cowards\n        resources:\n          requests:\n            memory: &quot;256Mi&quot;\n            cpu: &quot;100m&quot;\n          limits:\n            memory: &quot;512Mi&quot;\n            cpu: &quot;500m&quot;\n<\/code><\/pre>\n<p>If your <code>limit<\/code> is significantly higher than your <code>request<\/code>, you are creating &#8220;burstable&#8221; pods. That\u2019s fine for a dev environment, but in prod, it leads to bin-packing nightmares. If the node is overcommitted and your pod tries to hit its limit, the kernel\u2019s OOM score for your process skyrockets. I spent three hours debugging why the <code>order-processor<\/code> kept vanishing, only to find out it was trying to burst to 2GB on a node that only had 200MB of slack.<\/p>\n<hr \/>\n<h2><span class=\"ez-toc-section\" id=\"2_Probes_that_Kill_How_Your_Liveness_Check_Created_a_Cascading_Failure\"><\/span>2. Probes that Kill: How Your Liveness Check Created a Cascading Failure<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I am begging you: stop using <code>livenessProbes<\/code> as a crutch for bad code. A liveness probe is meant to catch a deadlocked process, not to restart an app because the database is slow.<\/p>\n<p>During the outage, the <code>payment-gateway<\/code> service started lagging because the API server was slow. What did your liveness probe do? It saw a 2-second delay, decided the pod was &#8220;dead,&#8221; and killed it. <\/p>\n<p>This happened across all 50 replicas simultaneously. <\/p>\n<p>Now, instead of a slow service, we had <em>no<\/em> service. And when the new pods came up, they had to perform their &#8220;startup routine&#8221;\u2014which involves hitting the database to cache schema. Fifty pods hitting the DB at once killed the DB. <\/p>\n<p><strong>The &#8220;Death Spiral&#8221; Probe:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">livenessProbe:\n  httpGet:\n    path: \/healthz\n    port: 8080\n  initialDelaySeconds: 3\n  periodSeconds: 5\n  failureThreshold: 1 # Why would you do this?\n<\/code><\/pre>\n<p>If your <code>failureThreshold<\/code> is 1, a single network hiccup or a long Garbage Collection (GC) pause kills your container. Congratulations, you\u2019ve built a self-destruct mechanism. Use <code>startupProbes<\/code> for heavy lifting and give your <code>readinessProbes<\/code> some breathing room. A pod that isn&#8217;t &#8220;Ready&#8221; just stops receiving traffic; a pod that isn&#8217;t &#8220;Live&#8221; gets executed. Learn the difference.<\/p>\n<hr \/>\n<h2><span class=\"ez-toc-section\" id=\"3_RBAC_is_Not_Optional_Stop_Giving_%E2%80%98Cluster-Admin_to_Your_CICD_Bot\"><\/span>3. RBAC is Not Optional: Stop Giving &#8216;Cluster-Admin&#8217; to Your CI\/CD Bot<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I found a ServiceAccount in the <code>dev-tools<\/code> namespace yesterday named <code>jenkins-deployer<\/code>. It had a ClusterRoleBinding to <code>cluster-admin<\/code>. <\/p>\n<p>Do you know what that means? It means if anyone compromises that Jenkins instance\u2014which, let\u2019s be honest, is running a version of Java from the Mesozoic era\u2014they own the entire cluster. They can delete the <code>kube-system<\/code> namespace. They can steal the TLS certs. They can spin up 5,000 crypto-miners on our Spot instances.<\/p>\n<p>RBAC (Role-Based Access Control) is painful because it requires you to actually know what your application does. I know that\u2019s a lot to ask. But &#8220;I need to list pods&#8221; does not mean you need <code>verbs: [\"*\"]<\/code> on <code>resources: [\"*\"]<\/code>.<\/p>\n<p><strong>The &#8220;I&#8217;m Lazy&#8221; RBAC:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">rules:\n- apiGroups: [&quot;&quot;]\n  resources: [&quot;*&quot;]\n  verbs: [&quot;*&quot;]\n<\/code><\/pre>\n<p><strong>The &#8220;I Actually Care About Security&#8221; RBAC:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">rules:\n- apiGroups: [&quot;&quot;]\n  resources: [&quot;pods&quot;, &quot;services&quot;]\n  verbs: [&quot;get&quot;, &quot;list&quot;, &quot;watch&quot;]\n<\/code><\/pre>\n<p>Stop using the <code>default<\/code> ServiceAccount for everything. It has no permissions by design. Don&#8217;t &#8220;fix&#8221; it by adding permissions to it. Create a scoped ServiceAccount for your deployment. It takes ten lines of YAML. Just do it.<\/p>\n<hr \/>\n<h2><span class=\"ez-toc-section\" id=\"4_The_Hidden_Cost_of_%E2%80%9Ckubernetes_best%E2%80%9D_Intentions\"><\/span>4. The Hidden Cost of &#8220;kubernetes best&#8221; Intentions<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Most of you treat <strong>kubernetes best<\/strong> practices like the &#8220;Terms and Conditions&#8221; of a software update\u2014you scroll past the reality of how distributed systems actually fail just to click &#8216;Accept&#8217; on a broken deployment. You read a medium article about &#8220;Service Meshes&#8221; and suddenly you&#8217;re injecting Istio sidecars into every namespace without looking at the overhead.<\/p>\n<p>We are paying a &#8220;Sidecar Tax&#8221; of 1.5 vCPUs per node just to handle the Envoy proxies for services that literally only talk to one other service over plain HTTP. You want &#8220;observability,&#8221; but you won&#8217;t even instrument your code with Prometheus metrics. You expect the infrastructure to magically tell you why your Python script is leaking memory.<\/p>\n<p>And don&#8217;t get me started on Helm. Helm is a great way to install software you don&#8217;t understand. I\u2019ve seen charts where the <code>values.yaml<\/code> is 4,000 lines long, and you\u2019re just changing <code>replicaCount<\/code> and hoping the other 3,999 lines don&#8217;t blow up the cluster. When the outage hit, I tried to use <code>helm rollback<\/code>, but the release was in a <code>PENDING_UPGRADE<\/code> state because the last five deploys failed their readiness checks. I had to manually delete the Secret holding the Helm release state just to get the cluster to stop trying to deploy a broken image.<\/p>\n<hr \/>\n<h2><span class=\"ez-toc-section\" id=\"5_Networking_Sprawl_When_Your_CNI_Decides_to_Stop_Routing_Traffic\"><\/span>5. Networking Sprawl: When Your CNI Decides to Stop Routing Traffic<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Kubernetes networking is a lie built on top of iptables and BGP. Our CNI (Container Network Interface) plugin, Cilium, is fantastic\u2014until you overwhelm the <code>conntrack<\/code> table on the underlying Linux nodes.<\/p>\n<p>During the &#8220;Great API Meltdown,&#8221; the number of orphaned connections reached 65,536. At that point, the kernel started dropping packets. Not just for the broken app, but for <em>everything<\/em>. CoreDNS couldn&#8217;t resolve internal names. The Kubelet couldn&#8217;t heartbeat to the control plane.<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\">$ kubectl describe pod coredns-78fcdf6894-q4v7b -n kube-system\nEvents:\n  Type     Reason     Age                From               Message\n  ----     ------     ----               ----               -------\n  Warning  Unhealthy  2m (x24 over 10m)  kubelet            Liveness probe failed: Get &quot;http:\/\/10.0.1.15:8080\/health&quot;: dial tcp 10.0.1.15:8080: connect: connection refused\n<\/code><\/pre>\n<p>The reason? You guys are using <code>ndots:5<\/code> in your <code>dnsConfig<\/code>. Every time your app tries to talk to <code>database.internal<\/code>, it first tries to resolve:<br \/>\n1. <code>database.internal.my-namespace.svc.cluster.local<\/code><br \/>\n2. <code>database.internal.svc.cluster.local<\/code><br \/>\n3. <code>database.internal.cluster.local<\/code><br \/>\n4. <code>database.internal.us-east-1.compute.internal<\/code><\/p>\n<p>That\u2019s four DNS queries for every single connection. Multiply that by 1,000 requests per second, and you\u2019re DDoS-ing our own CoreDNS. If you\u2019re talking to a service in the same namespace, just use the service name. If you&#8217;re talking to an external URL, use a FQDN with a trailing dot to skip the search path.<\/p>\n<hr \/>\n<h2><span class=\"ez-toc-section\" id=\"6_Storage_Classes_and_the_Myth_of_Statelessness\"><\/span>6. Storage Classes and the Myth of Statelessness<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>&#8220;We&#8217;re cloud-native,&#8221; you said. &#8220;Our apps are stateless,&#8221; you said. Then why do I see 400 PersistentVolumeClaims (PVCs) in the production namespace?<\/p>\n<p>The biggest headache during the recovery was the <code>ReadWriteOnce<\/code> (RWO) access mode. We have a multi-AZ cluster. When Node A in <code>us-east-1a<\/code> died, the scheduler tried to move the pod to Node B in <code>us-east-1b<\/code>. But the EBS volume is locked to <code>us-east-1a<\/code>. <\/p>\n<p>The pod stayed in <code>ContainerCreating<\/code> for 45 minutes because the volume was &#8220;already attached to another node.&#8221;<\/p>\n<p><strong>The Error Log from Hell:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-text\">Warning  FailedAttachVolume  5m  attachdetach-controller  Multi-Attach error for volume &quot;pvc-12345&quot; Volume is already used by pod &quot;old-app-pod&quot;\n<\/code><\/pre>\n<p>If you need state, use a managed service like RDS or S3. If you <em>insist<\/em> on running a database inside K8s, you better have a <code>StatefulSet<\/code> and a very good reason. And for the love of all that is holy, stop using <code>strategy: Recreate<\/code> for your deployments just because your app can&#8217;t handle two versions of the schema running at once. That&#8217;s not a deployment strategy; that&#8217;s an admission of failure.<\/p>\n<p>A <code>RollingUpdate<\/code> is the standard. If your app can&#8217;t handle a <code>RollingUpdate<\/code>, your app isn&#8217;t ready for Kubernetes.<\/p>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"The_%E2%80%9Cimage_latest%E2%80%9D_Fireable_Offense\"><\/span>The &#8220;image: latest&#8221; Fireable Offense<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>I am making this its own section because I am tired of explaining it. If I see <code>image: latest<\/code> in a production manifest again, I am revoking your <code>git push<\/code> rights.<\/p>\n<p>When you use <code>latest<\/code>, you have no idea what code is actually running. If a node reboots and pulls the image again, it might get a different version than the other nodes in the same deployment. Now you have a distributed system running two different versions of the code with the same version string. Debugging that is like trying to find a black cat in a coal cellar at midnight.<\/p>\n<p>Use immutable tags. Use the Git SHA. Use a semantic version. Just don&#8217;t use <code>latest<\/code>.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Why_Im_Still_Here\"><\/span>Why I&#8217;m Still Here<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>I spent 48 hours fixing this because I actually care about the uptime. I care about the fact that our customers couldn&#8217;t process payments for four hours. I care that the junior SRE on call was crying because he thought he&#8217;d deleted the production database when, in reality, the networking was just so hosed he couldn&#8217;t see it.<\/p>\n<p>Kubernetes is a tool, not a personality trait. It is incredibly powerful, but it is also a loaded gun pointed directly at your foot. Every line of YAML you write is a decision. If you make those decisions based on &#8220;it worked in my local Minikube,&#8221; you are going to keep waking me up.<\/p>\n<p>I&#8217;m going home. I&#8217;m going to sleep for 14 hours. When I come back, I expect to see every one of your deployments updated with resource requests, proper probes, and pinned image tags.<\/p>\n<p>If I see another <code>OOMKilled<\/code> event caused by a <code>BestEffort<\/code> pod, I\u2019m not fixing it. I\u2019m just going to let the cluster burn and send you the bill for the AWS bill.<\/p>\n<p><strong>Tools you should actually learn to use before your next PR:<\/strong><br \/>\n&#8211; <strong>Stern:<\/strong> To tail multiple pod logs without losing your mind.<br \/>\n&#8211; <strong>Kube-capacity:<\/strong> To see how much of the cluster you&#8217;re actually wasting.<br \/>\n&#8211; <strong>Kustomize:<\/strong> To manage your YAML without the Helm-chart-hell.<br \/>\n&#8211; <strong>Prometheus\/Grafana:<\/strong> Look at the <code>container_memory_working_set_bytes<\/code> metric. It\u2019s the only one that matters for OOMs.<\/p>\n<p>Fix your YAML. Leave me alone.<\/p>\n<p>\u2014 Your exhausted SRE.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Related_Articles\"><\/span>Related Articles<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Explore more insights and best practices:<\/p>\n<ul>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/top-artificial-intelligence-best-practices-for-success-5\/\">Top Artificial Intelligence Best Practices For Success 5<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/master-aws-best-practices-optimize-your-cloud-performance\/\">Master Aws Best Practices Optimize Your Cloud Performance<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/master-html-the-essential-guide-to-building-websites\/\">Master Html The Essential Guide To Building Websites<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>It is 4:42 AM. The sun isn\u2019t up, but my blood pressure is. I\u2019ve spent the last 48 hours staring at a Grafana dashboard that looked like a heart monitor of a patient in active cardiac arrest. My eyes feel like they\u2019ve been rubbed with sandpaper, and the smell of stale, burnt coffee\u2014the kind that\u2019s &#8230; <a title=\"10 Kubernetes Best Practices for Production Success\" class=\"read-more\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\" aria-label=\"Read more  on 10 Kubernetes Best Practices for Production Success\">Read more<\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-4866","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.0 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>10 Kubernetes Best Practices for Production Success - ITSupportWale<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"10 Kubernetes Best Practices for Production Success - ITSupportWale\" \/>\n<meta property=\"og:description\" content=\"It is 4:42 AM. The sun isn\u2019t up, but my blood pressure is. I\u2019ve spent the last 48 hours staring at a Grafana dashboard that looked like a heart monitor of a patient in active cardiac arrest. My eyes feel like they\u2019ve been rubbed with sandpaper, and the smell of stale, burnt coffee\u2014the kind that\u2019s ... Read more\" \/>\n<meta property=\"og:url\" content=\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\" \/>\n<meta property=\"og:site_name\" content=\"ITSupportWale\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-24T15:46:48+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png\" \/>\n\t<meta property=\"og:image:width\" content=\"512\" \/>\n\t<meta property=\"og:image:height\" content=\"512\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Techie\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Techie\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\"},\"author\":{\"name\":\"Techie\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\"},\"headline\":\"10 Kubernetes Best Practices for Production Success\",\"datePublished\":\"2026-08-24T15:46:48+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\"},\"wordCount\":1849,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\",\"name\":\"10 Kubernetes Best Practices for Production Success - ITSupportWale\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\"},\"datePublished\":\"2026-08-24T15:46:48+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/itsupportwale.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"10 Kubernetes Best Practices for Production Success\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"name\":\"ITSupportWale\",\"description\":\"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides\",\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\",\"name\":\"itsupportwale\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"contentUrl\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"width\":1119,\"height\":144,\"caption\":\"itsupportwale\"},\"image\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\",\"name\":\"Techie\",\"sameAs\":[\"https:\/\/itsupportwale.com\",\"iswblogadmin\"],\"url\":\"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"10 Kubernetes Best Practices for Production Success - ITSupportWale","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/","og_locale":"en_US","og_type":"article","og_title":"10 Kubernetes Best Practices for Production Success - ITSupportWale","og_description":"It is 4:42 AM. The sun isn\u2019t up, but my blood pressure is. I\u2019ve spent the last 48 hours staring at a Grafana dashboard that looked like a heart monitor of a patient in active cardiac arrest. My eyes feel like they\u2019ve been rubbed with sandpaper, and the smell of stale, burnt coffee\u2014the kind that\u2019s ... Read more","og_url":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/","og_site_name":"ITSupportWale","article_publisher":"https:\/\/www.facebook.com\/Itsupportwale-298547177495978","article_published_time":"2026-08-24T15:46:48+00:00","og_image":[{"width":512,"height":512,"url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png","type":"image\/png"}],"author":"Techie","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Techie","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#article","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/"},"author":{"name":"Techie","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d"},"headline":"10 Kubernetes Best Practices for Production Success","datePublished":"2026-08-24T15:46:48+00:00","mainEntityOfPage":{"@id":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/"},"wordCount":1849,"commentCount":0,"publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/","url":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/","name":"10 Kubernetes Best Practices for Production Success - ITSupportWale","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/#website"},"datePublished":"2026-08-24T15:46:48+00:00","breadcrumb":{"@id":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/itsupportwale.com\/blog\/10-kubernetes-best-practices-for-production-success-3\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/itsupportwale.com\/blog\/"},{"@type":"ListItem","position":2,"name":"10 Kubernetes Best Practices for Production Success"}]},{"@type":"WebSite","@id":"https:\/\/itsupportwale.com\/blog\/#website","url":"https:\/\/itsupportwale.com\/blog\/","name":"ITSupportWale","description":"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides","publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/itsupportwale.com\/blog\/#organization","name":"itsupportwale","url":"https:\/\/itsupportwale.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","contentUrl":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","width":1119,"height":144,"caption":"itsupportwale"},"image":{"@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/Itsupportwale-298547177495978"]},{"@type":"Person","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d","name":"Techie","sameAs":["https:\/\/itsupportwale.com","iswblogadmin"],"url":"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/"}]}},"_links":{"self":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4866","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/comments?post=4866"}],"version-history":[{"count":0,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4866\/revisions"}],"wp:attachment":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/media?parent=4866"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/categories?post=4866"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/tags?post=4866"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}