{"id":4873,"date":"2026-09-03T00:07:06","date_gmt":"2026-09-02T18:37:06","guid":{"rendered":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/"},"modified":"2026-09-03T00:07:06","modified_gmt":"2026-09-02T18:37:06","slug":"what-is-kubernetes-a-comprehensive-guide-to-orchestration","status":"publish","type":"post","link":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/","title":{"rendered":"What is Kubernetes? A Comprehensive Guide to Orchestration"},"content":{"rendered":"<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_80 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6a98875246d5e\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6a98875246d5e\"  aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#Kubernetes_The_Distributed_System_You_Probably_Dont_Need_But_Are_Stuck_With_Anyway\" >Kubernetes: The Distributed System You Probably Don&#8217;t Need (But Are Stuck With Anyway)<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#The_%E2%80%9CHello_World%E2%80%9D_Lie\" >The &#8220;Hello World&#8221; Lie<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#The_Control_Plane_Where_the_Bodies_are_Buried\" >The Control Plane: Where the Bodies are Buried<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#Networking_The_CNI_and_the_ndots_Disaster\" >Networking: The CNI and the ndots Disaster<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#Storage_The_%E2%80%9CStateful%E2%80%9D_Fallacy\" >Storage: The &#8220;Stateful&#8221; Fallacy<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#The_OOMKiller_and_the_Memory_Trap\" >The OOMKiller and the Memory Trap<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#YAML-Hell_and_the_Configuration_Gap\" >YAML-Hell and the Configuration Gap<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#The_%22Real_World%22_Gotcha_Image_Pull_Secrets_and_Rate_Limits\" >The \"Real World\" Gotcha: Image Pull Secrets and Rate Limits<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#RBAC_The_Illusion_of_Security\" >RBAC: The Illusion of Security<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#The_Cost_of_%22Cloud_Native%22\" >The Cost of \"Cloud Native\"<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#The_Reality_Check\" >The Reality Check<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#Related_Articles\" >Related Articles<\/a><\/li><\/ul><\/nav><\/div>\n<h2><span class=\"ez-toc-section\" id=\"Kubernetes_The_Distributed_System_You_Probably_Dont_Need_But_Are_Stuck_With_Anyway\"><\/span>Kubernetes: The Distributed System You Probably Don&#8217;t Need (But Are Stuck With Anyway)<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>It was 3:14 AM on a Tuesday in 2021. I was staring at a Grafana dashboard that looked like a heart monitor for someone having a massive coronary. We were running a fleet of microservices on Kubernetes v1.20 in <code>us-east-1<\/code>. I had just pushed a change to the <code>ingress-nginx<\/code> controller config to &#8220;optimize&#8221; the buffer sizes for a new API endpoint. Within ninety seconds, the <code>kube-apiserver<\/code> latency spiked to thirty seconds. The cluster was effectively lobotomized. Nodes started dropping into <code>NotReady<\/code> status because the <code>kubelet<\/code> couldn&#8217;t report heartbeats. The <code>controller-manager<\/code>, sensing blood in the water, started rescheduling pods from the &#8220;dead&#8221; nodes onto the &#8220;live&#8221; ones, which promptly collapsed under the sudden surge of traffic and I\/O pressure.<\/p>\n<p>The culprit wasn&#8217;t the config change itself. It was a cascading failure triggered by a misconfigured <code>livenessProbe<\/code>. When the ingress controller reloaded, it momentarily stopped responding to health checks. Kubernetes, being the dutiful soldier it is, killed the pods. But because we hadn&#8217;t set proper <code>readinessProbes<\/code>, the service endpoints were removed before the new pods were ready. Traffic hit a black hole. The retry logic in our <code>checkout-service<\/code> was too aggressive, creating a thundering herd that saturated the conntrack tables on every worker node. I spent four hours manually scaling the <code>aws-node<\/code> daemonset just to get the CNI to stop choking on its own tongue. That is the reality of Kubernetes: it is a force multiplier for both your productivity and your ability to destroy your own infrastructure.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_%E2%80%9CHello_World%E2%80%9D_Lie\"><\/span>The &#8220;Hello World&#8221; Lie<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Most Kubernetes documentation is written by people who want to sell you a managed service or a YAML-templating tool. They show you a 10-line <code>Deployment<\/code> manifest and tell you that you&#8217;ve achieved &#8220;web-scale.&#8221; They don&#8217;t tell you about the <code>ndots: 5<\/code> issue in <code>\/etc\/resolv.conf<\/code> that adds 20ms of latency to every external DNS lookup. They don&#8217;t mention that <code>Resources.Limits.CPU<\/code> is implemented via CFS quotas, which will throttle your application into the dirt even if the node has 90% idle CPU. They definitely don&#8217;t talk about the nightmare of managing stateful sets when your EBS volume is stuck in <code>attaching<\/code> state for twenty minutes because of a race condition in the CSI driver.<\/p>\n<p>We use Kubernetes because we want an API for our infrastructure. That\u2019s it. It\u2019s not about &#8220;containers&#8221; anymore; it\u2019s about having a standardized way to describe a desired state and letting a control loop try\u2014and often fail\u2014to reach it. If you are running three Go binaries and a Postgres instance, you are paying a massive &#8220;complexity tax&#8221; for features you will never use. You are managing a distributed database (etcd), a complex overlay network (Overlay\/VxLAN), and a sophisticated scheduler just to do what a <code>systemd<\/code> unit and a bash script could do in 50 lines of code.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_Control_Plane_Where_the_Bodies_are_Buried\"><\/span>The Control Plane: Where the Bodies are Buried<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The <code>kube-apiserver<\/code> is the only thing that matters. Everything else\u2014the <code>kube-scheduler<\/code>, the <code>kube-controller-manager<\/code>, the <code>kubelet<\/code>\u2014is just a client of the API. When people talk about &#8220;Kubernetes scaling,&#8221; they are usually talking about etcd performance. If your etcd <code>fsync<\/code> latency climbs above 10ms, your cluster is a ticking time bomb.<\/p>\n<blockquote><p>\n    <strong>Pro-tip:<\/strong> Never, ever run etcd on the same disks as your application logs. I\u2019ve seen <code>fluentbit<\/code> saturate disk I\/O during a log spike, causing etcd to lose quorum, which resulted in the entire cluster entering a read-only state. Use dedicated NVMe drives for etcd.\n<\/p><\/blockquote>\n<p>The scheduler is another area where &#8220;magic&#8221; happens. It\u2019s essentially a massive <code>for<\/code> loop that matches pods to nodes. But it\u2019s a greedy algorithm. It doesn&#8217;t look at actual usage; it looks at <code>Requests<\/code>. If you have a node with 32GB of RAM and you have 4 pods requesting 8GB each, that node is &#8220;full,&#8221; even if those pods are only actually using 512MB. This leads to the &#8220;Bin Packing&#8221; problem. You end up with 40% cluster utilization but you can&#8217;t schedule a new pod because of your <code>Requests<\/code> settings.<\/p>\n<pre><code>\n# A typical \"I don't know what I'm doing\" resource block\nresources:\n  requests:\n    cpu: \"500m\"\n    memory: \"1Gi\"\n  limits:\n    cpu: \"2\" # This will cause CFS throttling. Avoid it for latency-sensitive apps.\n    memory: \"1Gi\" # Keep requests and limits equal for Memory to get 'Guaranteed' QoS.\n<\/code><\/pre>\n<p>If you set <code>limits.cpu<\/code>, the Linux kernel will enforce a quota. If your app is multi-threaded (like a Java or Go runtime), it will burst, hit the quota, and the kernel will stop it from running for the remainder of the period. This looks like &#8220;random&#8221; latency spikes in your APM. My opinion? Don&#8217;t set CPU limits unless you&#8217;re in a multi-tenant environment where you don&#8217;t trust the developers. Set <code>requests<\/code> to what you actually need and let the scheduler do its job.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Networking_The_CNI_and_the_ndots_Disaster\"><\/span>Networking: The CNI and the <code>ndots<\/code> Disaster<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Kubernetes networking is a lie built on top of <code>iptables<\/code> or <code>IPVS<\/code>. When a pod at <code>10.2.4.5<\/code> wants to talk to <code>api.stripe.com<\/code>, it doesn&#8217;t just go out. It goes through a series of transformations that would make a mathematician weep. If you&#8217;re using the default <code>bridge<\/code> mode, you&#8217;re losing performance. If you&#8217;re using <code>Calico<\/code> with <code>BGP<\/code>, you&#8217;re managing a network topology you probably aren&#8217;t qualified for. If you&#8217;re using <code>Cilium<\/code>, you&#8217;re using eBPF, which is brilliant until you need to debug why a packet is being dropped and your standard <code>tcpdump<\/code> tools show you nothing.<\/p>\n<p>Let&#8217;s talk about <code>ndots<\/code>. This is the single most common &#8220;hidden&#8221; performance killer. By default, Kubernetes sets <code>ndots: 5<\/code> in <code>\/etc\/resolv.conf<\/code>. This means if you try to resolve <code>api.stripe.com<\/code>, the resolver will try:<\/p>\n<ul>\n<li><code>api.stripe.com.namespace.svc.cluster.local<\/code><\/li>\n<li><code>api.stripe.com.svc.cluster.local<\/code><\/li>\n<li><code>api.stripe.com.cluster.local<\/code><\/li>\n<li><code>api.stripe.com.us-east-1.compute.internal<\/code><\/li>\n<li>And finally: <code>api.stripe.com<\/code><\/li>\n<\/ul>\n<p>That is four failed DNS lookups before it even tries the real domain. If your <code>CoreDNS<\/code> is under load, or if you&#8217;re hitting AWS&#8217;s 1024 packets-per-second limit on the VPC DNS, your application will start throwing <code>UnknownHostException<\/code> or <code>ErrImagePull<\/code>. The fix? Use a fully qualified domain name (FQDN) with a trailing dot: <code>api.stripe.com.<\/code> or manually override the <code>dnsConfig<\/code> in your pod spec.<\/p>\n<pre><code>\nspec:\n  dnsConfig:\n    options:\n      - name: ndots\n        value: \"1\"\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"Storage_The_%E2%80%9CStateful%E2%80%9D_Fallacy\"><\/span>Storage: The &#8220;Stateful&#8221; Fallacy<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I have a very simple rule: <strong>Do not run your primary database on Kubernetes.<\/strong> Yes, I know about the <code>Postgres Operator<\/code>. Yes, I know about <code>Rook\/Ceph<\/code>. I don&#8217;t care. The complexity of managing a distributed storage layer on top of a distributed container orchestrator on top of a distributed cloud provider&#8217;s block storage is a recipe for data loss.<\/p>\n<p>When a node dies in Kubernetes, the <code>PersistentVolume<\/code> (PV) is still attached to that dead node. The <code>AttachDetachController<\/code> has to realize the node is gone, wait for the timeout, call the AWS\/GCP API to detach the volume, and then attach it to the new node. This can take anywhere from 2 to 15 minutes. If your database is down for 15 minutes because a node failed, your &#8220;High Availability&#8221; is a joke. Use RDS. Use Cloud SQL. Use a managed service until you are at a scale where the cost of the managed service is higher than the salary of two full-time DBAs. Most of you are not at that scale.<\/p>\n<p>If you <em>must<\/em> run stateful workloads, use <code>LocalPersistentVolumes<\/code> with NVMe drives. You lose the ability to move pods between nodes easily, but you gain predictable I\/O and you don&#8217;t have to deal with the &#8220;Stuck EBS Volume&#8221; dance. But then you have to manage replication at the application level. Are you ready to manage <code>Patroni<\/code> or <code>Galera<\/code>? Probably not.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_OOMKiller_and_the_Memory_Trap\"><\/span>The OOMKiller and the Memory Trap<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Memory is not like CPU. You can&#8217;t &#8220;throttle&#8221; memory. When you run out, someone has to die. In Kubernetes, this is the <code>OOMKiller<\/code>. But it&#8217;s not just about your pod. There is a hierarchy of who gets killed first based on the <code>oom_score_adj<\/code>.<\/p>\n<ul>\n<li><strong>Guaranteed:<\/strong> (Requests == Limits). These are the last to be killed.<\/li>\n<li><strong>Burstable:<\/strong> (Requests < Limits). These are killed if the node is under pressure.<\/li>\n<li><strong>BestEffort:<\/strong> (No requests or limits). These are the first to be sacrificed to the chaos gods.<\/li>\n<\/ul>\n<p>I once saw a cluster where the <code>kubelet<\/code> itself was OOM-killed because a developer deployed a &#8220;log-scraper&#8221; with no limits that leaked memory. Because the <code>kubelet<\/code> was in the same cgroup as other system processes but didn&#8217;t have its memory properly reserved via <code>--kube-reserved<\/code>, the kernel killed the most important process on the node. The node went <code>NotReady<\/code>, the pods were rescheduled, and the cycle repeated on the next node. A literal &#8220;virus&#8221; of a pod that killed every node it touched.<\/p>\n<blockquote><p>\n    <strong>Note to self:<\/strong> Always set <code>--system-reserved<\/code> and <code>--kube-reserved<\/code> on your worker nodes. If you don&#8217;t, the Linux kernel will prioritize a random Python script over the Kubelet when things get tight.\n<\/p><\/blockquote>\n<h2><span class=\"ez-toc-section\" id=\"YAML-Hell_and_the_Configuration_Gap\"><\/span>YAML-Hell and the Configuration Gap<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We&#8217;ve traded 50 lines of bash for 50,000 lines of YAML. Helm is a band-aid on a bullet wound. It\u2019s a text-templating engine that doesn&#8217;t understand Kubernetes semantics. You end up with <code>values.yaml<\/code> files that are 2,000 lines long, and a single indentation error causes a production outage. <\/p>\n<p>The real problem is that YAML is static, but infrastructure is dynamic. We use tools like <code>Kustomize<\/code> or <code>CDK8s<\/code> to try and manage the mess, but we&#8217;re just adding layers of abstraction. The &#8220;Pragmatic&#8221; approach? Keep it as flat as possible. Avoid deeply nested Helm charts. If you can&#8217;t explain what a manifest does by looking at it for 30 seconds, it&#8217;s too complex.<\/p>\n<p>Consider the <code>terminationGracePeriodSeconds<\/code>. The default is 30. If your app takes 35 seconds to flush its buffers to the database, you are losing data every time you deploy. Kubernetes sends a <code>SIGTERM<\/code>, waits 30 seconds, and then sends a <code>SIGKILL<\/code>. Most people don&#8217;t even have a signal handler in their code, so the app just dies immediately anyway. You need to handle <code>SIGTERM<\/code> gracefully.<\/p>\n<pre><code>\n\/\/ Example of what your code SHOULD do\nsigChan := make(chan os.Signal, 1)\nsignal.Notify(sigChan, syscall.SIGTERM)\n<-sigChan\nlog.Println(\"Received SIGTERM, shutting down gracefully...\")\nserver.Shutdown(ctx) \/\/ This gives the app time to finish active requests\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"The_%22Real_World%22_Gotcha_Image_Pull_Secrets_and_Rate_Limits\"><\/span>The \"Real World\" Gotcha: Image Pull Secrets and Rate Limits<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Here is a fun one that only happens when you're already having a bad day. You have a massive outage. You need to scale your pods from 10 to 100 to handle the recovery traffic. But you're using Docker Hub, and you haven't configured <code>imagePullSecrets<\/code> with a paid account. Suddenly, half your pods are stuck in <code>ImagePullBackOff<\/code> because you've been rate-limited. Or worse, you're using <code>imagePullPolicy: Always<\/code>, and your private registry is down. Even though the image is already on the node, the <code>kubelet<\/code> refuses to start the pod because it can't check if there's a newer version.<\/p>\n<p><strong>The Expert Move:<\/strong> Use <code>imagePullPolicy: IfNotPresent<\/code>. Use a local registry (like ECR, GCR, or an internal Artifactory). Never depend on the public internet for your production availability. If Docker Hub goes down, your cluster should still be able to scale.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"RBAC_The_Illusion_of_Security\"><\/span>RBAC: The Illusion of Security<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Most companies have \"Cluster Admin\" for everyone because RBAC is hard. This is how you get a junior dev accidentally deleting the <code>kube-system<\/code> namespace because they thought they were in <code>minikube<\/code>. <\/p>\n<p>Kubernetes RBAC is verbose and easy to get wrong. You have <code>Roles<\/code>, <code>ClusterRoles<\/code>, <code>RoleBindings<\/code>, and <code>ClusterRoleBindings<\/code>. If you give a service account the ability to <code>create pods<\/code>, you have effectively given it <code>root<\/code> on the node, because it can just create a privileged pod that mounts the host's <code>\/<\/code> filesystem.<\/p>\n<pre><code>\n# This is a security nightmare. Don't do this.\nkind: ClusterRole\napiVersion: rbac.authorization.k8s.io\/v1\nmetadata:\n  name: \"allow-everything\"\nrules:\n- apiGroups: [\"*\"]\n  resources: [\"*\"]\n  verbs: [\"*\"]\n<\/code><\/pre>\n<p>Instead, use tools like <code>Gatekeeper<\/code> or <code>Kyverno<\/code> to enforce policies. Block privileged containers. Block hostPath mounts. Block containers running as root. If you don't enforce these at the admission controller level, your RBAC is just theater.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"The_Cost_of_%22Cloud_Native%22\"><\/span>The Cost of \"Cloud Native\"<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We need to talk about the cloud bill. Managed Kubernetes (EKS, GKE, AKS) isn't expensive because of the control plane fee ($70-$100\/month). It's expensive because of the \"hidden\" costs:<\/p>\n<ol>\n<li><strong>Inter-AZ Data Transfer:<\/strong> Kubernetes doesn't care about Availability Zones by default. A pod in <code>us-east-1a<\/code> will happily talk to a service in <code>us-east-1b<\/code>, and AWS will charge you $0.01 per GB. At scale, this can be thousands of dollars. Use <code>topologyKeys<\/code> or <code>Service Topology<\/code> to keep traffic local.<\/li>\n<li><strong>Load Balancers:<\/strong> Every <code>type: LoadBalancer<\/code> service creates a new cloud LB. At $20\/month per LB, plus data processing, this adds up. Use an Ingress Controller and a single LB.<\/li>\n<li><strong>NAT Gateway:<\/strong> If your nodes are in private subnets, all their traffic (including pulling images!) goes through a NAT Gateway. This is often the most expensive item on an AWS bill. Use VPC Endpoints for ECR, S3, and STS.<\/li>\n<li><strong>Empty Nodes:<\/strong> If your HPA (Horizontal Pod Autoscaler) scales down pods but your Cluster Autoscaler doesn't scale down nodes, you're paying for idle compute.<\/li>\n<\/ol>\n<h2><span class=\"ez-toc-section\" id=\"The_Reality_Check\"><\/span>The Reality Check<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Kubernetes is a tool for managing complexity with more complexity. It solves the \"it works on my machine\" problem by making it \"it doesn't work in production for reasons I don't understand.\" If you are a Senior SRE, your job isn't to \"install Kubernetes.\" Your job is to protect the business from the inherent instability of a distributed system. This means setting up <code>PodDisruptionBudgets<\/code> so a node upgrade doesn't take down your entire API. It means setting <code>priorityClass<\/code> so your critical payment service kicks off the \"cat-picture-generator\" pod when resources are low. It means knowing that <code>iptables -L<\/code> is still your best friend when the network goes sideways.<\/p>\n<p>I've spent a decade in this industry, and I've seen the hype cycle move from VMs to Mesos to Swarm to Kubernetes. The tech changes, but the failure modes remain the same: exhausted resources, unhandled signals, and a lack of understanding of the underlying primitives. Kubernetes is just a very fancy way to run <code>exec()<\/code> on a remote machine. Don't let the YAML fool you into thinking it's anything more than that.<\/p>\n<p>Stop trying to build the \"perfect\" platform. Build a platform that is observable enough that you can fix it when it inevitably breaks at 3 AM. Because it will break. And when it does, no amount of \"transformative\" cloud-native AI-driven auto-scaling will save you. Only a deep understanding of the Linux kernel and a very fast <code>kubectl<\/code> finger will.<\/p>\n<p>If you're still using <code>latest<\/code> tags in production, you deserve the outage you're about to have.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Related_Articles\"><\/span>Related Articles<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Explore more insights and best practices:<\/p>\n<ul>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/what-is-machine-learning-guide\/\">What Is Machine Learning Guide<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\">Artificial Intelligence Best Practices A Complete Guide 6<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/getting-started-with-progressive-web-app\/\">Getting Started With Progressive Web App<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Kubernetes: The Distributed System You Probably Don&#8217;t Need (But Are Stuck With Anyway) It was 3:14 AM on a Tuesday in 2021. I was staring at a Grafana dashboard that looked like a heart monitor for someone having a massive coronary. We were running a fleet of microservices on Kubernetes v1.20 in us-east-1. I had &#8230; <a title=\"What is Kubernetes? A Comprehensive Guide to Orchestration\" class=\"read-more\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\" aria-label=\"Read more  on What is Kubernetes? A Comprehensive Guide to Orchestration\">Read more<\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-4873","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.0 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>What is Kubernetes? A Comprehensive Guide to Orchestration - ITSupportWale<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"What is Kubernetes? A Comprehensive Guide to Orchestration - ITSupportWale\" \/>\n<meta property=\"og:description\" content=\"Kubernetes: The Distributed System You Probably Don&#8217;t Need (But Are Stuck With Anyway) It was 3:14 AM on a Tuesday in 2021. I was staring at a Grafana dashboard that looked like a heart monitor for someone having a massive coronary. We were running a fleet of microservices on Kubernetes v1.20 in us-east-1. I had ... Read more\" \/>\n<meta property=\"og:url\" content=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\" \/>\n<meta property=\"og:site_name\" content=\"ITSupportWale\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-02T18:37:06+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png\" \/>\n\t<meta property=\"og:image:width\" content=\"512\" \/>\n\t<meta property=\"og:image:height\" content=\"512\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Techie\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Techie\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\"},\"author\":{\"name\":\"Techie\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\"},\"headline\":\"What is Kubernetes? A Comprehensive Guide to Orchestration\",\"datePublished\":\"2026-09-02T18:37:06+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\"},\"wordCount\":2248,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\",\"name\":\"What is Kubernetes? A Comprehensive Guide to Orchestration - ITSupportWale\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\"},\"datePublished\":\"2026-09-02T18:37:06+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/itsupportwale.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"What is Kubernetes? A Comprehensive Guide to Orchestration\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"name\":\"ITSupportWale\",\"description\":\"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides\",\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\",\"name\":\"itsupportwale\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"contentUrl\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"width\":1119,\"height\":144,\"caption\":\"itsupportwale\"},\"image\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\",\"name\":\"Techie\",\"sameAs\":[\"https:\/\/itsupportwale.com\",\"iswblogadmin\"],\"url\":\"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"What is Kubernetes? A Comprehensive Guide to Orchestration - ITSupportWale","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/","og_locale":"en_US","og_type":"article","og_title":"What is Kubernetes? A Comprehensive Guide to Orchestration - ITSupportWale","og_description":"Kubernetes: The Distributed System You Probably Don&#8217;t Need (But Are Stuck With Anyway) It was 3:14 AM on a Tuesday in 2021. I was staring at a Grafana dashboard that looked like a heart monitor for someone having a massive coronary. We were running a fleet of microservices on Kubernetes v1.20 in us-east-1. I had ... Read more","og_url":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/","og_site_name":"ITSupportWale","article_publisher":"https:\/\/www.facebook.com\/Itsupportwale-298547177495978","article_published_time":"2026-09-02T18:37:06+00:00","og_image":[{"width":512,"height":512,"url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png","type":"image\/png"}],"author":"Techie","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Techie","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#article","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/"},"author":{"name":"Techie","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d"},"headline":"What is Kubernetes? A Comprehensive Guide to Orchestration","datePublished":"2026-09-02T18:37:06+00:00","mainEntityOfPage":{"@id":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/"},"wordCount":2248,"commentCount":0,"publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/","url":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/","name":"What is Kubernetes? A Comprehensive Guide to Orchestration - ITSupportWale","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/#website"},"datePublished":"2026-09-02T18:37:06+00:00","breadcrumb":{"@id":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-comprehensive-guide-to-orchestration\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/itsupportwale.com\/blog\/"},{"@type":"ListItem","position":2,"name":"What is Kubernetes? A Comprehensive Guide to Orchestration"}]},{"@type":"WebSite","@id":"https:\/\/itsupportwale.com\/blog\/#website","url":"https:\/\/itsupportwale.com\/blog\/","name":"ITSupportWale","description":"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides","publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/itsupportwale.com\/blog\/#organization","name":"itsupportwale","url":"https:\/\/itsupportwale.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","contentUrl":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","width":1119,"height":144,"caption":"itsupportwale"},"image":{"@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/Itsupportwale-298547177495978"]},{"@type":"Person","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d","name":"Techie","sameAs":["https:\/\/itsupportwale.com","iswblogadmin"],"url":"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/"}]}},"_links":{"self":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4873","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/comments?post=4873"}],"version-history":[{"count":0,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4873\/revisions"}],"wp:attachment":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/media?parent=4873"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/categories?post=4873"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/tags?post=4873"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}