{"id":4857,"date":"2026-08-12T21:33:57","date_gmt":"2026-08-12T16:03:57","guid":{"rendered":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/"},"modified":"2026-08-12T21:33:57","modified_gmt":"2026-08-12T16:03:57","slug":"docker-best-practices-10-tips-for-faster-leaner-images","status":"publish","type":"post","link":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/","title":{"rendered":"Docker Best Practices: 10 Tips for Faster, Leaner Images"},"content":{"rendered":"<p><strong>INCIDENT LOG: 2024-05-14T03:14:22Z<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\">[03:14:22] Kubelet: Warning FailedScheduling - 0\/12 nodes are available: 12 Insufficient memory.\n[03:15:01] Node-04: Kernel: [124098.44] oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=\/,mems_allowed=0,global_oom,task_memcg=\/kubepods\/besteffort\/pod-abc,task=node,pid=14221,uid=0\n[03:15:01] Node-04: Kernel: Out of memory: Killed process 14221 (node) total-vm:4.2GB, anon-rss:1.8GB, file-rss:0B, shmem-rss:0B\n[03:16:45] PagerDuty: [CRITICAL] Production API - High Error Rate (98%)\n[03:17:10] me@workstation:~$ docker pull registry.internal\/app:latest\nError response from daemon: manifest for registry.internal\/app:latest not found: manifest unknown: manifest unknown\n[03:18:30] me@workstation:~$ docker inspect registry.internal\/app:latest\n[]\nError: No such object: registry.internal\/app:latest\n[03:20:00] me@workstation:~$ # Who the hell pushed to production at 3 AM?\n<\/code><\/pre>\n<p>I\u2019ve been awake for 48 hours. My eyes feel like they\u2019ve been scrubbed with steel wool, and my caffeine intake has reached levels that would make a cardiologist weep. While the rest of the &#8220;Engineering&#8221; team was dreaming about whatever &#8220;clever&#8221; new framework is trending on Hacker News, I was digging through the wreckage of our production cluster because someone decided that &#8220;it works on my machine&#8221; was a valid deployment strategy.<\/p>\n<p>The following is not just a post-mortem. It is a mandatory remediation guide. If I see another Dockerfile that looks like it was written by a toddler with a copy of &#8220;Docker for Dummies&#8221; from 2014, I will personally revoke your SSH access and move your desk to the basement next to the backup tape drives.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_80 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6a7e1e3a872b9\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6a7e1e3a872b9\"  aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#1_The_%E2%80%98Latest_Tag_is_a_Suicide_Note_SHA256_Pinning_or_Bust\" >1. The &#8216;Latest&#8217; Tag is a Suicide Note: SHA256 Pinning or Bust<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#2_The_3GB_Layer_Cake_Multi-Stage_Builds_or_Death\" >2. The 3GB Layer Cake: Multi-Stage Builds or Death<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#3_Root_is_for_Idiots_Implementing_Non-Privileged_Users\" >3. Root is for Idiots: Implementing Non-Privileged Users<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#4_Signal_Handling_Why_Your_Container_Wont_Die\" >4. Signal Handling: Why Your Container Won&#8217;t Die<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#5_The_%E2%80%9CSecret_Leak%E2%80%9D_ENV_Instructions_are_Not_for_Secrets\" >5. The &#8220;Secret Leak&#8221;: ENV Instructions are Not for Secrets<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#6_Healthchecks_That_Actually_Mean_Something\" >6. Healthchecks That Actually Mean Something<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#7_Ephemeral_Storage_and_the_tmp_Explosion\" >7. Ephemeral Storage and the \/tmp Explosion<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#8_Layer_Caching_Stop_Invalidating_the_World\" >8. Layer Caching: Stop Invalidating the World<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#9_The_Base_Image_Choice_Alpine_vs_Debian_Slim\" >9. The Base Image Choice: Alpine vs. Debian Slim<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#Summary_of_Required_Actions\" >Summary of Required Actions<\/a><\/li><\/ul><\/nav><\/div>\n<h2><span class=\"ez-toc-section\" id=\"1_The_%E2%80%98Latest_Tag_is_a_Suicide_Note_SHA256_Pinning_or_Bust\"><\/span>1. The &#8216;Latest&#8217; Tag is a Suicide Note: SHA256 Pinning or Bust<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The outage started because a junior developer\u2014who shall remain nameless but knows exactly who they are\u2014pushed a &#8220;quick fix&#8221; to the <code>latest<\/code> tag. Our CI\/CD pipeline, configured with the backbone of a jellyfish, pulled <code>node:latest<\/code> as the base image. Between the last successful build and this morning\u2019s disaster, the upstream <code>node<\/code> image updated from a stable Debian Bullseye base to a version that broke our native C++ bindings. <\/p>\n<p>Using <code>latest<\/code> is not a &#8220;docker best&#8221; practice; it is a declaration of technical bankruptcy. When you use <code>latest<\/code>, you are playing Russian Roulette with your infrastructure. You have no guarantee of what code is actually running. You have no reproducibility. You have no soul.<\/p>\n<p>From this moment forward, every <code>FROM<\/code> instruction in this organization will use specific version tags and, more importantly, SHA256 digests. <\/p>\n<p><strong>The Wrong Way:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\">FROM node:latest\n# This is a ticking time bomb.\n<\/code><\/pre>\n<p><strong>The SRE-Approved Way:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\"># Node v20.11.1 on Debian Bullseye Slim\nFROM node:20.11.1-bullseye-slim@sha256:7d1a3674681607567786687006886e3260840c67537330777176793666f27f0d\n<\/code><\/pre>\n<p>By pinning the SHA256 digest, we ensure that the bits we tested in staging are the exact same bits that hit production. Docker Engine v24.0.7 doesn&#8217;t care about your feelings; it cares about the hash. If the hash changes, the build fails. That\u2019s how we sleep at night.<\/p>\n<blockquote>\n<p><strong>SRE Rant:<\/strong> I don&#8217;t care if it&#8217;s &#8220;inconvenient&#8221; to update the hash. You know what&#8217;s inconvenient? Explaining to the CTO why we lost $40k in revenue because you were too lazy to copy-paste a string of hex characters.<\/p>\n<\/blockquote>\n<h2><span class=\"ez-toc-section\" id=\"2_The_3GB_Layer_Cake_Multi-Stage_Builds_or_Death\"><\/span>2. The 3GB Layer Cake: Multi-Stage Builds or Death<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>When I finally managed to pull the failing image to my local machine to debug, I had enough time to go make a sandwich, eat it, and contemplate my career choices. The image was 3.4GB. For a Node.js API. <\/p>\n<p>I inspected the layers using <code>docker history<\/code> and found the culprit. The image contained the entire build toolchain: <code>gcc<\/code>, <code>g++<\/code>, <code>make<\/code>, <code>python3<\/code>, and the entire <code>node_modules<\/code> directory including <code>devDependencies<\/code>. This isn&#8217;t just bloat; it\u2019s a performance killer. Large images increase pull times, which increases the &#8220;Mean Time To Recovery&#8221; (MTTR). During a scaling event, those extra gigabytes mean the difference between a healthy cluster and a cascading failure.<\/p>\n<p>We are moving to multi-stage builds. This is a non-negotiable &#8220;docker best&#8221; requirement. You build your binaries in a &#8220;heavy&#8221; image and then copy only the necessary artifacts into a &#8220;slim&#8221; runtime image.<\/p>\n<p><strong>The Disaster Dockerfile:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\">FROM node:20\nWORKDIR \/app\nCOPY . .\nRUN apt-get update &amp;&amp; apt-get install -y build-essential python3\nRUN npm install\nCMD [&quot;node&quot;, &quot;server.js&quot;]\n# Result: 3.4GB of garbage.\n<\/code><\/pre>\n<p><strong>The Remediation Dockerfile:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\"># Stage 1: Build\nFROM node:20.11.1-bullseye-slim@sha256:7d1a3... AS builder\nWORKDIR \/app\nCOPY package*.json .\/\nRUN apt-get update &amp;&amp; apt-get install -y --no-install-recommends build-essential python3 \\\n    &amp;&amp; npm ci \\\n    &amp;&amp; apt-get purge -y --auto-remove build-essential python3 \\\n    &amp;&amp; rm -rf \/var\/lib\/apt\/lists\/*\n\nCOPY . .\nRUN npm run build\n\n# Stage 2: Runtime\nFROM node:20.11.1-bullseye-slim@sha256:7d1a3...\nWORKDIR \/app\n# Only copy what is strictly necessary\nCOPY --from=builder \/app\/dist .\/dist\nCOPY --from=builder \/app\/node_modules .\/node_modules\nCOPY --from=builder \/app\/package.json .\/package.json\n\nUSER node\nCMD [&quot;node&quot;, &quot;dist\/server.js&quot;]\n<\/code><\/pre>\n<p>By using multi-stage builds, we dropped the image size from 3.4GB to 180MB. We also reduced the attack surface by removing the compiler and other utilities that a hacker would love to find in a compromised container. Also, notice the <code>rm -rf \/var\/lib\/apt\/lists\/*<\/code>. If you don&#8217;t clean up your apt cache in the <em>same layer<\/em> it was created, it stays in the image forever. Layer caching is a double-edged sword; learn how to use it or get out of the kitchen.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"3_Root_is_for_Idiots_Implementing_Non-Privileged_Users\"><\/span>3. Root is for Idiots: Implementing Non-Privileged Users<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I checked the running processes on the compromised pod. Everything was running as <code>root<\/code>. Do you realize what that means? If there is a container breakout vulnerability (and they happen more often than you think), the attacker has root access to the host kernel. <\/p>\n<p>&#8220;But it&#8217;s easier to handle file permissions as root!&#8221; <\/p>\n<p>I don&#8217;t care. Your laziness is a security liability. A &#8220;docker best&#8221; configuration always specifies a non-privileged user. Most official images (like Node, Python, and Alpine) already include a default non-root user, but you&#8217;re too busy &#8220;innovating&#8221; to use them.<\/p>\n<p>If you are using a base image that doesn&#8217;t have a user, you create one. Here is the exact syntax you will use. No exceptions.<\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\">FROM debian:bullseye-slim@sha256:6248bc...\n\n# Create a system group and user\nRUN groupadd -g 10001 appgroup &amp;&amp; \\\n    useradd -u 10000 -g appgroup -m -s \/bin\/bash appuser\n\nWORKDIR \/home\/appuser\/app\nCOPY --chown=appuser:appgroup . .\n\n# Switch to the non-privileged user\nUSER 10000\n\n# Now the application runs without root privileges\nCMD [&quot;.\/my-binary&quot;]\n<\/code><\/pre>\n<p>Why use IDs instead of names? Because Kubernetes and other orchestrators handle <code>runAsUser<\/code> directives better with UIDs. It avoids ambiguity when the host&#8217;s <code>\/etc\/passwd<\/code> doesn&#8217;t match the container&#8217;s. If I see <code>USER root<\/code> in a Dockerfile again, I will trigger a manual OOMKill on your local machine.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"4_Signal_Handling_Why_Your_Container_Wont_Die\"><\/span>4. Signal Handling: Why Your Container Won&#8217;t Die<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>During the outage, I tried to restart the pods. Kubernetes sent a <code>SIGTERM<\/code>. The pods ignored it. Kubernetes waited 30 seconds, got annoyed, and sent a <code>SIGKILL<\/code>. This resulted in corrupted database transactions and orphaned file locks because the application didn&#8217;t have time to shut down gracefully.<\/p>\n<p>The reason? The junior used the &#8220;shell form&#8221; of <code>CMD<\/code> instead of the &#8220;exec form.&#8221;<\/p>\n<p>When you write <code>CMD node server.js<\/code>, Docker wraps your command in <code>\/bin\/sh -c<\/code>. This makes <code>\/bin\/sh<\/code> PID 1. When the kernel sends a <code>SIGTERM<\/code> to the container, it goes to <code>\/bin\/sh<\/code>. And guess what? Shells don&#8217;t pass signals to their child processes unless you explicitly tell them to. Your Node.js app never even knew it was being asked to stop.<\/p>\n<p><strong>The Wrong Way (Shell Form):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\">CMD node server.js\n# PID 1 is \/bin\/sh. Signals are swallowed.\n<\/code><\/pre>\n<p><strong>The &#8220;Docker Best&#8221; Way (Exec Form):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\">ENTRYPOINT [&quot;\/usr\/bin\/tini&quot;, &quot;--&quot;]\nCMD [&quot;node&quot;, &quot;server.js&quot;]\n# PID 1 is tini (an init helper), which correctly reaps zombies and forwards signals.\n<\/code><\/pre>\n<p>We are now standardizing on using <code>tini<\/code> or <code>dumb-init<\/code> for all containers. It handles the PID 1 problem, ensures that <code>SIGTERM<\/code> actually reaches your application, and prevents zombie processes from clogging up the process table. If your app takes 30 seconds to die, it better be because it&#8217;s flushing a massive buffer, not because it&#8217;s ignoring the kernel.<\/p>\n<blockquote>\n<p><strong>Internal Note:<\/strong> If you don&#8217;t understand the difference between a signal and a syscall, go back to CS 101. I&#8217;m not here to tutor you; I&#8217;m here to keep the site alive.<\/p>\n<\/blockquote>\n<h2><span class=\"ez-toc-section\" id=\"5_The_%E2%80%9CSecret_Leak%E2%80%9D_ENV_Instructions_are_Not_for_Secrets\"><\/span>5. The &#8220;Secret Leak&#8221;: ENV Instructions are Not for Secrets<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>This was the highlight of my 48-hour nightmare. While inspecting the image layers to find out why it was so big, I ran <code>docker history --no-trunc<\/code>. And there it was, in plain text, for anyone with access to our internal registry to see:<\/p>\n<p><code>ENV STRIPE_API_KEY=sk_test_51Mz...<\/code><br \/>\n<code>ENV AWS_SECRET_ACCESS_KEY=AKIA...<\/code><\/p>\n<p>I had to spend four hours rotating every single production secret because someone thought <code>ENV<\/code> was a secure way to pass credentials. <\/p>\n<p>Let me be clear: <strong>Anything in a Dockerfile is public knowledge to anyone who has the image.<\/strong> Even if you <code>unset<\/code> the variable in a later layer, it is still there in the previous layer. Docker layers are immutable. You cannot &#8220;delete&#8221; a secret once it&#8217;s been baked in.<\/p>\n<p>To follow &#8220;docker best&#8221; practices, we use build-time secrets or runtime environment variables provided by the orchestrator (Kubernetes Secrets), never baked into the image.<\/p>\n<p><strong>How to use Build-Time Secrets (Docker BuildKit):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\"># syntax=docker\/dockerfile:1\nFROM node:20.11.1-bullseye-slim\nWORKDIR \/app\n\n# Mount the secret during build, it never touches the image layers\nRUN --mount=type=secret,id=my_secret \\\n    export SECRET_VAL=$(cat \/run\/secrets\/my_secret) &amp;&amp; \\\n    npm run build -- --api-key=$SECRET_VAL\n\nCMD [&quot;node&quot;, &quot;server.js&quot;]\n<\/code><\/pre>\n<p>When building, you pass the secret like this:<br \/>\n<code>docker build --secret id=my_secret,src=.\/secret.txt .<\/code><\/p>\n<p>The secret is mounted in a temporary filesystem (<code>tmpfs<\/code>) and never hits the disk or the image layers. If I find a secret in a <code>docker history<\/code> log again, I\u2019m not just revoking your access; I\u2019m calling HR.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"6_Healthchecks_That_Actually_Mean_Something\"><\/span>6. Healthchecks That Actually Mean Something<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The junior added a <code>HEALTHCHECK<\/code> instruction. Great, right? Wrong.<br \/>\n<code>HEALTHCHECK CMD curl -f http:\/\/localhost:8080\/ || exit 1<\/code><\/p>\n<p>During the outage, the Node.js event loop was completely blocked by a synchronous CPU-intensive task (another &#8220;clever&#8221; optimization). The port was &#8220;open,&#8221; so <code>curl<\/code> succeeded, but the application was effectively dead. It couldn&#8217;t process a single request. Kubernetes thought the pod was healthy, so it kept sending traffic to a black hole.<\/p>\n<p>A &#8220;docker best&#8221; healthcheck must validate the internal state of the application, not just the network stack. It should check database connectivity, cache availability, and whether the event loop lag is within acceptable limits.<\/p>\n<p><strong>The &#8220;I Actually Care About My App&#8221; Healthcheck:<\/strong><\/p>\n<p>Create a <code>healthcheck.js<\/code> script:<\/p>\n<pre class=\"codehilite\"><code class=\"language-javascript\">const http = require('http');\nconst options = {\n    host: 'localhost',\n    port: 8080,\n    path: '\/health',\n    timeout: 2000\n};\n\nconst request = http.request(options, (res) =&gt; {\n    if (res.statusCode === 200) {\n        process.exit(0);\n    } else {\n        process.exit(1);\n    }\n});\n\nrequest.on('error', (err) =&gt; {\n    process.exit(1);\n});\n\nrequest.end();\n<\/code><\/pre>\n<p>Then in your Dockerfile:<\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\">HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \\\n  CMD [ &quot;node&quot;, &quot;healthcheck.js&quot; ]\n<\/code><\/pre>\n<p>And for the love of all that is holy, make sure your <code>\/health<\/code> endpoint doesn&#8217;t just return <code>{\"status\": \"ok\"}<\/code>. It should actually <em>ping<\/em> the database. If the database is down, the app is down. Tell the truth.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"7_Ephemeral_Storage_and_the_tmp_Explosion\"><\/span>7. Ephemeral Storage and the \/tmp Explosion<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The final nail in the coffin for Node-04 was the ephemeral storage. The application was writing debug logs to <code>\/app\/logs\/debug.log<\/code> inside the container. Since containers use an overlay filesystem (OverlayFS), every write operation creates a &#8220;copy-on-write&#8221; action. <\/p>\n<p>The junior didn&#8217;t set a log rotation policy. The container&#8217;s writable layer grew until it consumed all available disk space on the host&#8217;s <code>\/var\/lib\/docker\/overlay2<\/code> directory. This triggered a DiskPressure eviction, which killed not just the offending pod, but three other critical services on the same node.<\/p>\n<p><strong>Remediation:<\/strong><br \/>\n1. <strong>Logs go to STDOUT\/STDERR.<\/strong> Period. No log files inside the container. Let the container runtime and the logging driver (Fluentd, Loki, etc.) handle the persistence.<br \/>\n2. <strong>Use Tmpfs for temporary files.<\/strong> If you absolutely must write temporary files, use a <code>tmpfs<\/code> mount so they stay in RAM and don&#8217;t bloat the image layers.<\/p>\n<p><strong>Example Kubernetes Pod Spec (since you clearly can&#8217;t be trusted with Docker alone):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">spec:\n  containers:\n  - name: app\n    image: registry.internal\/app:v1.2.3@sha256:...\n    resources:\n      limits:\n        memory: &quot;512Mi&quot;\n        cpu: &quot;500m&quot;\n        ephemeral-storage: &quot;1Gi&quot; # Limit the damage\n    volumeMounts:\n    - name: cache-volume\n      mountPath: \/tmp\n  volumes:\n  - name: cache-volume\n    emptyDir:\n      medium: Memory\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"8_Layer_Caching_Stop_Invalidating_the_World\"><\/span>8. Layer Caching: Stop Invalidating the World<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Every time you change a single line of code, our CI\/CD takes 15 minutes to build. Why? Because you put <code>COPY . .<\/code> at the top of your Dockerfile.<\/p>\n<p>Docker builds images in layers. If a layer changes, every subsequent layer must be rebuilt. By copying your entire source code before running <code>npm install<\/code>, you invalidate the <code>npm install<\/code> layer every time you change a comment in a README file.<\/p>\n<p><strong>The &#8220;I Value My Time&#8221; Layering:<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-dockerfile\">WORKDIR \/app\n\n# Copy only the dependency files first\nCOPY package.json package-lock.json .\/\n\n# This layer is cached unless package.json changes\nRUN npm ci\n\n# Now copy the rest of the code\nCOPY . .\n\n# This layer only runs if the code changes\nRUN npm run build\n<\/code><\/pre>\n<p>This simple change reduced our build time from 15 minutes to 2 minutes. That\u2019s 13 minutes of my life I get back every time you push a bug. Multiply that by 50 developers and 20 pushes a day, and you\u2019ve just saved the company thousands of dollars in compute costs and developer frustration.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"9_The_Base_Image_Choice_Alpine_vs_Debian_Slim\"><\/span>9. The Base Image Choice: Alpine vs. Debian Slim<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>I see a lot of you using <code>alpine<\/code> because it&#8217;s &#8220;small.&#8221; Alpine uses <code>musl<\/code> libc. Node.js and many Python libraries are built against <code>glibc<\/code> (standard on Debian\/Ubuntu). When you run a <code>glibc<\/code> application on <code>musl<\/code>, you often run into weird performance issues, DNS resolution bugs, or &#8220;File not found&#8221; errors that make no sense.<\/p>\n<p>Unless you are an expert in C-standard library compatibility, use <code>debian-slim<\/code>. It\u2019s slightly larger (maybe 30MB more), but it\u2019s much more stable and uses <code>glibc<\/code>. We are an enterprise, not a hobbyist <a href=\"https:\/\/itsupportwale.com\/blog\/\" title=\"Read more about blog\">blog<\/a>. Stability beats a few megabytes every single time.<\/p>\n<p><strong>Current Standard:<\/strong> <code>debian:bullseye-slim<\/code> or <code>node:20-bullseye-slim<\/code>. <\/p>\n<p>If I see an <code>alpine<\/code> image that hasn&#8217;t been thoroughly vetted for <code>musl<\/code> compatibility, I will reject the PR. I don&#8217;t have the energy to debug your segmentation faults at 4 AM.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Summary_of_Required_Actions\"><\/span>Summary of Required Actions<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<ol>\n<li><strong>Audit all Dockerfiles:<\/strong> Replace <code>latest<\/code> with SHA256 pinned tags.<\/li>\n<li><strong>Refactor for Multi-Stage:<\/strong> Remove build tools from runtime images.<\/li>\n<li><strong>Drop Privileges:<\/strong> Add <code>USER<\/code> instructions to every single image.<\/li>\n<li><strong>Fix Signal Handling:<\/strong> Use the exec form <code>[\"cmd\", \"arg\"]<\/code> and <code>tini<\/code>.<\/li>\n<li><strong>Purge Secrets:<\/strong> Move all <code>ENV<\/code> secrets to Kubernetes Secrets or build-time mounts.<\/li>\n<li><strong>Implement Real Healthchecks:<\/strong> Check the DB, not just the port.<\/li>\n<\/ol>\n<p>I am going to sleep now. If my pager goes off because someone ignored this guide, may the kernel have mercy on your soul, because I won&#8217;t. <\/p>\n<p><strong>SRE Signed Off: 2024-05-16T09:00:00Z<\/strong><br \/>\n<strong>Status: Irritable.<\/strong><\/p>\n","protected":false},"excerpt":{"rendered":"<p>INCIDENT LOG: 2024-05-14T03:14:22Z [03:14:22] Kubelet: Warning FailedScheduling &#8211; 0\/12 nodes are available: 12 Insufficient memory. [03:15:01] Node-04: Kernel: [124098.44] oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=\/,mems_allowed=0,global_oom,task_memcg=\/kubepods\/besteffort\/pod-abc,task=node,pid=14221,uid=0 [03:15:01] Node-04: Kernel: Out of memory: Killed process 14221 (node) total-vm:4.2GB, anon-rss:1.8GB, file-rss:0B, shmem-rss:0B [03:16:45] PagerDuty: [CRITICAL] Production API &#8211; High Error Rate (98%) [03:17:10] me@workstation:~$ docker pull registry.internal\/app:latest Error response from daemon: manifest for &#8230; <a title=\"Docker Best Practices: 10 Tips for Faster, Leaner Images\" class=\"read-more\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\" aria-label=\"Read more  on Docker Best Practices: 10 Tips for Faster, Leaner Images\">Read more<\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-4857","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.0 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Docker Best Practices: 10 Tips for Faster, Leaner Images - ITSupportWale<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Docker Best Practices: 10 Tips for Faster, Leaner Images - ITSupportWale\" \/>\n<meta property=\"og:description\" content=\"INCIDENT LOG: 2024-05-14T03:14:22Z [03:14:22] Kubelet: Warning FailedScheduling - 0\/12 nodes are available: 12 Insufficient memory. [03:15:01] Node-04: Kernel: [124098.44] oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=\/,mems_allowed=0,global_oom,task_memcg=\/kubepods\/besteffort\/pod-abc,task=node,pid=14221,uid=0 [03:15:01] Node-04: Kernel: Out of memory: Killed process 14221 (node) total-vm:4.2GB, anon-rss:1.8GB, file-rss:0B, shmem-rss:0B [03:16:45] PagerDuty: [CRITICAL] Production API - High Error Rate (98%) [03:17:10] me@workstation:~$ docker pull registry.internal\/app:latest Error response from daemon: manifest for ... Read more\" \/>\n<meta property=\"og:url\" content=\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\" \/>\n<meta property=\"og:site_name\" content=\"ITSupportWale\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-12T16:03:57+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png\" \/>\n\t<meta property=\"og:image:width\" content=\"512\" \/>\n\t<meta property=\"og:image:height\" content=\"512\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Techie\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Techie\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"12 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\"},\"author\":{\"name\":\"Techie\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\"},\"headline\":\"Docker Best Practices: 10 Tips for Faster, Leaner Images\",\"datePublished\":\"2026-08-12T16:03:57+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\"},\"wordCount\":1861,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\",\"name\":\"Docker Best Practices: 10 Tips for Faster, Leaner Images - ITSupportWale\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\"},\"datePublished\":\"2026-08-12T16:03:57+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/itsupportwale.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Docker Best Practices: 10 Tips for Faster, Leaner Images\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"name\":\"ITSupportWale\",\"description\":\"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides\",\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\",\"name\":\"itsupportwale\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"contentUrl\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"width\":1119,\"height\":144,\"caption\":\"itsupportwale\"},\"image\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\",\"name\":\"Techie\",\"sameAs\":[\"https:\/\/itsupportwale.com\",\"iswblogadmin\"],\"url\":\"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Docker Best Practices: 10 Tips for Faster, Leaner Images - ITSupportWale","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/","og_locale":"en_US","og_type":"article","og_title":"Docker Best Practices: 10 Tips for Faster, Leaner Images - ITSupportWale","og_description":"INCIDENT LOG: 2024-05-14T03:14:22Z [03:14:22] Kubelet: Warning FailedScheduling - 0\/12 nodes are available: 12 Insufficient memory. [03:15:01] Node-04: Kernel: [124098.44] oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=\/,mems_allowed=0,global_oom,task_memcg=\/kubepods\/besteffort\/pod-abc,task=node,pid=14221,uid=0 [03:15:01] Node-04: Kernel: Out of memory: Killed process 14221 (node) total-vm:4.2GB, anon-rss:1.8GB, file-rss:0B, shmem-rss:0B [03:16:45] PagerDuty: [CRITICAL] Production API - High Error Rate (98%) [03:17:10] me@workstation:~$ docker pull registry.internal\/app:latest Error response from daemon: manifest for ... Read more","og_url":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/","og_site_name":"ITSupportWale","article_publisher":"https:\/\/www.facebook.com\/Itsupportwale-298547177495978","article_published_time":"2026-08-12T16:03:57+00:00","og_image":[{"width":512,"height":512,"url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png","type":"image\/png"}],"author":"Techie","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Techie","Est. reading time":"12 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#article","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/"},"author":{"name":"Techie","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d"},"headline":"Docker Best Practices: 10 Tips for Faster, Leaner Images","datePublished":"2026-08-12T16:03:57+00:00","mainEntityOfPage":{"@id":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/"},"wordCount":1861,"commentCount":0,"publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/","url":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/","name":"Docker Best Practices: 10 Tips for Faster, Leaner Images - ITSupportWale","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/#website"},"datePublished":"2026-08-12T16:03:57+00:00","breadcrumb":{"@id":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/itsupportwale.com\/blog\/docker-best-practices-10-tips-for-faster-leaner-images\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/itsupportwale.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Docker Best Practices: 10 Tips for Faster, Leaner Images"}]},{"@type":"WebSite","@id":"https:\/\/itsupportwale.com\/blog\/#website","url":"https:\/\/itsupportwale.com\/blog\/","name":"ITSupportWale","description":"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides","publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/itsupportwale.com\/blog\/#organization","name":"itsupportwale","url":"https:\/\/itsupportwale.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","contentUrl":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","width":1119,"height":144,"caption":"itsupportwale"},"image":{"@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/Itsupportwale-298547177495978"]},{"@type":"Person","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d","name":"Techie","sameAs":["https:\/\/itsupportwale.com","iswblogadmin"],"url":"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/"}]}},"_links":{"self":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4857","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/comments?post=4857"}],"version-history":[{"count":0,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4857\/revisions"}],"wp:attachment":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/media?parent=4857"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/categories?post=4857"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/tags?post=4857"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}