Table of Contents
Docker is Not a Virtual Machine, and Your Images are Too Fat
It was 3:14 AM on a Tuesday in 2017. I was the “on-call hero” for a fintech startup that shall remain nameless. We had a monolithic Python API that we’d recently “containerized” to be modern. I pushed a change to the master branch, the CI/CD pipeline hummed along, and the deployment triggered. Ten minutes later, the PagerDuty alert started screaming. Not just one service. Everything. The Kubelet on our primary nodes was reporting DiskPressure and started evicting pods like a bouncer at a dive bar. I’d forgotten to add a .dockerignore file. My local venv, 4GB of temporary data science models, and a massive .git folder had been sucked into the build context, pushed to our private registry, and pulled onto every node simultaneously. The nodes ran out of disk space, the container runtime choked, and the entire cluster entered a death spiral.
I spent the next four hours manually cleaning up /var/lib/docker/overlay2 on six different instances while the CTO watched the Slack channel in silence. That’s the reality of Docker. It isn’t a “seamless” abstraction layer. It’s a leaky bucket of kernel namespaces, cgroups, and filesystem layers that will bite you the moment you stop respecting the underlying Linux primitives. If you treat it like a “lightweight VM,” you’ve already lost. Docker is a process wrapper with an identity crisis. Let’s stop pretending the “Hello World” tutorial is enough to run production systems.
The Documentation Lies to You
Most Docker documentation is written for developers who want to run a database on their laptop. It’s not written for SREs who have to manage 500 nodes in us-east-1. The docs tell you that FROM python:3.9 is a great starting point. It’s not. That image is nearly 900MB because it includes every build tool, header file, and obscure library you’ll never use. In production, every megabyte is a liability. It’s more time spent in docker pull, more money spent on NAT Gateway egress, and a larger attack surface for the next CVE-2024-whatever.
The industry is obsessed with the “Dockerize everything” hype, but we rarely talk about the cost of the abstraction. We’ve traded “it works on my machine” for “it works in the container but the kernel is OOM-killing the process because I didn’t set cgroup limits correctly.” Docker is a tool for packaging, not a substitute for understanding how Linux manages resources. If you don’t know what set -e does in a shell script, you shouldn’t be writing Dockerfiles.
The Meat: Layers, Cache, and the Union File System
Docker images are just a stack of tarballs. That’s it. When you see Step 4/10 : RUN apt-get update, Docker is creating a new layer. If you change a line at the top of your Dockerfile, every layer below it is invalidated. This is where most people fail at “Docker 101.”
# BAD DOCKERFILE
FROM node:18
COPY . /app
WORKDIR /app
RUN npm install
CMD ["node", "index.js"]
In the example above, every time you change a single character in a comment in index.js, Docker re-runs npm install. You’re wasting five minutes of CI time and downloading half the internet for no reason. You have to exploit the layer cache. You copy the dependency manifest first, install, and then copy the source code.
# BETTER DOCKERFILE
FROM node:18-slim
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci --production
COPY . .
USER node
CMD ["node", "index.js"]
But even this is amateur hour. If you’re a Senior SRE, you should be using multi-stage builds. There is zero reason for your production image to contain a compiler, a git client, or your ssh keys. You build the binary in one stage and copy it to a “distroless” or minimal base in the second.
Pro-tip: Use
--mount=type=cachewith BuildKit to persist your package manager’s cache between builds. It’s the difference between a 2-minute build and a 10-second build.
# THE ADULT WAY (Multi-stage + BuildKit Cache)
# syntax=docker/dockerfile:1.4
FROM golang:1.21-alpine AS builder
WORKDIR /src
RUN --mount=type=cache,target=/go/pkg/mod \
--mount=type=bind,source=go.sum,target=go.sum \
--mount=type=bind,source=go.mod,target=go.mod \
go mod download
COPY . .
RUN go build -o /bin/api ./cmd/api
FROM alpine:3.18
RUN apk add --no-cache ca-certificates tzdata
COPY --from=builder /bin/api /bin/api
USER 1000
ENTRYPOINT ["/bin/api"]
Notice the USER 1000. If I see root running your app in production, I’m revoking your SSH access. Containers don’t provide a security boundary by default. If your process is root inside the container and someone escapes via a kernel exploit, they’re root on the host. It’s that simple.
The Alpine Trap: glibc vs. musl
Everyone loves Alpine because it’s 5MB. It’s the “SRE’s darling.” But Alpine uses musl instead of glibc. If you’re running Go or Rust, you probably won’t notice until you try to use a C-binding that expects glibc. If you’re running Python, you’re in for a world of hurt. Most Python wheels (pre-compiled binaries) are built for manylinux (glibc). When you pip install pandas on Alpine, it can’t find a compatible wheel, so it starts compiling from source. Your 30-second build just became a 20-minute build, and your “small” image is now bloated with gcc and make just to get the thing to install.
I use debian-slim. It’s 30MB larger, but it uses glibc. It’s predictable. It doesn’t have weird DNS resolution bugs because musl handles /etc/resolv.conf differently than every other Linux distro. Stop chasing the 5MB dragon and start chasing the “I don’t want to debug DNS at 2 AM” dragon.
- Alpine: Great for static binaries (Go, Rust) or simple utilities (curl, jq).
- Debian-Slim: The gold standard for Python, Node, and Ruby.
- Distroless: The final boss of security. No shell, no package manager, just your binary.
- Ubuntu: Only if you really need a specific PPA or outdated library.
The Networking Nightmare
Docker networking is a mess of iptables rules that will make your head spin. By default, Docker uses the bridge driver. It creates a virtual bridge (docker0), assigns an IP range, and uses NAT to let containers talk to the outside world. This is fine for your local dev environment. It’s a disaster for high-performance networking.
Every packet going through that bridge has to be processed by the NAT engine. If you’re running a high-throughput database or a proxy like Nginx, you’ll see a measurable latency hit. This is why we use --network=host for performance-critical services, though it comes with the massive downside of sharing the host’s network namespace (no port isolation).
And don’t get me started on MTU (Maximum Transmission Unit). I once spent three days debugging why a container could curl a small JSON payload but timed out on a 10MB file. The host was on an AWS VPC with an MTU of 9001 (Jumbo Frames), but the Docker bridge was defaulted to 1500. The packets were being dropped silently. Note to self: Always check ip addr show docker0 when things get weird.
PID 1 and the Zombie Apocalypse
In Linux, PID 1 is special. It’s the init process. It’s responsible for reaping “zombie” processes (processes that have finished but haven’t been acknowledged by their parent). Most applications (Node, Python, Java) are not designed to be PID 1. They don’t handle signals like SIGTERM or SIGINT correctly, and they definitely don’t reap orphans.
If you run docker run my-app, and your app spawns subprocesses, those subprocesses will eventually become zombies when they die. They’ll stay in the process table until the container is restarted. Even worse, when you run docker stop, Docker sends SIGTERM to PID 1. If your app doesn’t explicitly catch that signal, Docker waits 10 seconds and then SIGKILLs it. Your app didn’t shut down gracefully. It didn’t close database connections. It didn’t finish the last request. It just died.
The fix is tini. It’s a tiny init binary that handles all this for you.
# The right way to handle signals
RUN apk add --no-cache tini
ENTRYPOINT ["/sbin/tini", "--"]
CMD ["node", "server.js"]
Or, if you’re using a modern version of Docker, just use the --init flag. But since you’re likely deploying to Kubernetes or ECS, you need to bake it into the image or ensure your entrypoint script handles signals correctly. Speaking of entrypoint scripts, never use the “shell form” of CMD.
# WRONG: Runs as /bin/sh -c "node server.js". Signals are lost.
CMD node server.js
# RIGHT: Runs as "node server.js" directly.
CMD ["node", "server.js"]
Storage: Volumes vs. Binds
I’ve seen people lose production data because they didn’t understand the difference between a bind mount and a volume. A bind mount (-v /host/path:/container/path) is a direct link to a directory on the host. It’s great for development because you can change code on your Mac and see it reflected in the container. In production, it’s a nightmare. It creates a hard dependency on the host’s file structure. If you move the container to a different node, the data isn’t there.
Volumes (-v my-data:/container/path) are managed by Docker. They live in /var/lib/docker/volumes/. They are abstracted away. But here’s the kicker: Docker never deletes them. You run docker rm -f my-container, and that volume stays there forever. Over six months, you’ll accumulate hundreds of gigabytes of “dangling” volumes. I’ve seen this take down entire build servers.
The “Real World” Gotcha: docker system prune is your friend, but docker system prune -a --volumes is a nuclear bomb. Use it with caution. I once saw a junior dev run it on a staging server and wipe out the persistent database volumes for the entire QA team. We had backups, but the “War Room” was not a fun place to be that afternoon.
The “Real World” Gotcha: The OOM Killer
Docker doesn’t limit memory by default. If your container has a memory leak, it will consume every byte of RAM on the host until the Linux kernel’s Out-Of-Memory (OOM) Killer wakes up. The OOM Killer is a blunt instrument. It looks for the process using the most memory and kills it. Often, that’s not the leaking container. Sometimes it’s the Docker daemon itself. Sometimes it’s the SSH daemon. Suddenly, you can’t even log into the box to fix it.
Always, always, always set memory limits. And no, --memory=1g is not enough. You need to understand the difference between hard limits and soft limits (--memory-reservation). A soft limit allows the container to use more RAM if the host has it, but pushes it back down when the host is under pressure. A hard limit kills the container the moment it touches the ceiling.
# Example of a responsible container run
docker run -d \
--name api-server \
--memory="1g" \
--memory-reservation="512m" \
--cpus="1.5" \
--restart=on-failure:5 \
my-api:v1.2.3
If you’re running Java, this gets even more complicated. Older versions of the JVM (pre-8u191) don’t realize they’re in a container. They look at the host’s total RAM to calculate their heap size. If your host has 64GB of RAM and you limit the container to 2GB, the JVM will try to allocate a 16GB heap and get OOM-killed immediately. Use -XX:+UseContainerSupport if you’re stuck on older Java versions, or just upgrade to a version that isn’t from the Stone Age.
Deep Dive: The Container Runtime Interface (CRI)
If you want to sound like you know what you’re talking about at a cocktail party (or a design doc review), stop saying “Docker” when you mean “the runtime.” Docker is actually a collection of tools. When you run a container, the Docker daemon (dockerd) talks to containerd, which then uses runC to actually talk to the kernel. runC is the thing that creates the namespaces and cgroups.
Why does this matter? Because Kubernetes deprecated Docker as a container runtime years ago. K8s now talks directly to containerd via the CRI. If you’re debugging a node in a modern cluster, docker ps won’t work. You have to use crictl ps. Understanding this stack helps you realize that Docker is just a UI. The real magic is in the kernel primitives:
- Namespaces: These provide isolation.
pid(processes),net(network),mnt(filesystems),uts(hostname),ipc(inter-process communication). - Cgroups (Control Groups): These provide resource constraints. CPU, Memory, I/O, Network bandwidth.
- Capability Sets: These define what the root user can actually do. Even as root, a container usually can’t load kernel modules or change the system clock unless you give it
--privilegedaccess (which you shouldn’t). - OverlayFS: The copy-on-write filesystem that makes layers possible. It merges the “lower” read-only layers with an “upper” writable layer.
If you want to see what’s actually happening under the hood, try running strace -f -e trace=clone,unshare docker run alpine echo "hi". You’ll see the clone() syscall with a bunch of flags like CLONE_NEWPID and CLONE_NEWNET. That’s Docker’s “secret sauce.” It’s just a very polished wrapper around clone().
Troubleshooting Like a Pro
When a container is failing, docker logs is your first stop, but it’s often useless. If the container is crashing before it can even start, the logs will be empty. This is where docker inspect comes in. Look at the State object. Look for the ExitCode. An exit code of 137 means it was OOM-killed. 139 means a segmentation fault. 127 means the command wasn’t found (usually a PATH issue or a missing dependency in a slim image).
If the container is running but behaving weirdly, don’t just docker exec -it bash. That’s the lazy way. If your image is properly minimized (like Distroless), there won’t even be a shell to exec into. Instead, use nsenter. This tool allows you to enter the namespaces of a running process from the host. It’s like teleporting into the container’s brain without needing a backdoor.
# Find the PID of the container's main process
PID=$(docker inspect --format '{{ .State.Pid }}' my-container)
# Enter the network namespace to run tcpdump from the host
sudo nsenter -t $PID -n tcpdump -i eth0
This is how you debug networking issues in production without installing tcpdump inside your 10MB production image. You keep your image clean and use the host’s tools to peer inside.
The Registry Bottleneck
We need to talk about the “Registry Death Spiral.” Imagine you have a 500-node cluster. You release a new version of your app. All 500 nodes start pulling a 1GB image at the same time. That’s 500GB of data hitting your registry in a matter of seconds. If you’re using a self-hosted Harbor or a small ECR instance, you’re going to have a bad time. I’ve seen registries fall over, causing deployments to hang and eventually timing out the entire CI/CD pipeline.
The solution isn’t “bigger registry servers.” The solution is better image management. Use docker-squash or multi-stage builds to keep images under 200MB. Use a P2P image distribution tool like Uber’s Kraken or Alibaba’s Dragonfly if you’re at massive scale. But for most of us, just being smart about layers is enough. If your “base” layer (the one with all your OS libraries) changes every day, you’re doing it wrong. That layer should be stable for weeks, so nodes only have to pull the tiny “app” layer on each deploy.
The Wrap-up
Docker is a tool, not a philosophy. It’s a way to package a process so it runs predictably, but it doesn’t absolve you from the responsibility of knowing how Linux works. Stop building bloated images, start using multi-stage builds, respect the PID 1 signal handling, and for the love of all that is holy, stop running your containers as root. Docker isn’t magic; it’s just tar files and iptables rules. Treat it with the skepticism it deserves, and it might actually work when you need it to.
Stop reading tutorials and start reading the man pages for namespaces and cgroups. That’s where the real senior-level knowledge lives.
Related Articles
Explore more insights and best practices: