[72439.123456] out_of_memory: Kill process 14201 (python3.11) score 942 or sacrifice child
[72439.123501] Killed process 14201 (python3.11) total-vm:16777216kB, anon-rss:14567424kB, file-rss:0kB, shmem-rss:0kB
[72439.123550] oom_reaper: reaped process 14201 (python3.11), now anon-rss:0kB, file-rss:0kB, shmem-rss:0kB
[72439.123600] traps: python3.11[14202] general protection fault ip:7f8a2c3d4e5f sp:7f8a1b2c3d4e error:0 in libc.so.6[7f8a2c300000+1a0000]
Segmentation fault (core dumped)
$ gdb python3.11 core.14201
(gdb) bt
Table of Contents
0 0x00007f8a2c3d4e5f in _int_malloc (av=0x7f8a2c50a020 , bytes=1024) at malloc.c:4123
1 0x00007f8a2c3d6789 in __GI___libc_malloc (bytes=1024) at malloc.c:3321
2 0x00000000005a2b3c in PyObject_Malloc ()
3 0x00000000005b1f2a in _PyObject_GC_Alloc ()
4 0x00000000005f3d12 in dict_resize ()
5 0x00000000005f4e9a in PyDict_SetItem ()
6 0x0000000000612c3d in _PyEval_EvalFrameDefault ()
… (142 levels of recursion and indirection omitted for sanity)
The Abstraction Tax and the Death of Determinism
I spent the morning staring at a hex dump of a corrupted heap because some “architect” decided that writing a high-throughput socket listener in Python 3.11.4 was a good idea. We are running on Linux Kernel 6.2.0-33-generic, a kernel that is perfectly capable of handling millions of interrupts per second, yet this “python code” managed to choke the life out of a 64-core Xeon box with only 5,000 concurrent connections.
The problem started with a junior developer’s “elegant” solution for a real-time data ingestion service. He used asyncio. He used pydantic. He used every buzzword-compliant library available on PyPI. He thought he was being efficient because the code looked clean. Clean code is a lie told by people who don’t understand how L1 caches work. When you write “python code”, you aren’t writing instructions for a CPU; you are writing suggestions for a massive, bloated C program that treats your hardware like an infinite resource.
The crash above wasn’t a fluke. It was the inevitable result of heap fragmentation and the sheer weight of the PyObject structure. In C, if I want to store a 64-bit integer, I use 8 bytes of memory. In this “python code”, that same integer is a PyLongObject. It has a reference count. It has a type pointer. It has a size field. It’s 28 bytes of overhead before you even get to the actual data. Now multiply that by a few million incoming packets, and you have a recipe for the OOM killer to come knocking on your door.
The junior dev’s “python code” was creating a new dictionary for every incoming JSON payload. Think about that. Every time a packet hits the wire, the runtime has to allocate a PyDictObject, calculate hashes, handle potential collisions, and manage the internal hash table resizing. On a kernel level, this translates to a constant stream of brk and mmap syscalls. The glibc 2.36 allocator is doing its best, but it can’t keep up with the sheer volume of small, short-lived allocations.
PyObject: The Twenty-Eight Byte Anchor
Let’s talk about the PyObject_HEAD. Every single thing in this “python code” is an object. You want to increment a counter? That’s not an add instruction. That’s a call to PyNumber_Add, which checks the types of both operands, creates a new object for the result, and decrements the reference count of the old one. If the reference count hits zero, the garbage collector might decide to wake up and ruin your latency tail.
$ strace -c -p 14201
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
45.21 2.123456 15 141563 futex
22.10 1.038472 8 129847 mmap
15.45 0.725123 12 60423 brk
10.12 0.475123 5 95024 recvfrom
7.12 0.334512 11 30421 epoll_wait
------ ----------- ----------- --------- --------- ----------------
100.00 4.696686 457278 total
Look at those futex calls. That is the sound of a multi-core processor screaming in agony. Because of the Global Interpreter Lock (GIL), even in Python 3.11.4, your “python code” is essentially a single-threaded bottleneck masquerading as a modern service. Those 64 cores are sitting idle, spinning on locks, waiting for their turn to touch the interpreter state. We are paying for 128 hardware threads and using one. It’s an insult to the engineers who designed the silicon.
The junior dev argued that “asyncio is non-blocking.” He’s half-right. It’s non-blocking for I/O, but it’s completely blocking for the CPU. While the event loop is busy deserializing a massive JSON blob into a forest of PyObjects, it can’t respond to new connections. The TCP backlog fills up. The kernel starts dropping SYN packets. The “python code” thinks it’s doing work, but it’s actually just shuffling pointers around while the rest of the system starves.
The GIL: A Multi-threaded Lie for the Masses
I’ve heard the rumors about PEP 703 and the “no-GIL” builds. They say it will save us. They are wrong. Even without the GIL, the fundamental problem remains: the “python code” is built on a foundation of reference counting. In a truly multi-threaded environment, every Py_INCREF and Py_DECREF has to be an atomic operation. Do you know what atomic operations do to a CPU cache? They invalidate the cache line across every other core. You end up with “cache line bouncing,” where the cores spend more time coordinating memory access than actually executing logic.
In our production disaster, the “python code” was attempting to share a global state dictionary across multiple “worker” threads (a mistake in itself). The result was a futex storm. Every time a thread wanted to update a counter, it had to lock the entire interpreter. I watched the top output. The CPU usage was at 100%, but the actual throughput was lower than a 486 running a Perl script.
The overhead of opcode dispatch is another silent killer. In a C program, a loop is a few instructions: mov, add, cmp, jne. In this “python code”, a loop involves the virtual machine fetching an opcode, looking up the function in a dispatch table, checking the argument types, and then finally executing the operation. It’s a massive amount of work just to do nothing. When you are dealing with high-concurrency sockets, every microsecond matters. This “python code” treats microseconds like they are free.
Socket Buffers and the Futility of Asyncio
The junior dev’s code used asyncio.streams. It looks pretty. It’s “idiomatic.” It’s also a disaster for memory pressure. When a packet arrives, the kernel puts it in a socket buffer. The “python code” then copies that data into a bytes object. Then it gets decoded into a string (another copy). Then it gets parsed into a JSON object (a whole tree of new objects). By the time the “python code” actually looks at the data, it has been copied and transformed four or five times.
$ lsof -p 14201
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NODE NAME
python3.1 14201 root mem REG 253,1 2024880 135421 /usr/lib/x86_64-linux-gnu/libc.so.6
python3.1 14201 root 0u CHR 136,0 0t0 3 /dev/pts/0
python3.1 14201 root 1u CHR 136,0 0t0 3 /dev/pts/0
python3.1 14201 root 2u CHR 136,0 0t0 3 /dev/pts/0
python3.1 14201 root 3u IPv4 456789 0t0 TCP *:8080 (LISTEN)
python3.1 14201 root 4u IPv4 456790 0t0 TCP 10.0.0.1:8080->10.0.0.5:54321 (ESTABLISHED)
... (4995 more ESTABLISHED connections)
Each of those 5,000 connections has its own set of buffers in the “python code”. Because Python’s memory management is non-deterministic, these buffers don’t get freed immediately when the connection closes. They linger in the “generation 0” or “generation 1” heap until the GC decides it’s time to do a sweep. Meanwhile, the rss (Resident Set Size) of the process keeps climbing.
I checked the tcp_rmem and tcp_wmem settings on the host. The kernel was configured to allow up to 16MB per socket. The “python code”, in its infinite wisdom, was also buffering data at the application level. We were double-buffering everything. When the traffic spiked, the “python code” couldn’t process the buffers fast enough, the heap exploded, and the OOM killer finally did what I should have done weeks ago: it killed the process.
Heap Fragmentation: A Slow Death by a Thousand Mallocs
One of the most insulting things about this “python code” is how it handles memory fragmentation. In a low-level language, I can use a pool allocator or a slab allocator for fixed-size structures. I can ensure that my memory is contiguous, which is friendly to the prefetcher. Python doesn’t give you that choice. It uses a private heap for small objects (under 512 bytes) and falls back to malloc for everything else.
As the “python code” runs, it constantly creates and destroys objects of varying sizes. This leaves holes in the heap. The glibc allocator tries to coalesce these holes, but it can’t always do it, especially if there’s a long-lived object sitting right in the middle of a free block. Over time, the virtual memory space becomes a Swiss cheese of allocated and unallocated chunks. The process might only be using 2GB of actual data, but its vsz is 16GB because the heap is so fragmented it can’t find a contiguous block for a new dictionary resize.
This is exactly what happened during the crash. The “python code” tried to resize a dictionary to hold more session data. The dict_resize function called PyObject_Malloc, which called malloc, which tried to expand the heap. But the heap was so fragmented that the kernel couldn’t satisfy the request, leading to a failure that the “python code” wasn’t prepared to handle.
The junior dev suggested we just “add more RAM.” This is the modern solution to everything: throw more hardware at bad software. But you can’t outrun physics. More RAM just means a longer pause when the garbage collector finally decides to do a full “stop-the-world” sweep. I’ve seen GC pauses in this “python code” that lasted for several seconds. In a real-time system, a several-second pause is an eternity. It’s not a glitch; it’s a total system failure.
The C-API: Where the Lies Finally Collapse
When the “python code” isn’t enough, people reach for C extensions. They think they can wrap a fast C library and get the best of both worlds. But the boundary between Python and C is a minefield. You have to manually manage reference counts. You have to handle the GIL. You have to convert between Python types and C types.
The segfault in the log above happened inside libc.so.6, but it was triggered by a C extension that the “python code” was using for “performance.” The extension was passed a Python list, and it tried to iterate over it. But while it was iterating, another thread (which had managed to grab the GIL) modified the list. The C extension was left holding a dangling pointer to a PyObject that had already been deallocated.
(gdb) frame 4
#4 0x00000000005f3d12 in dict_resize (mp=0x7f8a1400a120, minsize=12288) at Objects/dictobject.c:1234
1234 new_keys = (PyDictKeysObject *)dk_alloc(newsize);
(gdb) p *mp
$1 = {ob_refcnt = 1, ob_type = 0x8c5a20 <PyDict_Type>, ma_used = 8192, ma_version_tag = 12345, ma_keys = 0x7f8a1400b000, ma_values = 0x0}
This is the reality of the “python code” ecosystem. It’s a house of cards built on top of a swamp. We spend our time debugging issues that shouldn’t even exist. We worry about “thread safety” in a language that can’t even run two threads at the same time. We worry about “memory management” in a language that hides the heap from us.
I miss the days when a “segmentation fault” meant you actually did something wrong with a pointer, not that your interpreter’s internal state was corrupted by a race condition in a third-party library. I miss the days when you could look at a line of code and know exactly what instructions the CPU was going to execute. This “python code” is a layer of fog between the engineer and the machine.
The “elegant” solution provided by the junior dev is now in the trash. I’m rewriting the core socket handling logic in C. No asyncio. No pydantic. Just a clean epoll loop, a fixed-size buffer pool, and a deterministic memory layout. The “python code” can stay for the high-level configuration logic, but it has no business being anywhere near the data plane.
We’ve forgotten what it means to be engineers. We’ve traded understanding for convenience, and performance for “velocity.” But when the server is down and the heap is a mess of corrupted PyObjects, that velocity doesn’t look so good. It looks like a car crash. And I’m the one who has to clean up the glass.
The Manifesto of Regret is simple: if you care about your hardware, if you care about your latency, and if you care about your sanity, stop trying to solve every problem with “python code”. It is a tool for scripts and glue, not for high-concurrency systems. The hardware is fast. The kernel is efficient. The “python code” is the only thing standing in the way of a working system. It’s time we stopped pretending otherwise.
I’m going back to my editor. I have a struct to define, and it’s not going to have a 28-byte header. It’s going to have exactly what it needs, and not a bit more. That’s not “cynicism.” That’s engineering.
Related Articles
Explore more insights and best practices: