Python Tutorial for Beginners: Learn to Code Step-by-Step

POST-MORTEM REPORT: INCIDENT #8842-OMEGA
DATE: 2024-05-14
TIME: 03:14:22 EST
AUTHOR: Murphy, Senior Systems Engineer (Level IV)
SUBJECT: Total Production Collapse / Kernel Panic on Node-04


1. The Incident Log

I was woken up at 03:14 AM by the sound of a pager that has been in service longer than the junior developer who caused this mess. The primary load balancer started screaming about a 502 Gateway Timeout, and by the time I logged into the terminal, Node-04 was already a brick. The kernel was gasping for air.

[ 11642.314159] python3.12 invoked oom-killer: gfp_mask=0x100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0
[ 11642.314165] CPU: 14 PID: 14208 Comm: python3.12 Tainted: G        W          6.5.0-27-generic #28-Ubuntu
[ 11642.314168] Hardware name: Dell Inc. PowerEdge R740/0W23F1, BIOS 2.11.2 04/12/2021
[ 11642.314170] Call Trace:
[ 11642.314172]  <TASK>
[ 11642.314175]  dump_stack_lvl+0x48/0x70
[ 11642.314183]  dump_header+0x4a/0x210
[ 11642.314188]  oom_kill_process+0xec/0x180
[ 11642.314193]  out_of_memory+0x11b/0x5a0
[ 11642.314201]  __alloc_pages_slowpath.constprop.0+0xad1/0xd30
[ 11642.314210]  __alloc_pages+0x32d/0x350
[ 11642.314215]  wp_page_copy+0x129/0x8a0
[ 11642.314222]  handle_mm_fault+0x92f/0xe10
[ 11642.314234]  do_user_addr_fault+0x1d7/0x630
[ 11642.314241]  exc_page_fault+0x77/0x170
[ 11642.314247]  asm_exc_page_fault+0x26/0x30
[ 11642.314252] RIP: 0033:0x7f8a2c4b3a10
[ 11642.314258] Code: ... (omitted for brevity, but it was a disaster) ...
[ 11642.314260] RSP: 002b:00007fff5fbff618 EFLAGS: 00010202
[ 11642.314263] RAX: 0000000000000000 RBX: 000055a1d4e2b010 RCX: 00007f8a2c4b3a10
[ 11642.314265] RDX: 0000000000000000 RSI: 000055a1d4e2b010 RDI: 00007f8a2c4b3a10
[ 11642.314267] Out of memory: Kill process 14208 (python3.12) score 942 or sacrifice child.
[ 11642.314275] Killed process 14208 (python3.12) total-vm:64424508kB, anon-rss:62145820kB, file-rss:0kB, shmem-rss:0kB

The Out of Memory (OOM) killer did its job, but the damage was done. Sixty-four gigabytes of virtual memory vanished into the ether because someone thought they were being “efficient.”


2. The Autopsy

I tracked down the offending script. It was a “log processor” written by our newest junior hire, Kevin. Kevin likes to talk about “expressive syntax” and “functional paradigms.” He doesn’t like to talk about L1 cache misses or memory fragmentation.

Here is the “clever” snippet of Python 3.12.2 code that brought a $20,000 server to its knees:

# Kevin's "Masterpiece"
def process_logs(filename):
    # Load everything into a list comprehension because "it's faster"
    data = [line.strip().split(',') for line in open(filename, 'r').readlines()]

    # Perform a "clever" transformation
    processed_data = [
        { "id": int(row[0]), "payload": row[1:], "meta": "PROCESSED" }
        for row in data if len(row) > 1
    ]

    return processed_data

# Triggered on a 12GB log file
results = process_logs("/var/log/heavy_traffic.log")

At first glance, a novice might think this is fine. It’s “Pythonic,” right? Wrong. It’s a suicide note. Kevin used readlines(), which sucks the entire file into memory at once. Then he created a list of lists. Then he created a list of dictionaries. Each of these “convenient” abstractions carries a massive overhead in the CPython runtime.

By the time the interpreter reached the second list comprehension, the RSS (Resident Set Size) was climbing faster than a Falcon 9. The garbage collector couldn’t keep up because the references were all still in scope. The heap exploded, the stack was irrelevant, and the kernel had to put the process out of its misery.


3. The “Python Tutorial” for the Uninitiated

Since it appears we are now hiring people who think RAM is an infinite resource provided by the cloud gods, I am forced to write this python tutorial. Pay attention, because I am only going to explain how the machine actually works once.

Variables are not Boxes

In a real language like C, a variable is a memory address. In Python, every variable is a pointer to a PyObject struct on the heap. When you write x = 42, you aren’t putting the number 42 into a slot. You are creating a full-blown object that contains:
1. ob_refcnt: A reference count (8 bytes).
2. ob_type: A pointer to the type object (8 bytes).
3. ob_value: The actual data (for an integer, this is variable-sized).

In Python 3.12.1, even a small integer takes up 28 bytes of memory. If you have a list of a million integers, you aren’t using 4MB of RAM; you are using at least 28MB, plus the overhead of the list itself, which is just an array of 8-byte pointers (0x7fff5fbff618 style).

The List: A Leaky Bucket of Pointers

A Python list is a dynamic array of pointers. When you use a list comprehension like Kevin did, Python has to pre-allocate space. If the list grows, it over-allocates to avoid constant resizing.

import sys
empty_list = []
print(sys.getsizeof(empty_list)) # 56 bytes for nothing!

When Kevin ran open(filename).readlines(), he created a list where every element was a string object. Each string object has its own overhead (PEP 393 flexible string representation). Then he split those strings into more lists. He was effectively tripling the memory footprint of the raw data before he even started “processing” it.

The Global Interpreter Lock (GIL)

Kevin tried to “fix” the speed issue by wrapping this in a threading.Thread call. This is where I almost threw my monitor out the window.

The Global Interpreter Lock (GIL) is a mutex that prevents multiple native threads from executing Python bytecodes at once. This lock is necessary because CPython’s memory management is not thread-safe. If two threads tried to increment the ob_refcnt of the same object simultaneously, you’d get a race condition that would corrupt the heap faster than you can say “segmentation fault.”

While PEP 703 is making strides toward making the GIL optional, in our current Python 3.12 environment, it is very much alive. Using threads for CPU-bound tasks in Python doesn’t make things faster; it just adds context-switching overhead and makes the OOM killer’s job easier.


4. The Refactor

If Kevin had bothered to read a python tutorial written by someone who actually understands hardware, he would have used generators. A generator doesn’t store the whole result in memory; it yields one item at a time. It’s the difference between trying to swallow a whole cow and taking one bite at a time.

The Broken Code (Kevin’s Way)

# Memory complexity: O(N) where N is file size
# Time complexity: O(N) but with massive GC pressure
def bad_way(path):
    return [line for line in open(path).readlines()] 

The Production-Ready Code (The Murphy Way)

import sys

def good_way(path):
    """
    This is how a grown-up processes data. 
    We use a generator expression to keep the memory footprint 
    near zero, regardless of file size.
    """
    try:
        with open(path, 'r', encoding='utf-8') as f:
            for line in f:
                # 'yield' makes this a generator. 
                # We never load more than one line into RAM.
                yield line.strip().split(',')
    except OSError as e:
        print(f"Hardware failure or file missing: {e}", file=sys.stderr)

# Usage
for row in good_way("/var/log/heavy_traffic.log"):
    # Process one row at a time. The CPU stays cool. 
    # The RAM stays empty. Murphy stays asleep.
    pass

By using with open(path) as f: and iterating over the file object, we utilize the buffer protocol. The file is read in chunks (usually 4KB or 8KB, depending on the OS block size), and we only ever have one line in memory at a time. This is how you handle a 100GB log file on a machine with 8GB of RAM.


5. The Hardware Reality

Let’s talk about what happens at the silicon level. When you use “clever” high-level abstractions, you are spit-shining the face of the CPU.

The Heap and the Garbage Collector

Python uses a hybrid of Reference Counting and a Cyclic Garbage Collector.
– Reference Counting: As soon as an object’s ob_refcnt hits zero, the memory is deallocated. This is deterministic and efficient.
– Cyclic GC: This handles cases where object A points to object B, and object B points to object A. The GC periodically pauses the world to find these “islands” of unreachable objects.

When Kevin created millions of short-lived dictionary objects, he triggered the Generation 0 scan of the GC constantly. This isn’t free. The CPU has to stop doing actual work to walk the linked list of all tracked objects to see what can be deleted. This causes cache thrashing.

Cache Locality

A modern CPU like the Xeon in Node-04 relies on the L1, L2, and L3 caches. These caches are fast because they store contiguous blocks of memory. Python’s memory model is the enemy of cache locality. Because everything is a pointer to a PyObject scattered across the heap, the CPU is constantly waiting for the memory controller to fetch data from the slow DDR4 RAM. This is called a pointer chase.

When you use a list comprehension to build a massive list of dictionaries, you are creating a fragmented map of memory that the CPU hates. You are forcing the TLB (Translation Lookaside Buffer) to work overtime, and you are guaranteeing that the prefetcher has no idea what data is coming next.

The Stack

In a language with discipline, we use the stack for local variables. The stack is fast, local, and automatically cleaned up. In Python, almost everything lives on the heap. Even the frames of your function calls are objects on the heap. When you write deep recursive functions or massive loops with high-level objects, you are bloating the heap and inviting the kernel to kill your process.


6. The Ultimatum

I have spent thirty years watching languages come and go. I’ve seen Perl scripts that looked like line noise and Java enterprise “beans” that consumed more memory than the Apollo 11 guidance computer. Python is a fine tool for glue code and rapid prototyping, but it is a dangerous weapon in the hands of someone who doesn’t respect the hardware.

To the junior staff:
1. Stop using readlines(). If I see it in a PR again, I will revoke your git access.
2. Understand the Iterator Protocol. Learn what __iter__ and __next__ do. If you don’t know how a for loop actually works under the hood, you shouldn’t be writing one.
3. Respect the GIL. Don’t throw threads at a problem because you’re too lazy to write efficient code. If you need true parallelism, use the multiprocessing module or, better yet, write a C extension.
4. Measure your RSS. Use psutil or memory_profiler. If your script uses more than 500MB of RAM to process a text file, you have failed as an engineer.

If you had spent five minutes on a proper python tutorial—one that actually discusses memory management and the CPython source code—instead of looking for “one-liner tricks” on the internet, I wouldn’t have been awake at 3:00 AM fixing Node-04.

The next time production goes down because of a “clever” list comprehension, I won’t just restart the service. I’ll leave it down and let you explain to the CTO why we lost $50,000 in transactions because you wanted to save three lines of code.

Go back to basics. Learn how the memory controller works. Learn how the kernel allocates pages. And for the love of all that is holy, stop loading entire databases into Python lists.

END OF REPORT
Murphy
Senior Systems Engineer
Sent from my VT100 Terminal

Related Articles

Explore more insights and best practices:

Leave a Comment