{"id":4845,"date":"2026-07-27T22:35:15","date_gmt":"2026-07-27T17:05:15","guid":{"rendered":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/"},"modified":"2026-07-27T22:35:15","modified_gmt":"2026-07-27T17:05:15","slug":"artificial-intelligence-best-practices-a-complete-guide-6","status":"publish","type":"post","link":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/","title":{"rendered":"Artificial Intelligence Best Practices: A Complete Guide"},"content":{"rendered":"<p><strong>INCIDENT POST-MORTEM REPORT #RC-9902<\/strong><br \/>\n<strong>STATUS:<\/strong> CRITICAL \/ RESOLVED (FOR NOW)<br \/>\n<strong>DATE:<\/strong> OCTOBER 24, 2023<br \/>\n<strong>LEAD SRE:<\/strong> ELIAS (SRE-01)<br \/>\n<strong>DURATION:<\/strong> 18 HOURS, 14 MINUTES<br \/>\n<strong>SUBJECT:<\/strong> THE DAY THE WEIGHTS MELTED<\/p>\n<hr \/>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_80 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6a67c3134463f\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6a67c3134463f\"  aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#I_RAW_TERMINAL_LOG_DUMP_STDOUTSTDERR\" >I. RAW TERMINAL LOG DUMP (STDOUT\/STDERR)<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#II_SUMMARY_OF_THE_DISASTER\" >II. SUMMARY OF THE DISASTER<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#III_THE_CASCADING_FAILURE_OF_THE_INFERENCE_LAYER\" >III. THE CASCADING FAILURE OF THE INFERENCE LAYER<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#IV_THE_DEPENDENCY_ABYSS_PYTHON_PYTORCH_AND_CUDA\" >IV. THE DEPENDENCY ABYSS: PYTHON, PYTORCH, AND CUDA<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#V_MANIFESTO_OF_HARD-WON_LESSONS\" >V. MANIFESTO OF HARD-WON LESSONS<\/a><ul class='ez-toc-list-level-4' ><li class='ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#1_Observability_is_Not_Optional\" >1. Observability is Not Optional<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#2_The_Myth_of_%E2%80%9CSmart%E2%80%9D_Automation\" >2. The Myth of &#8220;Smart&#8221; Automation<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#3_Memory_is_a_Finite_Resource\" >3. Memory is a Finite Resource<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-4'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#4_Version_Pinning_is_a_Blood_Oath\" >4. Version Pinning is a Blood Oath<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#VI_THE_REMEDIATION_CODE_AND_LOGIC\" >VI. THE REMEDIATION: CODE AND LOGIC<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#VII_THE_REALITY_OF_GPU_MEMORY_AND_SHARDING\" >VII. THE REALITY OF GPU MEMORY AND SHARDING<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#VIII_THE_HUMAN_ELEMENT_%E2%80%9CMAGIC%E2%80%9D_VS_ENGINEERING\" >VIII. THE HUMAN ELEMENT: &#8220;MAGIC&#8221; VS. ENGINEERING<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#IX_FINAL_RECOMMENDATION_AND_SIGN-OFF\" >IX. FINAL RECOMMENDATION AND SIGN-OFF<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#Related_Articles\" >Related Articles<\/a><\/li><\/ul><\/nav><\/div>\n<h3><span class=\"ez-toc-section\" id=\"I_RAW_TERMINAL_LOG_DUMP_STDOUTSTDERR\"><\/span>I. RAW TERMINAL LOG DUMP (STDOUT\/STDERR)<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<pre class=\"codehilite\"><code class=\"language-bash\">[2023-10-24 03:12:01] [INFO] Starting Smart-Optimizer-v2.sh...\n[2023-10-24 03:12:05] [DEBUG] Scanning GPU memory... 8x NVIDIA A100-SXM4-80GB detected.\n[2023-10-24 03:12:10] [WARN] CUDA 12.1 context initialized. Memory fragmentation at 14%.\n[2023-10-24 03:14:22] [ERROR] torch.cuda.OutOfMemoryError: Tried to allocate 12.50 GiB (GPU 0; 79.35 GiB total capacity; 64.12 GiB already allocated; 10.22 GiB free; 65.01 GiB reserved in total by PyTorch)\n[2023-10-24 03:14:22] [CRITICAL] Kernel PID 4402 (python3.11) killed by OOM-Killer.\n[2023-10-24 03:14:23] [SYSTEM] Node k8s-gpu-node-04 status: NotReady.\n[2023-10-24 03:14:25] [NET] Ingress-Nginx: 502 Bad Gateway (Upstream: 10.244.1.12:8080)\n[2023-10-24 03:14:30] [MONITOR] Alert: p99 Latency &gt; 45000ms.\n[2023-10-24 03:15:00] [SYSTEM] Cascading failure detected. Sharding logic failing on nodes 01-08.\n[2023-10-24 03:15:12] [CRITICAL] Vector database connection refused. Embedding layer offline.\n[2023-10-24 03:15:45] [DEBUG] Attempting automated rollback... FAILED. Rollback script requires Python 3.10, current environment is 3.11.4.\n[2023-10-24 03:16:00] [FATAL] Total cluster blackout. 100% packet loss on inference API.\n<\/code><\/pre>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"II_SUMMARY_OF_THE_DISASTER\"><\/span>II. SUMMARY OF THE DISASTER<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>I am writing this through a haze of three double-espressos and the smell of ozone that I\u2019m fairly certain is coming from my own brain. My &#8216;A&#8217; key is missing its cap, so I\u2019m hitting the bare switch. It\u2019s fitting. Everything is bare today.<\/p>\n<p>At 03:12 UTC, a &#8220;smart&#8221; optimization script, written by someone who clearly thinks <strong>artificial intelligence<\/strong> is a magical genie that lives in a silicon lamp, was executed on the production inference cluster. This script was designed to &#8220;dynamically re-quantize&#8221; our weights from FP16 to INT8 on-the-fly to &#8220;save costs.&#8221; <\/p>\n<p>The script was built for Python 3.11.4, utilizing PyTorch 2.1.0 and CUDA 12.1. On paper, this is a modern stack. In reality, it was a suicide pact. The script didn&#8217;t account for the way CUDA 12.1 handles memory fragmentation when dealing with large-scale vector embeddings during a sharded inference pass. It triggered a massive memory leak, which led to an OOM (Out of Memory) cascade. Because our load balancer was configured to &#8220;fail open&#8221; (don&#8217;t ask me why, I didn&#8217;t write the YAML), it kept shoveling traffic into the dying nodes, which then died harder.<\/p>\n<p>The result? Eight A100s turned into very expensive space heaters. The p99 latency didn&#8217;t just spike; it left the atmosphere. We went from 120ms to &#8220;the request will be finished when the sun expands into a red dwarf.&#8221;<\/p>\n<p>The &#8220;magic&#8221; software people keep talking about? It\u2019s just math. And when the math is wrong, the servers melt. There is no &#8220;intelligence&#8221; here\u2014only a series of increasingly desperate attempts to keep a house of cards from blowing over in a light breeze.<\/p>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"III_THE_CASCADING_FAILURE_OF_THE_INFERENCE_LAYER\"><\/span>III. THE CASCADING FAILURE OF THE INFERENCE LAYER<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The failure started in the quantization logic. When you move from FP16 to INT8, you\u2019re essentially trying to fit a gallon of water into a pint glass. If you don&#8217;t do it with surgical precision, you get &#8220;drift.&#8221; In this case, the drift detection was nonexistent. The model started outputting garbage\u2014NaNs (Not a Number) everywhere. <\/p>\n<p>When the inference engine (PyTorch 2.1.0) encountered these NaNs during the sharding process, it didn&#8217;t just error out. It tried to re-calculate the embeddings. This caused a recursive memory allocation loop. Python 3.11.4 is fast, sure, but it\u2019s not fast enough to outrun a recursive OOM kill. <\/p>\n<p>The sharding logic, which splits the model across eight GPUs, lost synchronization. One GPU would be waiting for a tensor that didn&#8217;t exist because its neighbor was already dead. This created a &#8220;zombie&#8221; state where the processes were still alive but doing nothing but consuming 100% CPU while waiting for a NCCL (NVIDIA Collective Communications Library) timeout that was set to\u2014get this\u201430 minutes.<\/p>\n<p>Here is the broken YAML that allowed this nightmare to persist. Look at the resource limits. Or rather, the lack of them.<\/p>\n<p><strong>BROKEN CONFIGURATION (k8s-inference-deploy.yaml):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">apiVersion: apps\/v1\nkind: Deployment\nmetadata:\n  name: inference-engine-v2\nspec:\n  replicas: 8\n  template:\n    spec:\n      containers:\n      - name: pytorch-inference\n        image: internal-registry\/inference:latest-3.11.4-cuda12.1\n        resources:\n          # This is where the &quot;magic&quot; happens. \n          # No hard limits on memory, just &quot;requests.&quot;\n          requests:\n            nvidia.com\/gpu: 1\n            memory: &quot;64Gi&quot;\n          # limits: \n          #   memory: &quot;80Gi&quot;  &lt;-- Commented out by a &quot;senior&quot; dev to &quot;avoid unnecessary kills&quot;\n        env:\n        - name: TORCH_CUDA_ARCH_LIST\n          value: &quot;8.0&quot;\n        - name: OPTIMIZE_ON_FLY\n          value: &quot;true&quot; # The trigger for the disaster\n<\/code><\/pre>\n<p>By commenting out the limits, the developer essentially told the Linux kernel: &#8220;Feel free to let this process eat the entire node until the hardware screams.&#8221; And it did.<\/p>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"IV_THE_DEPENDENCY_ABYSS_PYTHON_PYTORCH_AND_CUDA\"><\/span>IV. THE DEPENDENCY ABYSS: PYTHON, PYTORCH, AND CUDA<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>We need to talk about the &#8220;dependency hell&#8221; that is modern <strong>artificial intelligence<\/strong> development. We are currently pinned to Python 3.11.4. Why? Because 3.12 broke the specific version of the vector database driver we use. We are pinned to PyTorch 2.1.0 because 2.2.0 has a regression in the way it handles <code>torch.compile<\/code> for Transformer architectures. We are on CUDA 12.1 because the drivers on the host machines are managed by a different team that thinks updating a driver is a &#8220;quarterly event.&#8221;<\/p>\n<p>When you mix these specific versions, you aren&#8217;t building a platform; you&#8217;re performing alchemy. <\/p>\n<p>The &#8220;Smart-Optimizer&#8221; script tried to use a feature in PyTorch 2.1.0 that was supposedly &#8220;stable&#8221; but had a known conflict with CUDA 12.1&#8217;s memory allocator. Specifically, when the script attempted to shard the vector embeddings across the NVLink interconnect, it triggered a race condition. <\/p>\n<p>The cold starts were the final nail. When we tried to bring the nodes back up, the model weights (all 175GB of them) had to be pulled from an S3 bucket. Because the entire cluster was trying to pull at once, we saturated the NAT gateway. 18 hours. It took 18 hours because we spent 4 of them just waiting for bytes to move across a wire.<\/p>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"V_MANIFESTO_OF_HARD-WON_LESSONS\"><\/span>V. MANIFESTO OF HARD-WON LESSONS<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>If I see one more LinkedIn post about how &#8220;easy&#8221; it is to deploy these models, I am going to throw my Model M keyboard through a window. Here is the reality of SRE work in the age of large-scale models. These are not &#8220;best practices.&#8221; These are survival tactics.<\/p>\n<h4><span class=\"ez-toc-section\" id=\"1_Observability_is_Not_Optional\"><\/span>1. Observability is Not Optional<span class=\"ez-toc-section-end\"><\/span><\/h4>\n<p>You cannot monitor a model like you monitor a web server. A web server returns a 200 or a 500. A model can return a 200 OK while actually outputting a &#8220;hallucination&#8221; or a string of NaNs that crashes the downstream service. You need drift detection at the embedding level. You need to monitor the p99 latency of the <em>individual<\/em> GPU kernels, not just the API endpoint.<\/p>\n<h4><span class=\"ez-toc-section\" id=\"2_The_Myth_of_%E2%80%9CSmart%E2%80%9D_Automation\"><\/span>2. The Myth of &#8220;Smart&#8221; Automation<span class=\"ez-toc-section-end\"><\/span><\/h4>\n<p>Stop letting scripts make decisions about hardware utilization. If you want to quantize a model, do it in the CI\/CD pipeline. Test it. Benchmark it. Verify the loss in precision. Do not\u2014under any circumstances\u2014let a script &#8220;optimize&#8221; weights on a live production node. &#8220;Smart&#8221; is just another word for &#8220;unpredictable.&#8221;<\/p>\n<h4><span class=\"ez-toc-section\" id=\"3_Memory_is_a_Finite_Resource\"><\/span>3. Memory is a Finite Resource<span class=\"ez-toc-section-end\"><\/span><\/h4>\n<p>GPU memory fragmentation is the silent killer. In PyTorch 2.1.0, the caching allocator is aggressive. If you don&#8217;t manually manage the cache or set <code>PYTORCH_CUDA_ALLOC_CONF<\/code>, you will eventually hit a wall where you have 10GB free but can&#8217;t allocate a 1GB contiguous block. <\/p>\n<h4><span class=\"ez-toc-section\" id=\"4_Version_Pinning_is_a_Blood_Oath\"><\/span>4. Version Pinning is a Blood Oath<span class=\"ez-toc-section-end\"><\/span><\/h4>\n<p>You don&#8217;t &#8220;upgrade&#8221; a stack like this. You rebuild it from the ground up and test it for a month. Python 3.11.4 to 3.11.5 might sound minor, but in the world of C-extensions and CUDA kernels, it\u2019s a potential landmine.<\/p>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"VI_THE_REMEDIATION_CODE_AND_LOGIC\"><\/span>VI. THE REMEDIATION: CODE AND LOGIC<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>To prevent this from happening again, I\u2019ve implemented a mandatory logging decorator that every inference call must use. It doesn&#8217;t just log &#8220;success.&#8221; It tracks the tensor health and memory delta. If it detects a NaN or a sudden 10% jump in memory usage, it kills the process immediately before it can contaminate the rest of the cluster.<\/p>\n<p><strong>FIXED LOGGING DECORATOR (telemetry.py):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-python\">import torch\nimport functools\nimport logging\nfrom time import perf_counter\n\nlogger = logging.getLogger(&quot;SRE-Sentinel&quot;)\n\ndef monitor_inference(func):\n    @functools.wraps(func)\n    def wrapper(*args, **kwargs):\n        if not torch.cuda.is_available():\n            raise RuntimeError(&quot;CUDA not available. SRE-Sentinel blocking execution.&quot;)\n\n        start_mem = torch.cuda.memory_allocated()\n        start_time = perf_counter()\n\n        try:\n            result = func(*args, **kwargs)\n\n            # Check for &quot;Garbage Weights&quot; (NaNs)\n            if torch.isnan(result).any():\n                logger.error(&quot;CRITICAL: NaN detected in output tensor. Terminating.&quot;)\n                raise ValueError(&quot;Inference produced NaNs. Potential weight corruption.&quot;)\n\n            end_time = perf_counter()\n            end_mem = torch.cuda.memory_allocated()\n\n            latency = (end_time - start_time) * 1000\n            mem_delta = (end_mem - start_mem) \/ 1024 \/ 1024 # MB\n\n            logger.info(f&quot;Inference complete. Latency: {latency:.2f}ms | MemDelta: {mem_delta:.2f}MB&quot;)\n            return result\n\n        except Exception as e:\n            logger.critical(f&quot;Inference failure: {str(e)}&quot;)\n            # Force clear cache on failure to prevent fragmentation\n            torch.cuda.empty_cache()\n            raise e\n\n    return wrapper\n<\/code><\/pre>\n<p>Furthermore, I have updated the Prometheus rules. We are no longer looking at &#8220;average&#8221; CPU usage. We are looking at the &#8220;Stall Ratio&#8221;\u2014the amount of time the GPU is waiting for the CPU to feed it data. If the stall ratio exceeds 15% for more than 2 minutes, we trigger a PagerDuty alert.<\/p>\n<p><strong>PROMETHEUS ALERT RULE (gpu-alerts.rules):<\/strong><\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">groups:\n- name: AI_Inference_Alerts\n  rules:\n  - alert: HighGPUMemoryFragmentation\n    expr: (sum(torch_cuda_memory_reserved_bytes) - sum(torch_cuda_memory_allocated_bytes)) \/ sum(torch_cuda_memory_reserved_bytes) &gt; 0.3\n    for: 5m\n    labels:\n      severity: critical\n    annotations:\n      summary: &quot;GPU Memory Fragmentation too high on {{ $labels.instance }}&quot;\n      description: &quot;Fragmentation is above 30%. OOM imminent. Python 3.11.4 allocator failing.&quot;\n\n  - alert: LatencyP99Spike\n    expr: histogram_quantile(0.99, sum(rate(inference_latency_seconds_bucket[5m])) by (le)) &gt; 2.0\n    for: 1m\n    labels:\n      severity: warning\n    annotations:\n      summary: &quot;p99 Latency exceeds 2 seconds.&quot;\n<\/code><\/pre>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"VII_THE_REALITY_OF_GPU_MEMORY_AND_SHARDING\"><\/span>VII. THE REALITY OF GPU MEMORY AND SHARDING<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Let\u2019s talk about sharding. Everyone thinks sharding is just &#8220;splitting the model.&#8221; It\u2019s not. It\u2019s a high-speed game of hot potato played with 80GB tensors. When we shard a model across eight A100s, we are using Tensor Parallelism. Every single layer of the model is split. This means for every single forward pass, the GPUs have to talk to each other multiple times.<\/p>\n<p>In our case, the &#8220;Smart-Optimizer&#8221; script messed with the memory alignment of the shards. When PyTorch 2.1.0 tried to perform an <code>all-reduce<\/code> operation (combining the results from all GPUs), it found that the tensors were no longer aligned in memory. <\/p>\n<p>This is where the &#8220;dependency hell&#8221; becomes a literal hell. CUDA 12.1 introduced a new memory management feature that was supposed to make this faster. Instead, it interacted with Python 3.11.4\u2019s new garbage collector in a way that caused the &#8220;zombie&#8221; processes I mentioned earlier. The garbage collector tried to free a tensor that the GPU was still using for an <code>all-reduce<\/code> operation. <\/p>\n<p>The result was a kernel panic\u2014not in the OS, but in the CUDA driver itself. Once the driver panics, the only way out is a hard reboot of the physical host. Do you know how long it takes to reboot a server with 2TB of RAM and 8 GPUs? 15 minutes. 15 minutes of staring at a black screen, praying the BIOS doesn&#8217;t hang.<\/p>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"VIII_THE_HUMAN_ELEMENT_%E2%80%9CMAGIC%E2%80%9D_VS_ENGINEERING\"><\/span>VIII. THE HUMAN ELEMENT: &#8220;MAGIC&#8221; VS. ENGINEERING<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The core of the problem isn&#8217;t the code. It&#8217;s the mindset. Management treats <strong>artificial intelligence<\/strong> like a utility\u2014like electricity or water. You turn the tap, and the &#8220;intelligence&#8221; flows out. They don&#8217;t want to hear about CUDA versions or memory fragmentation. They want &#8220;seamless&#8221; (a word I hate) optimization.<\/p>\n<p>But there is nothing &#8220;seamless&#8221; about running 175-billion-parameter models. It is a gritty, manual, and often violent process of forcing hardware to do things it wasn&#8217;t designed to do. We are using GPUs\u2014chips designed to draw triangles for video games\u2014to simulate the neural pathways of a brain. It\u2019s a hack. The whole industry is a hack.<\/p>\n<p>When someone says &#8220;the model is smart,&#8221; they are lying. The model is a file full of numbers. If those numbers are INT8 instead of FP16, and you didn&#8217;t calculate the scale factors correctly, the &#8220;smart&#8221; model will tell you that the capital of France is &#8220;3.14159.&#8221;<\/p>\n<p>We spent 18 hours fixing a problem that was caused by the desire to save a few dollars on compute costs through &#8220;smart&#8221; automation. We lost more money in downtime in the first ten minutes than that script would have saved us in a decade.<\/p>\n<hr \/>\n<h3><span class=\"ez-toc-section\" id=\"IX_FINAL_RECOMMENDATION_AND_SIGN-OFF\"><\/span>IX. FINAL RECOMMENDATION AND SIGN-OFF<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The cluster is back online. The &#8220;Smart-Optimizer&#8221; script has been deleted, purged from Git history, and the person who wrote it has been banned from touching the production environment until they can explain the difference between a row-major and a column-major matrix.<\/p>\n<p>We are staying on Python 3.11.4 and PyTorch 2.1.0 for now, but we are implementing a strict &#8220;No Dynamic Quantization&#8221; policy. If you want to change the weights, you do it in a lab, not in a live environment.<\/p>\n<p><strong>HARDWARE RECOMMENDATION:<\/strong><br \/>\nFor the next expansion of the inference rack, do not buy the standard 4U chassis. We need the <strong>Supermicro AS-4124GS-TNR<\/strong>. It has the airflow capacity to handle the thermal spikes we saw during the OOM cascade. If the software is going to melt, we might as well have fans that can blow the smoke out of the room before the fire alarm goes off.<\/p>\n<p>I\u2019m going home. If anyone pings me on Slack before noon, I will delete their SSH keys.<\/p>\n<p><strong>Elias<\/strong><br \/>\nLead SRE, Inference Operations<br \/>\n<em>Sent from a keyboard with a broken &#8216;A&#8217; key and a very tired soul.<\/em><\/p>\n<hr \/>\n<p><strong>END OF REPORT<\/strong><\/p>\n<h2><span class=\"ez-toc-section\" id=\"Related_Articles\"><\/span>Related Articles<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Explore more insights and best practices:<\/p>\n<ul>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/how-to-upgrade-to-python-3-7-on-ubuntu-18-10\/\">How To Upgrade To Python 3 7 On Ubuntu 18 10<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/master-docker-compose-simplify-multi-container-workflows\/\">Master Docker Compose Simplify Multi Container Workflows<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/what-is-kubernetes-a-simple-guide-to-k8s-orchestration\/\">What Is Kubernetes A Simple Guide To K8S Orchestration<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>INCIDENT POST-MORTEM REPORT #RC-9902 STATUS: CRITICAL \/ RESOLVED (FOR NOW) DATE: OCTOBER 24, 2023 LEAD SRE: ELIAS (SRE-01) DURATION: 18 HOURS, 14 MINUTES SUBJECT: THE DAY THE WEIGHTS MELTED I. RAW TERMINAL LOG DUMP (STDOUT\/STDERR) [2023-10-24 03:12:01] [INFO] Starting Smart-Optimizer-v2.sh&#8230; [2023-10-24 03:12:05] [DEBUG] Scanning GPU memory&#8230; 8x NVIDIA A100-SXM4-80GB detected. [2023-10-24 03:12:10] [WARN] CUDA 12.1 &#8230; <a title=\"Artificial Intelligence Best Practices: A Complete Guide\" class=\"read-more\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\" aria-label=\"Read more  on Artificial Intelligence Best Practices: A Complete Guide\">Read more<\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-4845","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.0 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Artificial Intelligence Best Practices: A Complete Guide - ITSupportWale<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Artificial Intelligence Best Practices: A Complete Guide - ITSupportWale\" \/>\n<meta property=\"og:description\" content=\"INCIDENT POST-MORTEM REPORT #RC-9902 STATUS: CRITICAL \/ RESOLVED (FOR NOW) DATE: OCTOBER 24, 2023 LEAD SRE: ELIAS (SRE-01) DURATION: 18 HOURS, 14 MINUTES SUBJECT: THE DAY THE WEIGHTS MELTED I. RAW TERMINAL LOG DUMP (STDOUT\/STDERR) [2023-10-24 03:12:01] [INFO] Starting Smart-Optimizer-v2.sh... [2023-10-24 03:12:05] [DEBUG] Scanning GPU memory... 8x NVIDIA A100-SXM4-80GB detected. [2023-10-24 03:12:10] [WARN] CUDA 12.1 ... Read more\" \/>\n<meta property=\"og:url\" content=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\" \/>\n<meta property=\"og:site_name\" content=\"ITSupportWale\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-27T17:05:15+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png\" \/>\n\t<meta property=\"og:image:width\" content=\"512\" \/>\n\t<meta property=\"og:image:height\" content=\"512\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Techie\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Techie\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\"},\"author\":{\"name\":\"Techie\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\"},\"headline\":\"Artificial Intelligence Best Practices: A Complete Guide\",\"datePublished\":\"2026-07-27T17:05:15+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\"},\"wordCount\":1830,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\",\"name\":\"Artificial Intelligence Best Practices: A Complete Guide - ITSupportWale\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\"},\"datePublished\":\"2026-07-27T17:05:15+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/itsupportwale.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Artificial Intelligence Best Practices: A Complete Guide\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"name\":\"ITSupportWale\",\"description\":\"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides\",\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\",\"name\":\"itsupportwale\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"contentUrl\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"width\":1119,\"height\":144,\"caption\":\"itsupportwale\"},\"image\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\",\"name\":\"Techie\",\"sameAs\":[\"https:\/\/itsupportwale.com\",\"iswblogadmin\"],\"url\":\"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Artificial Intelligence Best Practices: A Complete Guide - ITSupportWale","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/","og_locale":"en_US","og_type":"article","og_title":"Artificial Intelligence Best Practices: A Complete Guide - ITSupportWale","og_description":"INCIDENT POST-MORTEM REPORT #RC-9902 STATUS: CRITICAL \/ RESOLVED (FOR NOW) DATE: OCTOBER 24, 2023 LEAD SRE: ELIAS (SRE-01) DURATION: 18 HOURS, 14 MINUTES SUBJECT: THE DAY THE WEIGHTS MELTED I. RAW TERMINAL LOG DUMP (STDOUT\/STDERR) [2023-10-24 03:12:01] [INFO] Starting Smart-Optimizer-v2.sh... [2023-10-24 03:12:05] [DEBUG] Scanning GPU memory... 8x NVIDIA A100-SXM4-80GB detected. [2023-10-24 03:12:10] [WARN] CUDA 12.1 ... Read more","og_url":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/","og_site_name":"ITSupportWale","article_publisher":"https:\/\/www.facebook.com\/Itsupportwale-298547177495978","article_published_time":"2026-07-27T17:05:15+00:00","og_image":[{"width":512,"height":512,"url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png","type":"image\/png"}],"author":"Techie","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Techie","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#article","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/"},"author":{"name":"Techie","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d"},"headline":"Artificial Intelligence Best Practices: A Complete Guide","datePublished":"2026-07-27T17:05:15+00:00","mainEntityOfPage":{"@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/"},"wordCount":1830,"commentCount":0,"publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/","url":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/","name":"Artificial Intelligence Best Practices: A Complete Guide - ITSupportWale","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/#website"},"datePublished":"2026-07-27T17:05:15+00:00","breadcrumb":{"@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-a-complete-guide-6\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/itsupportwale.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Artificial Intelligence Best Practices: A Complete Guide"}]},{"@type":"WebSite","@id":"https:\/\/itsupportwale.com\/blog\/#website","url":"https:\/\/itsupportwale.com\/blog\/","name":"ITSupportWale","description":"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides","publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/itsupportwale.com\/blog\/#organization","name":"itsupportwale","url":"https:\/\/itsupportwale.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","contentUrl":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","width":1119,"height":144,"caption":"itsupportwale"},"image":{"@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/Itsupportwale-298547177495978"]},{"@type":"Person","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d","name":"Techie","sameAs":["https:\/\/itsupportwale.com","iswblogadmin"],"url":"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/"}]}},"_links":{"self":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4845","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/comments?post=4845"}],"version-history":[{"count":0,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4845\/revisions"}],"wp:attachment":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/media?parent=4845"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/categories?post=4845"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/tags?post=4845"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}