{"id":4841,"date":"2026-07-23T22:11:14","date_gmt":"2026-07-23T16:41:14","guid":{"rendered":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/"},"modified":"2026-07-23T22:11:14","modified_gmt":"2026-07-23T16:41:14","slug":"artificial-intelligence-best-practices-guide-2","status":"publish","type":"post","link":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/","title":{"rendered":"Artificial Intelligence Best Practices &#8211; Guide"},"content":{"rendered":"<p>It\u2019s 4:14 AM. The pager went off three days ago and I haven&#8217;t seen sunlight since. If I see one more &#8216;hallucination&#8217; treated as a feature, I\u2019m quitting.<\/p>\n<p>The following is the internal post-mortem for the &#8220;Project Chimera&#8221; collapse that wiped out our production inference cluster across three regions. If you are a junior dev looking for a &#8220;quick fix,&#8221; go back to Stack Overflow. This is for the people who have to clean up the blood.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_80 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<label for=\"ez-toc-cssicon-toggle-item-6a62b6824c62d\" class=\"ez-toc-cssicon-toggle-label\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/label><input type=\"checkbox\"  id=\"ez-toc-cssicon-toggle-item-6a62b6824c62d\"  aria-label=\"Toggle\" \/><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#I_INCIDENT_LOG_TRACEBACKS_AND_KERNEL_PANICS\" >I. INCIDENT LOG: TRACEBACKS AND KERNEL PANICS<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#II_ROOT_CAUSE_ANALYSIS_THE_DRIFT_OF_THE_UNMONITORED_STOCHASTIC_PARROT\" >II. ROOT CAUSE ANALYSIS: THE DRIFT OF THE UNMONITORED STOCHASTIC PARROT<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#III_VECTOR_DATABASE_OPTIMIZATION_BEYOND_THE_DEFAULT_INDEX\" >III. VECTOR DATABASE OPTIMIZATION: BEYOND THE DEFAULT INDEX<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#Hard-Learned_Rule_1_Tune_your_HNSW_Parameters\" >Hard-Learned Rule #1: Tune your HNSW Parameters<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#Hard-Learned_Rule_2_Index_Compaction_and_Garbage_Collection\" >Hard-Learned Rule #2: Index Compaction and Garbage Collection<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#IV_QUANTIZATION_STRATEGIES_AND_GPU_MEMORY_FRAGMENTATION\" >IV. QUANTIZATION STRATEGIES AND GPU MEMORY FRAGMENTATION<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#The_4-bit_vs_8-bit_Tradeoff\" >The 4-bit vs 8-bit Tradeoff<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#GPU_Memory_Management\" >GPU Memory Management<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#V_DATA_SANITIZATION_AND_THE_FALLACY_OF_THE_RAW_PROMPT\" >V. DATA SANITIZATION AND THE FALLACY OF THE RAW PROMPT<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#The_Sanitization_Pipeline\" >The Sanitization Pipeline<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#VI_OBSERVABILITY_METRICS_THAT_ACTUALLY_MEAN_SOMETHING\" >VI. OBSERVABILITY: METRICS THAT ACTUALLY MEAN SOMETHING<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#Critical_Metrics_to_Track\" >Critical Metrics to Track:<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#Prometheus_Configuration_for_vLLM\" >Prometheus Configuration for vLLM<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#VII_RATE_LIMITING_AND_THE_%E2%80%9CCIRCUIT_BREAKER%E2%80%9D_PATTERN\" >VII. RATE LIMITING AND THE &#8220;CIRCUIT BREAKER&#8221; PATTERN<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-15\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#Docker_Compose_for_a_Resilient_Inference_Node\" >Docker Compose for a Resilient Inference Node<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-16\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#VIII_THE_FALLACY_OF_%E2%80%9CAUTO-SCALING%E2%80%9D_AI\" >VIII. THE FALLACY OF &#8220;AUTO-SCALING&#8221; AI<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-17\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#IX_MODEL_VERSIONING_AND_SHADOW_DEPLOYMENTS\" >IX. MODEL VERSIONING AND SHADOW DEPLOYMENTS<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-18\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#X_CONCLUSION_STOP_TREATING_IT_LIKE_MAGIC\" >X. CONCLUSION: STOP TREATING IT LIKE MAGIC<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-19\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#Related_Articles\" >Related Articles<\/a><\/li><\/ul><\/nav><\/div>\n<h2><span class=\"ez-toc-section\" id=\"I_INCIDENT_LOG_TRACEBACKS_AND_KERNEL_PANICS\"><\/span>I. INCIDENT LOG: TRACEBACKS AND KERNEL PANICS<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The failure started at 02:14 UTC on Tuesday. It wasn&#8217;t a sudden crash; it was a slow, agonizing degradation of the inference service that eventually cascaded into a total cluster lockout.<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\"># Extract from \/var\/log\/k8s\/inference-pod-a7xf2.log\n[2023-10-24 02:14:08] INFO: Starting inference request 0x9f22\n[2023-10-24 02:14:10] ERROR: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 512.00 MiB (GPU 0; 40.00 GiB total capacity; 38.21 GiB already allocated; 128.00 MiB free; 38.50 GiB reserved in total by PyTorch) If reserved memory is &gt;&gt; allocated memory try setting max_split_size_mb to avoid fragmentation.\n[2023-10-24 02:14:10] CRITICAL: Service terminated with exit code 139 (Segmentation Fault)\n[2023-10-24 02:14:12] DEBUG: K8s Liveness probe failed. Restarting container...\n[2023-10-24 02:14:45] WARN: Node gke-gpu-pool-92bc-x1 is under heavy memory pressure.\n[2023-10-24 02:15:01] FATAL: kernel: [192834.12] oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=\/,mems_allowed=0,global_oom,task_memcg=\/kubepods\/besteffort\/pod...,task=python3,pid=4122,uid=1000\n<\/code><\/pre>\n<p>I ran a quick <code>grep<\/code> across the logs to see how widespread this was. It was everywhere.<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\"># Checking for the frequency of OOM kills across the namespace\nkubectl logs -n ml-prod -l app=inference-engine --tail=10000 | grep -c &quot;OutOfMemoryError&quot;\n&gt; 4,821\n\n# Checking the status of the vector DB sidecar\nkubectl get pods -n ml-prod | grep &quot;vector-db-proxy&quot; | awk '{print $3}' | uniq -c\n&gt; 142 CrashLoopBackOff\n&gt; 12 Running\n<\/code><\/pre>\n<p>The system was eating itself. The &#8220;artificial intelligence&#8221; we spent six months building turned into a recursive loop of garbage data that bloated the KV cache until the GPUs choked.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"II_ROOT_CAUSE_ANALYSIS_THE_DRIFT_OF_THE_UNMONITORED_STOCHASTIC_PARROT\"><\/span>II. ROOT CAUSE ANALYSIS: THE DRIFT OF THE UNMONITORED STOCHASTIC PARROT<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The marketing department calls this &#8220;artificial intelligence,&#8221; but let\u2019s be honest: it\u2019s a fragile stack of Python libraries, C++ bindings, and unoptimized CUDA kernels held together by hope. <\/p>\n<p>The root cause was a silent failure in our data ingestion pipeline. We are running <code>Transformers 4.38.0<\/code> on <code>PyTorch 2.2.1<\/code> with <code>CUDA 12.1<\/code>. Last Friday, the upstream &#8220;User Feedback&#8221; schema was modified without a corresponding update to our preprocessing script. Specifically, a field that used to be a truncated string was replaced with a raw, multi-megabyte JSON blob.<\/p>\n<p>Our model, a quantized 4-bit Llama-2-70b variant running via <code>bitsandbytes<\/code>, wasn&#8217;t configured with hard token limits on the input side. It tried to ingest the entire JSON blob into the context window. Because we were using <code>flash-attention-2<\/code>, the memory growth was supposed to be linear, but the sheer volume of concurrent requests with 32k+ tokens caused the GPU memory fragmentation to skyrocket.<\/p>\n<p>Furthermore, we encountered a massive &#8220;model drift&#8221; issue. The &#8220;artificial intelligence&#8221; started generating repetitive, high-entropy tokens because the input distribution shifted so far from the training set. These high-entropy outputs caused the <code>vLLM<\/code> scheduler to miscalculate the required block allocation in the PagedAttention mechanism, leading to the segmentation faults seen in the logs.<\/p>\n<p>We treated the model as a black box. We didn&#8217;t monitor the Kolmogorov-Smirnov test results for our input features. We didn&#8217;t have a circuit breaker for token counts. We just assumed the &#8220;intelligence&#8221; would handle it. It didn&#8217;t. It just died.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"III_VECTOR_DATABASE_OPTIMIZATION_BEYOND_THE_DEFAULT_INDEX\"><\/span>III. VECTOR DATABASE OPTIMIZATION: BEYOND THE DEFAULT INDEX<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We are using Milvus 2.3.10 for our vector embeddings. The outage was exacerbated by the fact that our vector search latency went from 40ms to 12s. Why? Because someone thought it was a good idea to use the default HNSW (Hierarchical Navigable Small World) parameters without considering the scale of our metadata.<\/p>\n<p>When the model drift occurred, the queries being sent to Milvus became increasingly nonsensical, hitting cold areas of the index.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Hard-Learned_Rule_1_Tune_your_HNSW_Parameters\"><\/span>Hard-Learned Rule #1: Tune your HNSW Parameters<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Stop using the defaults. For a production-grade Milvus or Pinecone deployment, you must explicitly define your <code>M<\/code> (max degree of the node) and <code>efConstruction<\/code> (search scope during index construction).<\/p>\n<pre class=\"codehilite\"><code class=\"language-python\"># Example of a sane Milvus index configuration\nindex_params = {\n    &quot;metric_type&quot;: &quot;IP&quot;, # Inner Product for normalized embeddings\n    &quot;index_type&quot;: &quot;HNSW&quot;,\n    &quot;params&quot;: {\n        &quot;M&quot;: 16,             # Range 4-64. Higher = more memory, better accuracy\n        &quot;efConstruction&quot;: 128 # Range 8-512. Higher = slower build, better search\n    }\n}\n<\/code><\/pre>\n<p>If you are using Pinecone, ensure you are not saturating your pod&#8217;s IOPS by sending massive metadata payloads. We found that stripping the metadata and storing it in a sidecar Redis instance (version 7.2.4) reduced our vector search latency by 60%. The &#8220;artificial intelligence&#8221; doesn&#8217;t need to see the raw JSON; it only needs the vector ID and the pre-sanitized context.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Hard-Learned_Rule_2_Index_Compaction_and_Garbage_Collection\"><\/span>Hard-Learned Rule #2: Index Compaction and Garbage Collection<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Milvus doesn&#8217;t just &#8220;handle&#8221; deleted vectors. If you are constantly updating embeddings due to model retraining or drift correction, you must trigger manual compactions. We had 40% &#8220;tombstone&#8221; data in our segments, which forced the search engine to scan dead memory.<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\"># Manual compaction trigger via Milvus CLI\nmilvus_cli &gt; compact -c collection_chimera_v2\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"IV_QUANTIZATION_STRATEGIES_AND_GPU_MEMORY_FRAGMENTATION\"><\/span>IV. QUANTIZATION STRATEGIES AND GPU MEMORY FRAGMENTATION<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We were running 4-bit quantization using <code>bitsandbytes<\/code> to save on VRAM. This is a trap if you don&#8217;t understand how <code>NF4<\/code> (NormalFloat 4) works. While it reduces the model footprint, the dequantization step during the forward pass happens in <code>FP16<\/code> or <code>BF16<\/code>. <\/p>\n<p>If your <code>PYTORCH_CUDA_ALLOC_CONF<\/code> is not tuned, PyTorch will keep hold of these intermediate buffers, leading to the &#8220;CUDA Out of Memory&#8221; errors we saw.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_4-bit_vs_8-bit_Tradeoff\"><\/span>The 4-bit vs 8-bit Tradeoff<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>In our post-mortem testing, we found that 8-bit quantization (<code>int8<\/code>) is significantly more stable for long-running inference services than 4-bit. The 4-bit weights are too sensitive to high-variance inputs. If you must use 4-bit, you need to implement a strict <code>max_new_tokens<\/code> limit and a <code>max_input_length<\/code> validator at the API gateway level.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"GPU_Memory_Management\"><\/span>GPU Memory Management<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Add this to your environment variables immediately. Do not pass go. Do not collect $200.<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\">export PYTORCH_CUDA_ALLOC_CONF=&quot;max_split_size_mb:128,garbage_collection_threshold:0.8&quot;\n<\/code><\/pre>\n<p>This forces PyTorch to be more aggressive about releasing memory blocks. Without this, the &#8220;artificial intelligence&#8221; will eventually fragment your VRAM until even a small 512-token request fails.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"V_DATA_SANITIZATION_AND_THE_FALLACY_OF_THE_RAW_PROMPT\"><\/span>V. DATA SANITIZATION AND THE FALLACY OF THE RAW PROMPT<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The most embarrassing part of this outage was the lack of basic input validation. We allowed raw user input to be concatenated into a f-string and sent to the model. This isn&#8217;t just a security risk (prompt injection); it\u2019s a stability nightmare.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_Sanitization_Pipeline\"><\/span>The Sanitization Pipeline<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Every input must pass through a multi-stage pipeline before it touches the <code>Transformers<\/code> tokenizer:<br \/>\n1.  <strong>Length Validation:<\/strong> Reject anything over a hard character limit.<br \/>\n2.  <strong>Regex Scrubbing:<\/strong> Remove non-printable characters and excessive whitespace that can confuse the tokenizer&#8217;s BPE (Byte Pair Encoding) logic.<br \/>\n3.  <strong>Schema Enforcement:<\/strong> Use Pydantic (version 2.6.1) to ensure the input matches the expected structure.<\/p>\n<pre class=\"codehilite\"><code class=\"language-python\">from pydantic import BaseModel, Field, validator\n\nclass InferenceRequest(BaseModel):\n    prompt: str = Field(..., max_length=4096)\n    temperature: float = Field(0.7, ge=0.0, le=1.5)\n\n    @validator('prompt')\n    def no_garbage_tokens(cls, v):\n        if &quot;[[REDACTED_INTERNAL_KEY]]&quot; in v:\n            raise ValueError(&quot;Potential leak detected&quot;)\n        return v.strip()\n<\/code><\/pre>\n<p>If the &#8220;artificial intelligence&#8221; receives garbage, it produces garbage. If it produces garbage at 100 requests per second, it crashes the cluster.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"VI_OBSERVABILITY_METRICS_THAT_ACTUALLY_MEAN_SOMETHING\"><\/span>VI. OBSERVABILITY: METRICS THAT ACTUALLY MEAN SOMETHING<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Our Grafana dashboards were full of useless metrics like &#8220;CPU Usage&#8221; and &#8220;Disk I\/O.&#8221; These are irrelevant for LLM ops. You need to monitor the internals of the model&#8217;s behavior.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Critical_Metrics_to_Track\"><\/span>Critical Metrics to Track:<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<ol>\n<li><strong>KV Cache Utilization:<\/strong> If this hits 90%, start dropping requests.<\/li>\n<li><strong>Token Throughput (Tokens\/Sec):<\/strong> A sudden drop indicates bottlenecking in the attention kernels.<\/li>\n<li><strong>Per-Token Latency (P99):<\/strong> Not just total request latency. You need to know how long each token takes to generate to identify &#8220;stuck&#8221; sequences.<\/li>\n<li><strong>Output Entropy:<\/strong> This is the most important metric for detecting model drift. If the average log-probability of the generated tokens starts to diverge significantly from your baseline, the model is &#8220;hallucinating&#8221; or looping.<\/li>\n<\/ol>\n<h3><span class=\"ez-toc-section\" id=\"Prometheus_Configuration_for_vLLM\"><\/span>Prometheus Configuration for vLLM<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>If you are using <code>vLLM 0.3.3<\/code>, ensure you are scraping the <code>\/metrics<\/code> endpoint and looking for <code>vllm:avg_generation_throughput<\/code>.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\"># prometheus-config.yaml snippet\n- job_name: 'vllm-inference'\n  static_configs:\n    - targets: ['inference-service.ml-prod.svc.cluster.local:8000']\n  metrics_path: \/metrics\n  relabel_configs:\n    - source_labels: [__address__]\n      target_label: instance\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"VII_RATE_LIMITING_AND_THE_%E2%80%9CCIRCUIT_BREAKER%E2%80%9D_PATTERN\"><\/span>VII. RATE LIMITING AND THE &#8220;CIRCUIT BREAKER&#8221; PATTERN<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We failed because we didn&#8217;t have backpressure. When the inference service slowed down, the upstream services just queued more requests, leading to a classic thundering herd problem.<\/p>\n<p>You must implement a token-bucket rate limiter that is aware of the <em>token count<\/em>, not just the <em>request count<\/em>. A single request with 4,000 tokens is more expensive than 40 requests with 10 tokens.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Docker_Compose_for_a_Resilient_Inference_Node\"><\/span>Docker Compose for a Resilient Inference Node<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>This is how the node should have been configured. Note the resource limits and the health checks that actually check the model&#8217;s responsiveness, not just the HTTP port.<\/p>\n<pre class=\"codehilite\"><code class=\"language-yaml\">version: '3.8'\nservices:\n  inference-engine:\n    image: vllm\/vllm-openai:v0.3.3\n    command: &gt;\n      --model \/models\/llama-2-70b-hf\n      --quantization bitsandbytes\n      --max-model-len 8192\n      --gpu-memory-utilization 0.90\n    deploy:\n      resources:\n        reservations:\n          devices:\n            - driver: nvidia\n              device_ids: ['0']\n              capabilities: [gpu]\n    healthcheck:\n      test: [&quot;CMD-SHELL&quot;, &quot;curl -f http:\/\/localhost:8000\/v1\/models || exit 1&quot;]\n      interval: 30s\n      timeout: 10s\n      retries: 3\n    environment:\n      - CUDA_VISIBLE_DEVICES=0\n      - PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:128\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"VIII_THE_FALLACY_OF_%E2%80%9CAUTO-SCALING%E2%80%9D_AI\"><\/span>VIII. THE FALLACY OF &#8220;AUTO-SCALING&#8221; AI<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>K8s Horizontal Pod Autoscaler (HPA) is useless for &#8220;artificial intelligence&#8221; workloads if you base it on CPU or Memory. By the time the memory usage spikes, the GPU is already OOM. <\/p>\n<p>You need to scale based on <strong>Available KV Cache Slots<\/strong>. We are now writing a custom metrics adapter that queries the <code>vLLM<\/code> engine for <code>vllm:num_requests_running<\/code>. If the number of running requests exceeds 80% of our total slot capacity, we spin up a new node. If there are no more nodes available, we return a <code>503 Service Unavailable<\/code> with a <code>Retry-After<\/code> header. It is better to fail gracefully than to take down the entire cluster.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"IX_MODEL_VERSIONING_AND_SHADOW_DEPLOYMENTS\"><\/span>IX. MODEL VERSIONING AND SHADOW DEPLOYMENTS<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We pushed the updated quantization config directly to production. This was a mistake. From now on, no model change\u2014not even a change in the <code>temperature<\/code> parameter\u2014goes live without a 24-hour &#8220;Shadow Deployment.&#8221;<\/p>\n<p>In a shadow deployment, the production traffic is mirrored to the new model version. We compare the outputs. If the Kullback-Leibler (KL) divergence between the production model and the shadow model exceeds a threshold, the deployment is automatically aborted.<\/p>\n<pre class=\"codehilite\"><code class=\"language-bash\"># Example of mirroring traffic using Istio\napiVersion: networking.istio.io\/v1alpha3\nkind: VirtualService\nmetadata:\n  name: inference-mirror\nspec:\n  hosts:\n    - inference-service\n  http:\n  - route:\n    - destination:\n        host: inference-service-v1\n      weight: 100\n    mirror:\n      host: inference-service-v2-shadow\n    mirror_percentage:\n      value: 10.0\n<\/code><\/pre>\n<h2><span class=\"ez-toc-section\" id=\"X_CONCLUSION_STOP_TREATING_IT_LIKE_MAGIC\"><\/span>X. CONCLUSION: STOP TREATING IT LIKE MAGIC<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The failure of this &#8220;artificial intelligence&#8221; implementation was entirely predictable given the lack of input sanitization and the disregard for low-level GPU memory management. We treated the model like a web server. It isn&#8217;t. It\u2019s a high-performance computing workload that is extremely sensitive to input variance.<\/p>\n<p>If you are going to work on this stack, you need to understand the difference between <code>FP16<\/code> and <code>BF16<\/code>. You need to know why <code>PagedAttention<\/code> is better than standard attention. You need to be able to read a CUDA traceback.<\/p>\n<p>I\u2019m going to sleep for four hours. If the pager goes off because someone changed a hyperparameter in production without testing it, don&#8217;t bother calling me. Just start updating your resume.<\/p>\n<p><strong>Document Version:<\/strong> 1.0.4<br \/>\n<strong>Author:<\/strong> SRE Lead (Cluster 4)<br \/>\n<strong>Status:<\/strong> Exhausted.<br \/>\n<strong>Approved for internal distribution only.<\/strong><\/p>\n<h2><span class=\"ez-toc-section\" id=\"Related_Articles\"><\/span>Related Articles<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Explore more insights and best practices:<\/p>\n<ul>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/tuning-nextcloud-for-better-performance\/\">Tuning Nextcloud For Better Performance<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/what-is-machine-learning-guide\/\">What Is Machine Learning Guide<\/a><\/li>\n<li><a href=\"https:\/\/itsupportwale.com\/blog\/cybersecurity-best-practices-guide\/\">Cybersecurity Best Practices Guide<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>It\u2019s 4:14 AM. The pager went off three days ago and I haven&#8217;t seen sunlight since. If I see one more &#8216;hallucination&#8217; treated as a feature, I\u2019m quitting. The following is the internal post-mortem for the &#8220;Project Chimera&#8221; collapse that wiped out our production inference cluster across three regions. If you are a junior dev &#8230; <a title=\"Artificial Intelligence Best Practices &#8211; Guide\" class=\"read-more\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\" aria-label=\"Read more  on Artificial Intelligence Best Practices &#8211; Guide\">Read more<\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-4841","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.0 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Artificial Intelligence Best Practices - Guide - ITSupportWale<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Artificial Intelligence Best Practices - Guide - ITSupportWale\" \/>\n<meta property=\"og:description\" content=\"It\u2019s 4:14 AM. The pager went off three days ago and I haven&#8217;t seen sunlight since. If I see one more &#8216;hallucination&#8217; treated as a feature, I\u2019m quitting. The following is the internal post-mortem for the &#8220;Project Chimera&#8221; collapse that wiped out our production inference cluster across three regions. If you are a junior dev ... Read more\" \/>\n<meta property=\"og:url\" content=\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\" \/>\n<meta property=\"og:site_name\" content=\"ITSupportWale\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-23T16:41:14+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png\" \/>\n\t<meta property=\"og:image:width\" content=\"512\" \/>\n\t<meta property=\"og:image:height\" content=\"512\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Techie\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Techie\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"10 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#article\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\"},\"author\":{\"name\":\"Techie\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\"},\"headline\":\"Artificial Intelligence Best Practices &#8211; Guide\",\"datePublished\":\"2026-07-23T16:41:14+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\"},\"wordCount\":1524,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\",\"name\":\"Artificial Intelligence Best Practices - Guide - ITSupportWale\",\"isPartOf\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\"},\"datePublished\":\"2026-07-23T16:41:14+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/itsupportwale.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Artificial Intelligence Best Practices &#8211; Guide\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#website\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"name\":\"ITSupportWale\",\"description\":\"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides\",\"publisher\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#organization\",\"name\":\"itsupportwale\",\"url\":\"https:\/\/itsupportwale.com\/blog\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"contentUrl\":\"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png\",\"width\":1119,\"height\":144,\"caption\":\"itsupportwale\"},\"image\":{\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/Itsupportwale-298547177495978\"]},{\"@type\":\"Person\",\"@id\":\"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d\",\"name\":\"Techie\",\"sameAs\":[\"https:\/\/itsupportwale.com\",\"iswblogadmin\"],\"url\":\"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Artificial Intelligence Best Practices - Guide - ITSupportWale","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/","og_locale":"en_US","og_type":"article","og_title":"Artificial Intelligence Best Practices - Guide - ITSupportWale","og_description":"It\u2019s 4:14 AM. The pager went off three days ago and I haven&#8217;t seen sunlight since. If I see one more &#8216;hallucination&#8217; treated as a feature, I\u2019m quitting. The following is the internal post-mortem for the &#8220;Project Chimera&#8221; collapse that wiped out our production inference cluster across three regions. If you are a junior dev ... Read more","og_url":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/","og_site_name":"ITSupportWale","article_publisher":"https:\/\/www.facebook.com\/Itsupportwale-298547177495978","article_published_time":"2026-07-23T16:41:14+00:00","og_image":[{"width":512,"height":512,"url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2021\/05\/android-chrome-512x512-1.png","type":"image\/png"}],"author":"Techie","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Techie","Est. reading time":"10 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#article","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/"},"author":{"name":"Techie","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d"},"headline":"Artificial Intelligence Best Practices &#8211; Guide","datePublished":"2026-07-23T16:41:14+00:00","mainEntityOfPage":{"@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/"},"wordCount":1524,"commentCount":0,"publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/","url":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/","name":"Artificial Intelligence Best Practices - Guide - ITSupportWale","isPartOf":{"@id":"https:\/\/itsupportwale.com\/blog\/#website"},"datePublished":"2026-07-23T16:41:14+00:00","breadcrumb":{"@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/itsupportwale.com\/blog\/artificial-intelligence-best-practices-guide-2\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/itsupportwale.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Artificial Intelligence Best Practices &#8211; Guide"}]},{"@type":"WebSite","@id":"https:\/\/itsupportwale.com\/blog\/#website","url":"https:\/\/itsupportwale.com\/blog\/","name":"ITSupportWale","description":"Tips, Tricks, Fixed-Errors, Tutorials &amp; Guides","publisher":{"@id":"https:\/\/itsupportwale.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/itsupportwale.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/itsupportwale.com\/blog\/#organization","name":"itsupportwale","url":"https:\/\/itsupportwale.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","contentUrl":"https:\/\/itsupportwale.com\/blog\/wp-content\/uploads\/2023\/09\/cropped-Logo-trans-without-slogan.png","width":1119,"height":144,"caption":"itsupportwale"},"image":{"@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/Itsupportwale-298547177495978"]},{"@type":"Person","@id":"https:\/\/itsupportwale.com\/blog\/#\/schema\/person\/8c5a2b3d36396e0a8fd91ec8242fd46d","name":"Techie","sameAs":["https:\/\/itsupportwale.com","iswblogadmin"],"url":"https:\/\/itsupportwale.com\/blog\/author\/iswblogadmin\/"}]}},"_links":{"self":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4841","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/comments?post=4841"}],"version-history":[{"count":0,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/posts\/4841\/revisions"}],"wp:attachment":[{"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/media?parent=4841"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/categories?post=4841"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/itsupportwale.com\/blog\/wp-json\/wp\/v2\/tags?post=4841"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}