Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Chris Fregly

Rating No ratings yet

In today's era of ever-growing generative models, AI Systems Performance Engineering equips professionals with actionable strategies to co-optimize hardware, software, and algorithms for high-performance and cost-effective AI systems. Authored by Chris Fregly, a performance-focused engineering and product leader, this comprehensive resource transforms complex systems into streamlined, high-impact AI solutions. Whether you're an engineer, researcher, or developer, this book offers a holistic roadmap for building resilient, scalable, and cost-effective AI systems that excel in both training and inference.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# AI Systems Performance Engineering ## 【One-Line Pitch】 A practical, full-stack guide to co-optimizing hardware, software, and algorithms for AI systems—covering everything from NVIDIA's latest GPU architectures to Kubernetes networking tweaks—for engineers, researchers, and developers who want to squeeze maximum performance and cost-efficiency out of training and inference workloads. ## 【Book Arc】 - **Opening (~0%–10%)**: Establishes the "why" of AI systems performance engineering through the DeepSeek case study—a $6M training run that rivaled models costing orders of magnitude more—and introduces core principles like profile-driven optimization, holistic system views, and striving for order-of-magnitude impact rather than incremental gains. - **Early (~10%–23%)**: Dives deep into NVIDIA's hardware roadmap, explaining the Grace-Blackwell superchip architecture with its unified memory (up to 864 GB per superchip), the NVL72 rack-scale system with 72 GPUs interconnected via NVLink, and the progression toward GB300 with Blackwell Ultra GPUs pushing into exascale territory. - **Early (~23%–32%)**: Transitions from hardware to software optimization, covering CPU and OS tuning techniques—NUMA awareness, CPU pinning with taskset, memory pinning for faster DMA transfers, huge pages, and disabling C-states to eliminate latency-inducing OS interference. - **Middle (~32%–42%)**: Explores GPU sharing and partitioning strategies including Multi-Process Service (MPS) for overlapping work from multiple processes, Multi-Instance GPU (MIG) for hardware-level partitioning, and Kubernetes time-slicing—plus container optimization and host networking for multi-node GPU workloads. - **Middle (~42%–48%)**: Focuses on data flow optimization—prefetching with PyTorch's DataLoader, batching I/O operations, overlapping communication with computation using CUDA streams, and network tuning for Ethernet and InfiniBand interconnects. - **Middle (~48%–end)**: Covers inference-specific optimizations including KV-cache management for Transformer models, the distinction between compute-bound prefill and memory-bound decode stages, and advanced techniques like NVMe-based KV-cache offloading for distributed inference clusters. ## 【Key Takeaways】 - **Performance engineering is a multiplier, not a luxury** (Opening): DeepSeek's $6M training run versus billion-dollar competitors proves that clever system optimizations can upend AI economics—at scale, every bit of performance translates to millions saved. (Early) - **Unified memory architectures are game-changers** (Early): The Grace-Blackwell superchip's cache-coherent NVLink-C2C interconnect lets CPU and GPUs share nearly a terabyte of memory as one pool, eliminating explicit data copies and enabling models that previously couldn't fit on a single GPU. (Early) - **Lower precision formats double throughput at each step** (Early): FP8 doubles FP16 throughput, and Blackwell's new FP4 doubles FP8 again—one Blackwell GPU hits ~9 PFLOPS with FP4, roughly 4× its FP16 rate, making mixed-precision training a core performance lever. (Early) - **CPU tuning is as important as GPU selection** (Early): Pinning processes to cores with taskset, aligning memory allocations within NUMA nodes, using pinned memory (2–3× faster GPU transfers), and disabling C-states can eliminate the "bubbles" where GPUs sit idle waiting for data. (Early) - **GPU sharing requires matching strategy to workload** (Middle): MPS overlaps execution from multiple processes for throughput, MIG provides hardware-isolated partitions for multi-tenant environments, and Kubernetes time-slicing offers simple isolation at the cost of idle time—choose based on whether you prioritize isolation or utilization. (Middle) - **Host networking eliminates container overhead for GPU clusters** (Middle): Setting hostNetwork: true in Kubernetes lets containers access InfiniBand directly without NAT or overlay translation—critical for MPI jobs and performance-sensitive multi-node workloads. (Middle) - **Overlapping communication and computation keeps GPUs fed** (Middle): Using CUDA streams to run all-reduce gradient aggregation alongside matrix multiplications, plus prefetching data batches ahead of time, reduces idle GPU time and improves overall throughput. (Middle) - **Inference has two distinct stages requiring different optimizations** (Middle): Prefill is compute-bound (building KV-cache from prompts), while decode is memory-throughput-bound (gathering weights for token generation)—advanced systems like NVIDIA Dynamo and vLLM run these stages on separate GPUs, and KV-cache offloading to NVMe handles idle sessions. (Middle) ## 【Reading Tips】 - **Skim the hardware roadmap sections** (~10%–23%) if you're not selecting GPU infrastructure—the key takeaway is understanding how unified memory and lower precision create optimization opportunities, not memorizing specific specs. - **Deep-read the CPU/OS tuning chapter** (~29%–32%)—the practical techniques (taskset, numactl, pin_memory=True, ulimit settings) are immediately applicable to any GPU server you'll work with. - **Pay special attention to the GPU sharing section** (~32%–39%)—MPS vs. MIG vs. time-slicing is a decision you'll face in any multi-tenant environment, and the trade-offs are clearly explained. - **The inference section** (~48%+) is essential if you're deploying models for production—understanding prefill vs. decode stages and KV-cache management is critical for building efficient inference servers. - **Take away the profiling mindset**: the book repeatedly emphasizes measuring everything, trusting data over assumptions, and using profilers to identify true bottlenecks before applying targeted optimizations. ## 【Coverage Limits】 Excerpts do not cover specific code examples for implementing the discussed techniques, nor do they detail the mathematical foundations of Transformer architectures or provide step-by-step benchmarking methodologies. ##
Page 16
eepSeek claims that their DeepSeek-R1 model was trained for only $6 million in compute – an order of magnitude lower than models like GPT-4 and Gemini – whil...
View in text
Excerpt 2
l GPUs - and 36 Grace CPUs - all interconnected via NVLink. The GB200 NVL72 is built as 18 compute nodes (each 1U in size), where each node contains two GB20...
View in text
Excerpt 3
g that might introduce unpredictable latency such as excess context switching, frequency scaling, and memory-to-disk swapping. The result should be that your...
View in text
Excerpt 4
based network, you can use the Linux networking /etc/sysctl.conf parameters: net.core.rmem_max , net.core.wmem_max to set max buffer sizes, and net.ipv4.tcp_...
View in text
Excerpt 5
onal headroom to scale the number of end users even higher. With the same number of GPUs, Dynamo’s dynamic batching, GPU load balancing, and disaggregated in...
View in text
Excerpt 6
educe the learning rate” as shown in Figure 6-3. Figure 6-4. Dynamic graph execution based on input data (Source: https://developer.nvidia.com/blog/enabling-...
View in text
Excerpt 7
ectivity for multi-GPU workloads. NVLink 5 provides up to 1.8 TB/s GPU-to-GPU Ensure your GPU servers run a recent, stable Linux kernel configured for high-p...
View in text
Excerpt 8
over FP16. On newest GPUs, also explore and evaluate FP8 or INT4 for certain models to further boost throughput for inference. Fused Activation + Scaling. Wh...
View in text
Tags
AI categories
Artificial IntelligencePerformance EngineeringGPU Computing
Publish Year: 2025
Language: English
File Format: PDF
File Size: 12.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…