Share E-Book

Quantization and Fast Inference MEAP V01 (Vivek Kalyanarangan)(Z-Library)

Author

Programming
Language English

Today's AI models demand a lot of memory, compute, and server horsepower—which quickly translates into cost. Quantization and Fast Inference show you how you can optimize AI models without architectural redesigns or task-specific compression. It reveals practical techniques for quantization, systematically reducing numerical precision to achieve faster inference, lower memory usage, and cheaper deployment—all with minimal accuracy loss. From quantization fundamentals to runtime packaging, the book gives you a complete and comprehensive overview of the full quantization pipeline. It starts by deriving quantization mapping from first principles, and then builds your knowledge and skill through techniques for production-tested PTQ and QAT workflows and a fully compressed deployment. You'll learn to apply post-training quantization to production models, run quantization-aware training using fake quantization and straight-through estimators, and handle subtle tradeoffs like activation outliers in LLMs, KV cache pressure, and sub-8-bit formats like NF4 and FP4. what's inside Applying post-training quantization to production models Deploying efficiently on CPUs, edge devices, and mobile Framework-agnostic techniques and real cross-framework parity testing Flowcharts and checklists for efficient decision making about the reader For ML engineers and researchers experienced in Python.

Format PDF
Size 8.1 MB
9
Views
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
MEAP Edition Manning Early Access Program Quantization and Fast Inference A practitioner’s guide to eficient AI Version 1 Copyright 2026 Manning Publications For more information on this and other Manning titles go to manning.com. © Manning Publications Co. To comment go to liveBook
Page 3
welcome Thank you for purchasing the MEAP for Quantization and Fast Inference. To get the most out of this book, you'll want to be comfortable with Python and PyTorch, and have built and trained a few neural networks. Some exposure to GPU execution will help, but you don't need to be a kernel author. The book is written for ML engineers, infrastructure engineers, and applied researchers with roughly two to six years of experience. If you've shipped a model to a real system and felt the weight of its latency, memory footprint, or serving cost, you're in the right place. Quantization has become one of the load-bearing techniques of modern AI. A 70B- parameter model in FP16 is an expensive thing to serve; the same model in 4-bit can fit on a single consumer GPU and run fast enough to hold a conversation. The gap between "this model is interesting" and "this model is deployable" is, more often than not, a quantization problem. And yet most of the material on the subject is scattered across arXiv papers, framework-specific tutorials, and blog posts that hand-wave the math and the numerical gotchas that actually bite you in production. I wrote this book to fill that gap. Every technique in these pages is grounded in real measurements from real models running on real hardware — models you and I both use at work, not toy tensors. When you see a perplexity number, a throughput figure, or a memory breakdown, it comes out of a companion script that you can run and modify yourself. The Github repository that accompanies the book is where the prose meets the silicon: if something in the text makes you skeptical, you can re-run the experiment and check. I think this is the only honest way to write about quantization, because the field is full of claims that sound right and fall apart the moment you profile them. The book builds up from the foundations — fixed-point arithmetic, affine mappings, the choices that govern how you map a continuous weight distribution onto a discrete grid — through the standard toolkit of post-training quantization and quantization-aware training, and into the techniques that define the large language model era. You'll work through the formats that matter in practice (INT8, FP8, FP4, NF4, ternary), the calibration and fine- tuning recipes that make them work, the cross-framework export paths that get a quantized model out of research and into a runtime, and the deployment stacks where the speedups are actually realized. Along the way we'll confront the places where intuition misleads: why symmetric and asymmetric quantization behave differently for activations, why some "faster" formats are slower in practice, and why a model can pass every numerical check and still regress on downstream tasks. © Manning Publications Co. To comment go to liveBook
Page 4
The field moves quickly, and I'll update these chapters through the MEAP as new techniques stabilize and as your feedback shapes the book. Please post questions, corrections, and suggestions in the liveBook Discussion forum — a book like this gets sharper with every reader who pushes back on it. —Vivek Kalyanarangan © Manning Publications Co. To comment go to liveBook
Page 5
brief contents 1 Facing the Eciency Wall 2 Building Quantization from First Principles 3 Choosing What to Quantize and at What Granularity 4 Applying Post-Training Quantization and Calibration 5 Recovering Accuracy with Quantization-Aware Training 6 Porting Workows Across Frameworks 7 Quantizing Large Language Models in Practice 8 Pushing into Low Bits Safely 9 Shipping with the Right Toolchain 10 Running on CPUs and Using GGUF Well 11 Targeting Edge and Mobile Devices 12 Delivering Proven Results with Two Capstones © Manning Publications Co. To comment go to liveBook
Page 6
1 Facing the Efficiency Wall  This chapter covers The memory-bandwidth bottleneck  Why quantization targets the dominant cost  The floating-point to integer transition  For most of the history of machine learning, efficiency was a secondary concern. Models were small enough to fit comfortably in memory. Inference was fast enough to feel instantaneous. When performance lagged, the usual remedies—better hardware, modest architectural tweaks, or more aggressive batching—were generally sufficient. Accuracy was the main currency, and the cost of getting there was often treated as an operational detail. That era has ended. The modern generation of models, especially large transformers, has pushed inference across a qualitative threshold. Parameter counts exploded, context lengths stretched by orders of magnitude, and workloads that once behaved like ordinary applications now behave like infrastructure. Latency flattens even on powerful GPUs. Utilization looks suspiciously low. Power draw and memory bandwidth, not arithmetic throughput, become the binding constraints. Quantization is the technique of representing neural network weights and activations using fewer bits—typically moving from 16-bit or 32-bit floating point down to 8-bit or 4-bit integers. By reducing the number of bits that must be stored and moved through the memory hierarchy, quantization directly attacks the dominant cost of modern inference: data movement. 1 © Manning Publications Co. To comment go to liveBook
Page 7
Quantization occupies a unique position among efficiency techniques. Unlike architectural redesigns or task-specific compression, it directly targets bytes moved per token without changing model structure or semantics. It reduces memory footprint, bandwidth, and energy in one stroke, and it composes cleanly with almost every other optimization you might already be using. The math is straightforward; the engineering judgment required to apply it well is not. This book is for ML engineers, infrastructure teams, and applied researchers who need to deploy models in production. If you've stared at a GPU utilization dashboard wondering why your expensive hardware sits half-idle, or watched cloud bills climb faster than traffic, or tried to squeeze a model onto an edge device and failed, this book is for you. We assume familiarity with PyTorch and basic neural network concepts, but no prior background in numerical methods or hardware architecture. Before we learn how to apply quantization to your infrastructure, we need a clear picture of why efficiency has become unavoidable, why quantization is the central lever, and what kind of trade-offs you are about to navigate. That clarity will anchor every technical decision for the rest of the book. Let's start by making the cost crisis concrete. 1.1 The cost crisis in memory, latency, and power Running AI systems today involves a different set of constraints than many engineers expect. Models are larger, inference workloads are less predictable, and performance can no longer be evaluated independently of cost, latency, and operational complexity. Where earlier systems could often be trained, fine-tuned, and deployed on a single machine, modern large language models frequently push the limits of available hardware. They may require substantial memory, exhibit noticeable latency even on high-end accelerators, and incur ongoing costs with every request they serve. These constraints are not temporary growing pains, nor are they the result of poor engineering practices. They are the natural consequence of a shift toward more general- purpose models designed to handle a wide range of tasks with minimal specialization. As a result, familiar optimization strategies—such as focusing exclusively on model accuracy or architectural elegance—are no longer sufficient. Engineers must now think in terms of systems: how models are deployed, how they scale under load, how latency affects user experience, and how costs accumulate over time. Since 2022, progress in LLMs has largely followed one blunt but effective strategy: make the model bigger and train it longer. That strategy worked spectacularly for capability—but it changed the economics of inference. According to data tracked by Our World in Data, the number of parameters in frontier AI models has approximately doubled every year since 2010. Context lengths expanded by four to sixteen times. Those multipliers don't add—they multiply. Here is the punchline up front: modern LLM systems are limited by memory movement and power, not by raw compute. That single fact explains why GPUs sit underutilized, why latency plateaus even on faster hardware, and why cloud bills scale faster than traffic. 2 © Manning Publications Co. To comment go to liveBook
Page 8
At inference time, a model is not "thinking." It is doing two very mundane things over and over: reading numbers from memory and performing simple arithmetic on them. The surprise for most people is not what happens—but which of these dominates cost. COMPUTE IS CHEAP. DATA MOVEMENT IS NOT Before we go further, let's sanity-check the energy numbers you're about to see. These are order-of-magnitude, hardware-agnostic figures grounded in published computer- architecture measurements. Exact values vary by process node and vendor, but the ratios are stable across decades of silicon. The numbers below illustrate the approximate energy cost of common operations in modern hardware. Arithmetic operations such as multiply- adds require relatively little energy, while the cost of accessing memory increases dramatically as data moves farther from the processor—from on-chip caches to off-chip DRAM. Approximate energy costs per operation: an FP32 fused multiply-add consumes roughly 3–5 picojoules (pJ), an FP16 fused multiply-add roughly 1–2 pJ, and an INT8 (8-bit integer) multiply-add roughly 0.2–0.5 pJ. These numbers come from measured ALU energy in modern CMOS designs, summarized across industry and academic surveys. Approximate energy costs per memory access: an L1 cache access costs roughly 1–5 pJ, an L2 cache access roughly 10–30 pJ, and an HBM (High Bandwidth Memory—the stacked DRAM used on modern GPUs) or off-chip DRAM access roughly 300–1000 pJ. You can see the numbers for yourself in figure 1.1, where we directly plot out the energy required for every operation. Notice the clear wall between memory access and compute. 3 © Manning Publications Co. To comment go to liveBook
Page 9
Figure 1.1 Fetching data from HBM costs roughly 1,700× more energy than an INT8 multiply-add. This gap explains why inference systems are bottlenecked by memory, not compute. The exact values matter less than the ratio. Fetching data from DRAM costs roughly two to three orders of magnitude more energy than arithmetic. This gap is structural. Compute energy scales aggressively with process improvements; memory access energy does not. A RUNNING EXAMPLE: SERVING A 7B-PARAMETER MODEL Let's work with a model size that many teams actively deploy today—models like Llama 2 7B or Mistral 7B represent this class well. At 7 billion parameters, the raw model size varies by precision: at FP32, the model occupies roughly 28 GB (7 billion parameters times 4 bytes); at FP16, roughly 14 GB; at INT8, roughly 7 GB. So far, this looks like a simple storage problem. It isn't. The real cost shows up when the model runs. 4 © Manning Publications Co. To comment go to liveBook
Page 10
NOTE A few terms we'll use throughout this section. The KV cache (key-value cache) stores the attention keys and values for every token generated so far; the model re-reads it at each step to maintain context. K and V refer to these key and value matrices individually. Hidden size is the width of each layer's internal representation—4,096 for a typical 7B model. When we refer to the "FP16 baseline" or "half-precision model" throughout this book, we mean whichever 16-bit floating-point format the model was trained and distributed in. Two 16-bit formats coexist in practice: IEEE FP16 (5 exponent bits, 10 mantissa bits, max value ~65,504) and BF16 (8 exponent bits, 7 mantissa bits, max value ~3.4×10³⁸). BF16 keeps FP32's dynamic range at the cost of coarser precision and has become the default for LLM training since roughly 2020—when you load a Hugging Face checkpoint with torch_dtype=torch.bfloat16, you are getting BF16 weights. Both formats occupy 2 bytes per value, so every memory traffic and energy calculation in this chapter applies identically regardless of which 16-bit format is in use. The precision/range tradeoff between the two formats has consequences for quantization error that we examine in Chapter 8. For each generated token, a transformer moves data through memory in three places: the model's weights, the KV cache, and activations and intermediate buffers between layers. Together, these determine the energy cost of inference. For a realistic 7B-parameter transformer with a roughly 2,048-token context, the dominant contributors to memory traffic per generated token are: model weights at FP16 (roughly 14 GB), KV cache reads and writes (roughly 1.0 GB), and activations and intermediates (roughly 0.5 MB, negligible by comparison). This already includes both K and V, FP16 storage, the model's 32 layers, its hidden size, and the full context length (because attention reads the entire cache on each new token). Taken together, this results in approximately 15 GB of memory traffic per generated token. Using a conservative HBM energy estimate of 500 picojoules per 4 bytes transferred, this corresponds to roughly 1.9 joules per generated token. SCALING TO REAL WORKLOADS Now let's translate that into something operational. Assume a modest but realistic serving rate for one replica: 1,000 generated tokens per second. Using roughly 1.9 joules per token, the power draw is approximately 1.9 kW per replica. This is steady-state power, not a peak. Run continuously for a year: 1.9 kW times 24 hours times 365 days equals roughly 16,600 kWh per year. For context, a typical U.S. household consumes about 11,000 kWh per year—roughly four times what a typical Indian household consumes. In practice, no production system runs a single replica in isolation. As capacity planning scales out, the power footprint grows in ways that are easy to underestimate. A small deployment of a handful of replicas already rivals the continuous draw of a residential block. A regional deployment begins to resemble the load profile of a commercial facility. A globally replicated service accumulates energy usage comparable to a small town. 5 © Manning Publications Co. To comment go to liveBook
Page 11
This is how inference crosses an invisible line—from an application concern to an infrastructure and energy-planning problem. Three trends amplify everything we just calculated. First, models keep growing— parameters increase faster than memory bandwidth. Second, context lengths keep expanding—the KV cache adds even more memory traffic. Third, test-time compute with reasoning models requires producing more tokens for the same query. Figure 1.2 makes the second trend concrete by showing how memory traffic per token changes as context length scales from 512 tokens to 128K. The model weights remain fixed, but the KV cache grows linearly—and quickly dominates total traffic. Figure 1.2 Memory traffic per token as context length grows. Model weights (filled) stay constant, but KV cache (hatched) scales linearly with context. At 128K tokens, total memory traffic reaches 78 GB per token— 5.5× more than at 512 tokens. Hardware improves—but power budgets, memory bandwidth, and thermal limits do not scale nearly as fast. Eventually, you hit a wall where latency stops improving, GPU utilization looks suspiciously low, and cost grows faster than usage. That wall is not accidental. 6 © Manning Publications Co. To comment go to liveBook
Page 12
NOTE The first mental anchor of this book. Throughout this book, we will pause at key moments to establish mental anchors—foundational ideas worth returning to when the details get complex. Here is the first: Bit precision is not primarily an accuracy decision—it is an energy decision. Every extra bit increases memory footprint, increases memory traffic, increases energy per token, and caps throughput long before compute does. You've now seen the wall—but not yet the escape hatch. The natural next question is: if lower precision saves this much energy, why doesn't it destroy model quality? Answering that carefully is the rest of this book. 1.2 Why quantization is the practical response We have now arrived at an uncomfortable but unavoidable conclusion: modern inference systems are constrained not by arithmetic, but by memory movement and power. Each generated token drags gigabytes through the memory hierarchy, and every extra bit of precision compounds that cost. At that point, the natural question is not whether efficiency matters, but which lever actually moves the needle. When faced with rising costs, it is tempting to assume that the problem will be solved the way previous ones were: wait for the next generation of hardware. More FLOPs. Faster GPUs. Wider memory buses. The analysis from Section 1.1 already hints at why this instinct fails. Hardware progress overwhelmingly favors compute density, not energy per byte moved. Arithmetic units shrink, pipelines deepen, and peak throughput grows rapidly—but the cost of fetching data from off-chip memory (any memory physically separate from the processor die, such as the HBM stacks on a GPU or conventional DRAM) improves far more slowly. Even high-bandwidth memory remains orders of magnitude more expensive, energetically, than a multiply-accumulate (a single multiply-then-add operation—the fundamental arithmetic step in neural network inference). As a result, GPUs get faster, yet inference latency plateaus. Peak FLOPs rise, yet utilization remains low. Power draw tracks memory traffic, not compute. This is not a short-term inefficiency. It is a structural mismatch between what models demand and what hardware can cheaply provide. Any response that does not directly reduce memory traffic is, at best, a partial solution. Before committing to quantization, it is worth acknowledging the other techniques commonly proposed to make models cheaper to run. Many of them are valuable. None of them, on their own, address the dominant term. Architectural changes such as Mixture-of-Experts, conditional computation, or sparse activation patterns reduce the amount of work performed per token. They can be highly effective—but they require architectural changes, complicate training and serving, and often shift cost rather than eliminate it. Most importantly, they are not drop-in: you cannot apply them to an existing deployed model without retraining or redesign. 7 © Manning Publications Co. To comment go to liveBook
Page 13
Attention optimizations and kernel tricks—such as FlashAttention (an algorithm that restructures attention computation to minimize memory reads), fused kernels (GPU programs that combine multiple operations into a single pass to avoid extra memory round- trips), and improved scheduling—dramatically reduce overheads inside attention blocks. They are essential optimizations—and they are already widely used. However, these techniques optimize how data moves, not how much data must move. The model weights still need to be read. The KV cache still needs to be accessed. The dominant memory traffic remains. Smarter batching and caching strategies—dynamic batching, speculative decoding, and request coalescing—improve throughput and tail latency under load. They are operationally critical, but they are amortization strategies, not reductions in per-token cost. When load rises, the same underlying energy bill eventually asserts itself. Knowledge distillation trains a smaller "student" model to reproduce the outputs of a larger "teacher" model, effectively compressing the teacher's learned behavior into fewer parameters. These approaches can work extremely well, but they are explicitly out of scope for this book. The reason is not lack of importance, but lack of generality: distillation produces a different model, not a different representation of the same one. It requires retraining, data access, and task-specific evaluation. The resulting model often trades flexibility and transferability for efficiency. The common limitation across all these approaches is that none of them systematically reduce the number of bits that must be moved per token across the entire model. Quantization does. Quantization is different in one crucial way: it scales down the dominant term linearly. Reducing precision from 16 bits to 8 bits cuts model weight size in half, memory bandwidth per token in half, and energy per token in half. The same is true when going to 4 bits, or mixed-precision schemes. No architectural changes. No full-scale retraining required (bar some techniques that use tiny calibration datasets). No changes to the computation graph. The operations are the same. The bytes are fewer. This is why quantization appears everywhere in practical systems—from mobile inference to data-center LLM serving. It is not because it is theoretically elegant, but because it is brutally effective. At this point, a reasonable objection arises: if precision drops so much, why doesn't everything break? The answer lies in what neural networks actually compute. Weights in a trained network do not encode exact values in the way a physics simulation does. They encode directions in a high-dimensional space (which features matter), correlations between inputs (which features co-occur), and relative influence (how strongly each connection contributes to the output). Most parameters tolerate small perturbations without materially changing the output. Neural networks are redundant, noisy by design, and trained under stochastic approximations. Quantization does not introduce a fundamentally new kind of error. It introduces a bounded, structured approximation—one the network is often already equipped to absorb. This is why reduced precision degrades quality gradually rather than catastrophically. 8 © Manning Publications Co. To comment go to liveBook
Page 14
In practice, the degradation is remarkably gentle at moderate bit-widths. As demonstrated by Frantar et al. in "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers" (2023), weight-only INT4 quantization of Llama 2 7B using GPTQ adds only ~0.1 perplexity points on WikiText-2 (from 5.47 to 5.61), while INT8 is nearly indistinguishable from FP16. Below 4-bit, quality degrades more noticeably and method choice matters enormously — naive 3-bit quantization can cost several perplexity points, while specialized methods like AWQ (Lin et al., 2023) or OmniQuant (Shao et al., 2023) narrow that gap significantly. Another way to frame this is to recognize that inference already lives in an approximate world. Training noise, dropout, numerical non-determinism, and finite batch statistics all introduce variation. Quantization adds one more controlled source of error. When done carefully, errors distribute across layers, biases can be corrected, and outliers can be handled explicitly. The result is not "wrong answers," but slightly different answers—often well within acceptable tolerances. In practice, model quantization is a series of recurring decisions that all trade accuracy, performance, and deployment constraints against one another. You must decide what to quantize—weights, activations, or the key-value cache—and how to do it, whether per- tensor, per-channel, in groups, or in blocks. You must also decide when quantization happens: after training (post-training quantization, or PTQ), during training (quantization- aware training, or QAT), or through low-bit adaptation techniques like QLoRA. Just as importantly, you must choose how far to push precision—8-bit, 4-bit, mixed precision, or lower—and where the model will ultimately run, from GPUs and CPUs to edge devices and specialized runtimes. There is no single correct configuration. Each choice reshapes the system's behavior and exposes different constraints. The purpose of working through these decisions is not to memorize techniques, but to develop the judgment needed to reason about trade-offs and select the right approach for a given context. NOTE The second mental anchor of this book. Quantization works not because models are simple— but because they are redundant. Or said differently: the objective is not maximum precision—it is sufficient precision at minimum energy. You have now seen why quantization is the practical response. The next section shows what it actually means, numerically, to leave the floating-point world behind and enter the integer domain—where every value sits on a fixed grid. 9 © Manning Publications Co. To comment go to liveBook
Page 15
1.3 Mapping Floating Point vs Integer at a high level Before we talk about how quantization works, we need to understand what changes when we move from floating point to integer arithmetic—that is, from a number system where each value carries its own scale (like scientific notation) to one where all values share a fixed, uniform spacing. This is not merely a precision reduction—it is a shift between two fundamentally different ways of representing numbers, each with distinct trade-offs for neural network inference. Quantization is not just about using fewer bits. It is about crossing a boundary between two very different numerical and computational worlds: floating point and integer arithmetic. This section builds a high-level mental map of that boundary. There will be no formulas or code yet. No scale factors. No zero points. Just the conceptual terrain you need in order to reason clearly when the math does arrive. Floating point numbers were not designed for machine learning. They were designed for scientific computing. Their original goal was to represent numbers that vary across enormous ranges—astronomical distances, microscopic measurements, probabilities near zero—while preserving relative precision. To do this, floating point representations combine a mantissa that stores significant digits and an exponent that scales those digits dynamically. The result is a system that automatically adjusts precision based on magnitude. Large numbers get coarse resolution. Small numbers get fine resolution. The spacing between representable values stretches and shrinks as needed. This flexibility is extraordinarily powerful. It is also expensive. Floating point optimizes for numerical expressiveness—the ability to represent a wide range of magnitudes with precision that adapts to each value's scale—not efficiency. That trade-off made perfect sense when compute were scarce and memory traffic was not the dominant concern. In modern inference systems, the balance has flipped. At inference time, models are no longer solving delicate numerical problems. They are applying learned transformations repeatedly and predictably. Yet floating point brings along the full machinery it was designed for: variable precision (spacing between representable values changes depending on magnitude), exponent alignment (before two floating-point numbers can be added, the hardware must shift one to match the other's scale), normalization and rounding (after every operation, the result must be adjusted to fit back into the standard floating-point format), and complex arithmetic units built to handle all of this transparently. Each floating point operation requires multiple internal steps just to prepare the numbers for arithmetic. Each value occupies more memory. Each byte moved costs energy. As we’ve seen, inference systems are dominated not by computation, but by data movement and power. Floating point inflates both. The key insight here is subtle but important: floating point precision is often unused precision during inference. The model does not demand it. The hardware still pays for it. 10 © Manning Publications Co. To comment go to liveBook
Page 16
If floating point feels flexible and adaptive, integers are deliberately rigid. An integer representation fixes three things up front: the range of values that can be represented, the spacing between representable values, and the exact arithmetic used to combine them. Nothing rescales itself. Nothing adapts at runtime. Every number lives on a fixed grid. At first glance, this feels like a limitation. Why willingly give up dynamic range? The answer is that this rigidity is exactly what makes integers attractive for inference. An integer format defines a closed numerical world. INT8, for example, uses 8 bits, which yields exactly 256 distinct values, all evenly spaced. There is no concept of "very large" or "very small" beyond the chosen range. Every value costs the same amount of storage. Every operation behaves predictably. Integers do not try to be clever. They are explicit and bounded about what they can represent. This strictness shows up immediately in hardware. A floating point operation must align exponents, perform mantissa arithmetic, normalize results, and handle rounding and edge cases. An integer operation does none of this. It operates on fixed-width values, produces a result in a predictable number of cycles, and requires far less control logic. As a result, integer arithmetic typically consumes fewer cycles, integer units draw less power per operation, and more integer ALUs can fit on the same silicon area. This is not a micro- optimization. It is a structural advantage. On modern CPUs and GPUs, integer pipelines are shallower, denser, and often capable of issuing more operations per cycle. This leads to higher throughput and lower energy per token. A natural objection arises at this point: if integers are so restrictive, why doesn't accuracy collapse? The answer lies in how neural networks actually use numbers. Neural networks do not rely on exact numeric values in the way physical simulations do. They care about relative influence, direction, and correlation. Several properties make them inherently tolerant to approximation: parameter redundancy, distributed representations, and training under stochastic noise. During training, models already experience variation from stochastic gradient descent, dropout and regularization, and finite-precision arithmetic. Quantization does not introduce a foreign type of error. It introduces a bounded, structured approximation—one the model is often already equipped to absorb. This is why reduced precision degrades quality gradually, not catastrophically. At the heart of quantization lies a single unavoidable trade-off. You must decide how much range you need and how much resolution you can afford to lose. Floating point preserves both by paying a high system cost. Integers force you to choose. You cannot represent everything. You must decide what matters. This tension—range versus resolution —will appear in every quantization decision you make, whether you are choosing bit-widths, grouping schemes, or calibration strategies. A useful way to think about this transition is through analogy. Floating point is like a zoom lens—it constantly adjusts to keep objects in focus, no matter how far or close they are. Integers are like a fixed grid—once you place it, everything must align to it. Quantization is the act of choosing where to place that grid. Once placed, the grid does not move. The numbers must fit. 11 © Manning Publications Co. To comment go to liveBook
Page 17
Figure 1.3 makes this distinction concrete by showing how floating-point adapts its spacing to where values cluster, while integer quantization commits to a single, fixed grid over a chosen range. Figure 1.3 Floating point concentrates precision near zero (top), leaving large values sparsely represented. Integers use uniform spacing across your chosen range (bottom). Quantization is the act of deciding where to place that fixed grid. NOTE The third mental anchor of this book. Floating point maximizes expressiveness. Integers maximize efficiency. Quantization is the art of deciding how much expressiveness you actually need. We’re now ready to leave intuition behind and build quantization from first principles. The next chapter introduces the numerical machinery that makes this transition possible—and the specific kinds of error it creates. 12 © Manning Publications Co. To comment go to liveBook
Page 18
1.4 Summary Modern LLM inference is constrained by memory bandwidth and power consumption, not raw compute—a 7B parameter model moves roughly 15 GB through memory for every token generated, consuming nearly 2 joules of energy per token. Quantization reduces precision (typically from 16-bit to 8-bit or 4-bit integers), directly cutting memory footprint, bandwidth, and energy consumption in proportion to the bit reduction. Neural networks tolerate quantization because they encode directions and correlations rather than exact values—their inherent redundancy absorbs the bounded approximation error that lower precision introduces. The core trade-off in quantization is range versus resolution: integers force you to choose a fixed grid where every number must fit, unlike floating point which auto-scales at the cost of hardware complexity. Quantization decisions span what to quantize (weights, activations, KV cache), when to quantize (post-training or during training), how aggressively to quantize (8-bit, 4-bit, mixed), and where the model runs (GPU, CPU, edge). 13 © Manning Publications Co. To comment go to liveBook
Page 19
2 Building Quantization from First Principles   This chapter covers Fixed-point number representation  Affine quantization mapping  Scale and zero-point parameters  Quantization error analysis  Every number inside a neural network—every weight, every activation, every gradient— lives on the real number line. Training happens in floating-point, where that line stretches as far as the mathematics requires. Deployment is a different story. The devices that run inference at scale—mobile phones, edge accelerators, data-center INT8 cores—speak a cruder language: small integers packed into 8-bit or 4-bit containers. Quantization is the engineering discipline that translates between these two worlds, mapping continuous floating-point values onto a finite grid of discrete integer levels. The previous chapter established why this translation matters: smaller types mean less memory, faster arithmetic, and lower energy per inference. This chapter builds the translation itself. We start from the hardware up—how a fixed-point machine actually stores and manipulates numbers—and work toward the affine mapping that modern frameworks use to compress neural network tensors into integers. Along the way, we derive the scale and zero-point parameters, analyze the two kinds of error that quantization introduces, and confront the engineering trade-off between symmetric and asymmetric schemes. 14 © Manning Publications Co. To comment go to liveBook
Page 20
If you are a machine-learning engineer preparing to quantize a production model, this chapter gives you the mathematical vocabulary and geometric intuition to understand what your quantization toolkit is doing under the hood. If you are a systems engineer designing inference hardware, it explains why the software stack makes the choices it does. And if you are simply curious about how a 32-bit floating-point weight becomes an 8-bit integer without destroying the model, this is where that story begins. By the end of the chapter you will be able to derive the quantization parameters for any tensor, predict whether symmetric or asymmetric quantization is appropriate by inspecting a histogram, and reason about the error budget that every quantized layer must live within. 2.1 The fixed-point world In the world of high-level software (Python, PyTorch), numbers are abstract entities. But in the fixed-point world of embedded DSPs and AI accelerators, numbers are physical states of hardware switches. To understand how a machine "thinks" about 1.5 or −0.5, we must first understand how it counts. THE FIXED-POINT ABSTRACTION (Q-FORMAT) Fixed-point is simply an agreement between you and the hardware: "We will store integers, but we will mathematically treat them as if they are divided by a constant Scale Factor (S)." Equation 2.1 In binary hardware, S is always a power of two (2-n). We use Q-format notation to track this. For example, Q3.4 (signed) means 3 integer bits and 4 fractional bits, giving a scale factor of S = 2-4 = 1/16 = 0.0625. WALKTHROUGH 1: THE "1.5" EXAMPLE Let's encode the number 1.5. First, quantize: 1.5÷S=1.5×16=24. The hardware stores the integer 24, which in binary is 0001 1000. If we visualize the bits, we can see the "imaginary" binary point: 15 © Manning Publications Co. To comment go to liveBook
The above is a preview of the first 20 pages. Register to read the complete e-book.

Recommended for You

Loading recommended books...
Failed to load, please try again later

Tip the Site

Scan the WeChat Pay or Alipay code to tip. No login required.

WeChat Pay
Alipay
← Back to List