AI guide
【One-Line Pitch】
A production-minded field guide to turning a raw pre-trained model into a reliable, well-behaved system—covering SFT, RLHF, DPO/KTO/GRPO, quantization, domain adaptation, and agentic training. Best for AI engineers and developers who already use LLMs and now need to shape their behavior, not just prompt them.
【Book Arc】
- **Opening (~0%–9%)**: Frames post-training as the discipline that converts "capable by default" into "reliable by design," and lays out the book's four-part structure (Foundation, Tools, Craft, Frontier) plus reading routes.
- **Early (~9%–23%)**: The Foundation and core Tools—when to fine-tune at all, then SFT, reinforcement learning and preference optimization (PPO, DPO, KTO, GRPO, and relatives), closing with evaluation strategy and its statistical traps.
- **Early–Middle (~23%–31%)**: The Craft begins—efficiency techniques (quantization, compression, batching) and domain adaptation, including catastrophic forgetting, data sourcing, and compliance-driven training.
- **Middle (~31%–46%)**: Agentic models and reasoning: tool calling, memory management, sandboxing and least privilege, then training for complex thought; the author's framing of post-training as the "everything technology" where capability becomes character.
- **Late (~46%–54%)**: Synthetic training (self-play, generated data) and multimodal post-training beyond text—vision, speech, video, sensor data—plus the practicalities of code, compute, and confidentiality.
- **Ending (~54%+)**: Future directions: mixture of experts, long context, continual learning, build-vs-buy, and the talent bottleneck, closing on what remains when techniques change.
【Key Takeaways】
- **Post-training is where behavior is decided, not just capability** (Opening): pre-training supplies raw ability; post-training shapes instruction-following, refusals, and domain fit—the part the book argues almost no one explains.
- **Fine-tuning is a decision, not a default** (Early): the book opens its practical section by asking whether you should fine-tune at all, treating method choice as a function of the constraint you're actually under.
- **Know the method landscape well enough to debug it** (Early): SFT, RLHF/PPO, DPO, KTO, GRPO, IPO, ORPO, and SimPO are presented with their trade-offs and a practical decision tree rather than as a menu of buzzwords.
- **Catastrophic forgetting is the central risk of domain adaptation** (Early–Middle): specializing a model can abruptly overwrite what it knew; the book pairs domain data sourcing with techniques to specialize "without amnesia."
- **Evaluation is where trust is earned or lost** (Early): Goodhart's Law, benchmark zoos, LLM-as-judge, regression and A/B testing, and evaluator disagreement are treated as first-class engineering problems, not afterthoughts.
- **Agents need safety engineering, not just tool-calling skill** (Middle): tool calls fail eventually, so sandboxing, least privilege, memory management, and training models to say no are part of the craft.
- **Memory is a design constraint, not an excuse** (Early–Middle): quantization and compression let you run larger models within the hardware you actually have.
- **The frontier keeps moving** (Ending): mixture of experts, long context, synthetic data, and continual learning are framed as directions to prepare for—while the underlying judgment about purpose stays human.
【Reading Tips】
- Use the book's own reading routes: skim Part I if you already fine-tune, and deep-read Chapters 5–6 (preference optimization and evaluation) if you need to choose methods and prove they work.
- Treat the math as debugging equipment—the author's stated reason for including it is to let you diagnose failures, so slow down where a technique's assumptions are spelled out.
- Keep the companion codebase (referenced in the excerpts) open alongside later chapters; the text is selective and the full training scripts live online.
- Expect later examples to need multi-GPU or multinode setups; plan compute access (the book points to neoclouds and academic discounts) before the efficiency and agentic chapters.
- Read the domain adaptation and agentic chapters together—forgetting, compliance, and adversarial reliability are the same production problem viewed from two angles.
【Coverage Limits】
These excerpts cover the front matter, table of contents, preface, and author's framing plus scattered chapter material; they do not include the actual technical content of most chapters, so this guide describes the book's structure and stated aims rather than its detailed methods.
Passage locations
Excerpt 1
isco The Craft of Post-Training THE CRAFT OF POST-TRAINING. Copyright © 2026 by Chris von Csefalvay. All rights reserved. No part of this work may be reprodu...
View in text
Excerpt 2
the Objective Training the RL Cascade Does It Actually Work? PPO: The Art of Cautious Improvement The Intuition Behind the Policy Gradient Clipping: How PPO...
View in text
Excerpt 3
ance of life (or “bounded autonomy,” as we would say today). This is, of course, the story of the Golem of Prague: Immensely powerful, tireless, and obedient...
View in text
Excerpt 4
ractive course, with exercises and code examples throughout. Training scripts can be extensive, and to keep the book manageable, I have been selective with t...
View in text