Capable by default. Reliable by design.
A pre-trained model has read most of the internet—and can be trusted with almost none of it. Post-training is the work that changes that: where you take a raw, general model and shape it into something that behaves, follows instructions, refuses what it shouldn’t do, and handles the specific job you need. It’s the human hand on the machine, and the part almost no one explains.
Chris von Csefalvay has spent his career building production ML systems in industry, from clinical language to legal text. In The Craft of Post-Training, he shows you the decisions behind every technique: when to fine-tune and when not to, why a model quietly gets worse, and which method fits the constraint you’re actually under. The math is here, because knowing why a technique works is what lets you debug it when it breaks.
You’ll know how to:
Choose among the main post-training methods, from SFT and RLHF to DPO, KTO, and GRPO, well enough to fix failures instead of guessing
Adapt a model to your domain without catastrophic forgetting—the tendency of a network to abruptly overwrite what it already knew when you train it on something new
Run larger models with the memory you have by using new quantization
Train agentic systems to act reliably under adversarial pressure
Measure what matters in your deployment, beyond standard benchmarks
When you’ve used LLMs long enough, you start to wonder what was done to make them behave. The secret is in the post-training that shaped them. The Craft of Post-Training shows you how that’s done.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A production-minded field guide to turning a raw pre-trained model into a reliable, well-behaved system—covering SFT, RLHF, DPO/KTO/GRPO, quantization, domain adaptation, and agentic training. Best for AI engineers and developers who already use LLMs and now need to shape their behavior, not just prompt them.
【Book Arc】
- **Opening (~0%–9%)**: Frames post-training as the discipline that converts "capable by default" into "reliable by design," and lays out the book's four-part structure (Foundation, Tools, Craft, Frontier) plus reading routes.
- **Early (~9%–23%)**: The Foundation and core Tools—when to fine-tune at all, then SFT, reinforcement learning and preference optimization (PPO, DPO, KTO, GRPO, and relatives), closing with evaluation strategy and its statistical traps.
- **Early–Middle (~23%–31%)**: The Craft begins—efficiency techniques (quantization, compression, batching) and domain adaptation, including catastrophic forgetting, data sourcing, and compliance-driven training.
- **Middle (~31%–46%)**: Agentic models and reasoning: tool calling, memory management, sandboxing and least privilege, then training for complex thought; the author's framing of post-training as the "everything technology" where capability becomes character.
- **Late (~46%–54%)**: Synthetic training (self-play, generated data) and multimodal post-training beyond text—vision, speech, video, sensor data—plus the practicalities of code, compute, and confidentiality.
- **Ending (~54%+)**: Future directions: mixture of experts, long context, continual learning, build-vs-buy, and the talent bottleneck, closing on what remains when techniques change.
【Key Takeaways】
- **Post-training is where behavior is decided, not just capability** (Opening): pre-training supplies raw ability; post-training shapes instruction-following, refusals, and domain fit—the part the book argues almost no one explains.
- **Fine-tuning is a decision, not a default** (Early): the book opens its practical section by asking whether you should fine-tune at all, treating method choice as a function of the constraint you're actually under.
- **Know the method landscape well enough to debug it** (Early): SFT, RLHF/PPO, DPO, KTO, GRPO, IPO, ORPO, and SimPO are presented with their trade-offs and a practical decision tree rather than as a menu of buzzwords.
- **Catastrophic forgetting is the central risk of domain adaptation** (Early–Middle): specializing a model can abruptly overwrite what it knew; the book pairs domain data sourcing with techniques to specialize "without amnesia."
- **Evaluation is where trust is earned or lost** (Early): Goodhart's Law, benchmark zoos, LLM-as-judge, regression and A/B testing, and evaluator disagreement are treated as first-class engineering problems, not afterthoughts.
- **Agents need safety engineering, not just tool-calling skill** (Middle): tool calls fail eventually, so sandboxing, least privilege, memory management, and training models to say no are part of the craft.
- **Memory is a design constraint, not an excuse** (Early–Middle): quantization and compression let you run larger models within the hardware you actually have.
- **The frontier keeps moving** (Ending): mixture of experts, long context, synthetic data, and continual learning are framed as directions to prepare for—while the underlying judgment about purpose stays human.
【Reading Tips】
- Use the book's own reading routes: skim Part I if you already fine-tune, and deep-read Chapters 5–6 (preference optimization and evaluation) if you need to choose methods and prove they work.
- Treat the math as debugging equipment—the author's stated reason for including it is to let you diagnose failures, so slow down where a technique's assumptions are spelled out.
- Keep the companion codebase (referenced in the excerpts) open alongside later chapters; the text is selective and the full training scripts live online.
- Expect later examples to need multi-GPU or multinode setups; plan compute access (the book points to neoclouds and academic discounts) before the efficiency and agentic chapters.
- Read the domain adaptation and agentic chapters together—forgetting, compliance, and adversarial reliability are the same production problem viewed from two angles.
【Coverage Limits】
These excerpts cover the front matter, table of contents, preface, and author's framing plus scattered chapter material; they do not include the actual technical content of most chapters, so this guide describes the book's structure and stated aims rather than its detailed methods.
the Objective Training the RL Cascade Does It Actually Work? PPO: The Art of Cautious Improvement The Intuition Behind the Policy Gradient Clipping: How PPO...
ance of life (or “bounded autonomy,” as we would say today). This is, of course, the story of the Golem of Prague: Immensely powerful, tireless, and obedient...
ractive course, with exercises and code examples throughout. Training scripts can be extensive, and to keep the book manageable, I have been selective with t...
tips, tricks, and tribal wisdom that pervade the profession. When a fine-tuning run produces unexpected results, the practitioner who understands the underly...
uch of this is tedious rather than intellectually demanding. In the time since I began writing this book, a new generation of code agents and “vibe coding” t...
average frontier model’s training run is estimated to cost. Despite their economic size, foundation labs differ not so much in methods but in objectives (wit...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
The Craft of Post-Training A Practical Guide for AI Engineers and Developers (for . .) (Chris von Csefalvay)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
The Craft of Post-Training A Practical Guide for AI Engineers and Developers (for . .) (Chris von Csefalvay)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment