Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Sebastian Raschka

Rating No ratings yet

LLM reasoning models have the power to tackle truly challenging problems that require finding the right path through multiple steps. In this book you’ll learn how to build a working reasoning model from the ground up. You will start with an existing pre-trained LLM and then implement reasoning-focused improvements from scratch. Sebastian Raschka, the bestselling author of Build a Large Language Model (From Scratch), is your guide on this exciting journey. Sebastian mentors you every step of the way with clear explanations, practical code, and a keen focus on what really matters. Understand LLM reasoning by creating your own reasoning model–from scratch! In Build A Reasoning Model (From Scratch) you’ll learn how • Implement core reasoning improvements for LLMs • Evaluate models using judgment-based and benchmark-based methods • Improve reasoning without updating model weights • Use reinforcement learning to integrate external tools like calculators • Apply distillation techniques to learn from larger reasoning models • Understand the full reasoning model development pipeline Reasoning models break problems into steps, producing more reliable answers in math, logic, and code. These improvements aren’t just a curiosity–they’re already integrated into top models like Grok 4 and GPT-5. Build A Reasoning Model (From Scratch) demystifies these complex models with a simple the best way to learn how something works is to build it yourself! You’ll begin with a pre-trained LLM, adding and improving its reasoning capabilities in ways you can see, test, and understand.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide to turning a pretrained LLM into a reasoning model, teaching you the full pipeline—evaluation, inference-time scaling, reinforcement learning, and distillation—by building each piece in code. Best for developers and ML practitioners who already know basic LLM mechanics and want to understand reasoning models by constructing one. 【Book Arc】 - **Opening (~0%–10%)**: Frames what reasoning models are, how they differ from standard instruction-tuned LLMs, and why building one from scratch is the clearest path to understanding. Introduces the token/pretraining/post-training vocabulary the rest of the book relies on. - **Early (~10%–35%)**: Sets up the coding environment and walks through generating text with a pretrained LLM—tokenization, input preparation, streaming generation, and performance touches like `torch.compile`. - **Middle (~35%–50%)**: Shifts to evaluation, building a math verifier that extracts, normalizes, and grades answers against reference solutions, and introducing verifiable rewards as groundwork for later RL. - **Late (~50%–85%)**: Covers reasoning improvements without weight updates (inference-time scaling, voting, self-refinement) and then reasoning training, including reinforcement learning with external tools like calculators. - **Ending (~85%–100%)**: Applies distillation to learn from larger reasoning models and ties the stages together into the full reasoning-model development pipeline. 【Key Takeaways】 - **Reasoning models are built on top of a pretrained LLM, not from zero** (Opening): the book starts from existing weights and layers reasoning-focused improvements on top, which keeps the scope practical and testable. - **Pretraining, instruction tuning, and preference tuning are distinct stages** (Opening): understanding this separation clarifies where reasoning improvements fit and why an instruction-tuned model is still not a chatbot. - **Evaluation is a first-class engineering problem** (Middle): the book implements a math verifier that extracts, normalizes, and grades answers, handling fractions, LaTeX, decimals, and percentages—robust grading is what makes later training signals trustworthy. - **Verifiable rewards are the bridge from evaluation to reinforcement learning** (Middle): the same verifier logic used to score math answers becomes the reward signal for RL in later chapters. - **Reasoning can improve without updating model weights** (Late): inference-time scaling techniques such as advanced text generation, voting, and self-refinement offer gains before any training is involved. - **Reinforcement learning can integrate external tools** (Late): the book shows how RL can teach a model to use tools like calculators, extending reasoning beyond what the model alone can do. - **Distillation transfers reasoning ability from larger models** (Ending): smaller models can learn reasoning behavior by learning from larger reasoning models rather than training from scratch. - **The full pipeline is the point** (Ending): evaluation, inference-time techniques, RL, and distillation are presented as connected stages, not isolated tricks. 【Reading Tips】 - Deep-read the evaluation chapter even if you care mainly about training—the verifier and grading logic underpin the RL rewards later, so skimming it will make chapter 6 feel arbitrary. - Treat the environment setup and text-generation chapters as a working baseline: get the code running end-to-end before moving on, since later chapters reuse the wrapper and generation utilities. - Skim the tokenizer internals (BPE, token ID printing) if you already know them; the conceptual payoff is small compared to the verifier and RL sections. - Pay attention to the stage diagram (evaluation → inference-time → training) as a mental map; it is the book's organizing spine and helps you place each technique. - When reading the RL and distillation chapters, keep asking "what signal is the model learning from?"—that question connects verifiable rewards, tool use, and distillation into one story. 【Coverage Limits】 The excerpts cover the book's framing, environment setup, tokenization, generation, and the math verifier in detail, but the later chapters on inference-time scaling, reinforcement learning, and distillation are represented mainly by chapter listings and brief mentions. Specific implementation details, code, and results for those later stages are not covered here.
Page 8
57 4 ■ Improving reasoning with inference-time scaling 94 5 ■ Inference-time scaling via self-refinement 136 6 ■ Training reasoning models with reinforcement...
View in text
Excerpt 2
esolution more reliably than pip. It also creates isolated environments automatically and comes with its own Python executable (but will use the system Pytho...
View in text
Excerpt 3
nizer.eos_token_id Pass end-of-sequence (eos) token ID. token_id = token.squeeze(0).tolist() print( Licensed to THIAGO BANDEIRA <thiago@lar.ifce.edu.br> 50 c...
View in text
Excerpt 4
False | False | PASS check_10 | True | True | PASS check_11 | True | True | PASS check_12 | False | False | PASS check_13 | True | True | PASS check_14 | Tru...
View in text
Excerpt 5
If a token such as "Berlin" stands out more strongly than the alternatives, it is more likely to be selected later. If the differences shrink, lower- ranked...
View in text
Excerpt 6
is loaded correctly, let’s use it with the temperature and top-p sampling code from the previous chapter on a MATH-500 prompt. Listing 5.2 Generating text wi...
View in text
Excerpt 7
int log-probability: tensor(-29.8750, dtype=torch.bfloat16) As you can see, the difference between the "Berlin" and "Bridge" is now much more pronounced (-16...
View in text
Excerpt 8
ceived quality. The idea is that the reward model can auto- matically score new model outputs, eliminating the need for human annotation at every training st...
View in text
Tags
AI categories
Artificial IntelligenceProgrammingPython
ISBN: 1633434672
Publish Year: 2026
Language: English
Pages: 440
File Format: PDF
File Size: 10.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…