Share E-Book

AuthorRichard Heimann

The goal of this book is to illuminate the connections between key breakthroughs, placing each into a narrative that spans the deep learning revolution of the 2010s and early 2020s. Rather than treating each paper as an isolated achievement, the chapters weave them together to tell a larger story that connects technological innovation with shifting paradigms, cultural milestones, and evolving philosophies within AI. In the book, you will discover what problems each solved, what doors they opened, and even what controversies or questions they raised. Examining the landmark research through the eyes of Ilya Sutskever, the book offers a cohesive framework for understanding how we arrived at today’s state of AI and where we might be heading. More than can be said for most books in machine learning and AI, generally to be referenced, not read, the book is written for a broad but technically curious audience. Its primary audience is software engineers, data scientists, and machine learning practitioners, enriching the understanding of why those models exist in their current form and the key engineering patterns that enabled them.

AI Reading Assistant

Summary and highlights from this book's index; jump to passages in the text

Passage locations
Tags
No tags
ISBN: 1633434796
Publish Year: 2026
Language: 英文
Pages: 313
File Format: PDF
File Size: 8.6 MB
Support Statistics
¥.00 · 0times
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

M A N N I N G Richard Heimann Foreword by Sebastian Raschka Foundational ideas of modern AI
Praise for Sutskever’s List If you want to learn where today’s methods came from, what motivated them, and how a relatively small set of ideas helped define the trajectory of the field, this is an excellent place to begin. —From the Foreword by Sebastian Raschka author of Build a Large Language Model (From Scratch) Provides a clear, concise and well-contextualized overview of the most important break- throughs in AI research since the 1990s. It’s a highly enjoyable read, with the author weaving the personalities of Sutskever and his peers into the broader narrative. —Nico Smuts, Data Science Executive The author refuses to treat Sutskever’s List as what it superficially appears to be— just a reading list, a paper after a paper that you have to treat in isolation. Instead, the author reconstructs it as a carefully argued intellectual journey. —Francisco Perez-Sorrosal, independent AI Engineer and Advisor The book is instrumental in forming a mental picture of the main ideas that shaped today’s achievements in AI and their interplay. —Stefano Lottini, Software Engineer The author has an exquisitely deep, detailed, and nuanced knowledge of the subject area. —Wendy Langer, Good Stuff! Tutoring Sutskever’s List is all you need if you want to learn the history, and breadth and depth of deep learning by understanding foundational research papers! Author’s intuitions at the end of each chapter have so much wisdom and insights—you will love reading them. —Bhavin Thaker, New Relic I like how the author made this book into a story. I enjoyed it thoroughly. —Jay Kelkar, Kelkar Systems Chapter 8 gave me a clear conceptualization of complexity and grokking research I didn’t have before. —Leo Huovinen, Tampere University
ii
Sutskever’s List FOUNDATIONAL IDEAS OF MODERN AI RICHARD HEIMANN FOREWORD BY SEBASTIAN RASCHKA M A N N I N G SHELTER ISLAND
For online information and ordering of this and other Manning books, please visit www.manning.com. The publisher offers discounts on this book when ordered in quantity. For more information, please contact Special Sales Department Manning Publications Co. 20 Baldwin Road PO Box 761 Shelter Island, NY 11964 Email: orders@manning.com ©2026 by Manning Publications Co. All rights reserved. No part of this publication may be reproduced, stored in a retrieval system, or transmitted, in any form or by means electronic, mechanical, photocopying, or otherwise, without prior written permission of the publisher. Many of the designations used by manufacturers and sellers to distinguish their products are claimed as trademarks. Where those designations appear in the book, and Manning Publications was aware of a trademark claim, the designations have been printed in initial caps or all caps. Recognizing the importance of preserving what has been written, it is Manning’s policy to have the books we publish printed on acid-free paper, and we exert our best efforts to that end. Recognizing also our responsibility to conserve the resources of our planet, Manning books are printed on paper that is at least 15 percent recycled and processed without the use of elemental chlorine. The author and publisher have made every effort to ensure that the information in this book was correct at press time. The author and publisher do not assume and hereby disclaim any liability to any party for any loss, damage, or disruption caused by errors or omissions, whether such errors or omissions result from negligence, accident, or any other cause, or from any usage of the information herein. Manning Publications Co. Development editor: Doug Rudder 20 Baldwin Road Technical editor: Tulika Bhatt PO Box 761 Review editor: Kishor Rit Shelter Island, NY 11964 Production editor: Kathy Rossland Copy editor: Kari Lucke Proofreader: Olga Milanko Typesetter and cover designer: Marija Tudor ISBN 9781633434790 Printed in the United States of America
To my wife and three boys: Thank you for feigning credulity every time I said, “just one more hour of writing.”
brief contents 1 ■ What did Ilya see? 1 2 ■ The AlexNet moment 17 3 ■ ResNet revolution 42 4 ■ Deep learning accelerates 66 5 ■ Attention is all you need 95 6 ■ The birth of hyperscale 129 7 ■ The pivot to reasoning 153 8 ■ Simplicity, hidden in complexity 187 9 ■ Safe superintelligence 216 epilogue The missing pieces 245 appendix Design patterns for engineers 261 references 267 index 307vi
contents foreword xi preface xiii acknowledgments xiv about this book xvi about the author xix about the cover illustration xx 1 What did Ilya see? 1 1.1 Ilya’s rise 3 1.2 The GPT-2 controversy 7 1.3 The list 12 2 The AlexNet moment 17 2.1 Feature engineering 20 2.2 Pre-AlexNet skepticism 22 2.3 ImageNet 25 2.4 Why training is hard 27 2.5 AlexNet 30 Network architecture 30 ■ Training innovations 33 Data augmentation 36 ■ Efficient and scalable training 37 ■ The effect 39vii
CONTENTSviii3 ResNet revolution 42 3.1 Telephone game 43 3.2 Comparison with other major architectures 49 ResNet vs. AlexNet and ZFNet 49 ■ ResNet vs. VGGNet 49 ■ ResNet vs. GoogLeNet 50 3.3 Real-world adoption 51 3.4 ResNet v2 51 Post-activation to pre-activation 52 ■ Ablation studies 52 3.5 Extending ResNet with dense prediction 55 3.6 Human measuring sticks 59 3.7 CS231n 62 3.8 Scale what? 63 4 Deep learning accelerates 66 4.1 The unreasonable effectiveness of RNNs 67 4.2 Understanding LSTM networks 74 4.3 Recurrent neural network regularization 79 4.4 Deep Speech 2 83 Core architecture 83 ■ Training techniques 85 Language models and decoding 87 ■ Significance and broader influence 88 ■ An engineering shift 90 5 Attention is all you need 95 5.1 Transformers 96 Empirical results 102 ■ Universal transformer 103 5.2 The Annotated Transformer 106 5.3 Bahdanau attention 107 5.4 Just point to it! 113 Experimental results 114 ■ Significance 117 5.5 Neural Turing Machines 117 Significance 119 5.6 Order matters 119 5.7 Reversing input sentences 122 5.8 Effects 125
CONTENTS ix6 The birth of hyperscale 129 6.1 Scaling canon 130 Kaplan scaling laws 131 ■ Does shape matter? 133 Scaling with data and model size 134 ■ Scaling compute 135 ■ Architecture comparisons 137 The scaling hypothesis 137 ■ Chinchilla and post-Chinchilla 138 6.2 Scaling detours 140 Jaggedness 141 ■ Contamination 143 6.3 Parallelism 145 GPipe 146 ■ Results 147 ■ Wise to pipeline? 149 6.4 Rounding errors 152 7 The pivot to reasoning 153 7.1 Variational lossy autoencoder 154 What Is VLAE? 156 ■ Results 157 Autoregressive prior 158 7.2 Relational reasoning 159 Benchmark evaluation 160 7.3 Neural message passing for quantum chemistry 162 Accuracy as a bottleneck 163 ■ Message passing neural networks 164 ■ Training and results 166 Towers architecture 167 ■ Influence 168 7.4 Relational recurrent neural networks 169 Experimental methodology and empirical results 172 Historical context 175 7.5 Reasoning models 176 Distributional brittleness 177 ■ Reversal curse 179 ■ Prompting as a reasoning interface 180 7.6 Paper doubts 182 8 Simplicity, hidden in complexity 187 8.1 Coffee automaton 188 Methodology 189 ■ Results 191
CONTENTSx8.2 Kolmogorov complexity and algorithmic randomness 194 Algorithmic statistics 195 ■ Significance 198 8.3 A tutorial introduction to the minimum description length principle 198 8.4 Keeping artificial neural networks simple 202 Methodology 202 ■ Results 203 ■ Limitations 204 Cultural influence 205 8.5 Grokking 205 Compression 208 ■ Theory to practice to theory 209 Double descent 212 8.6 Blurry JPEG 215 9 Safe superintelligence 216 9.1 Machine superintelligence 217 Defining intelligence 217 ■ Universal intelligence measure 219 ■ AIXI 221 ■ Environment taxonomy 223 ■ Agents and limits 224 9.2 AI safety 225 Intelligence explosion 226 ■ Technological singularity 226 ■ AGI (as a field) is born 227 Superintelligence 228 ■ AI foom 229 ■ Effects 229 9.3 Explaining too little and promising too much 231 Legg’s functionalism 232 ■ A hint of behaviorism 234 ■ Meaning in use 235 ■ Turing’s test 236 ■ Definitions masquerade as explanations 238 ■ Public tests 240 ■ Dirty hands, clean(er) claims 241 epilogue The missing pieces 245 appendix Design patterns for engineers 261 references 267 index 307
foreword I have always liked reading papers. In my view, a good paper does more than present a result. It explains what problems the researchers wanted to solve, what assumptions they made, what was not working at the time, and why a new idea seemed worth pursuing. In that sense, papers are one of the best ways to understand a field, because they not only present the methods but also the motivations that produced them. That is one of the reasons I enjoyed and recommend this book. It takes a good mix of influential papers and turns them into a guided tour through the development of modern AI. And rather than treating these works as a mere list, it shows how they connect. You can see how one line of work created the conditions for the next and how practical issues shaped research priorities that ultimately led to the methods we use today. This perspective is especially useful today, because modern AI moves quickly, and it is easy to encounter current methods only in their applied form. For example, AlexNet was a response to the limitations of hand- crafted features and the difficulty of making deep learning work for com- puter vision at scale. Residual networks addressed the optimization problems that came with the increasing depth of these deep neural nets. Sequence models, attention mechanisms, and Transformers emerged because earlier approaches ran into concrete bottlenecks in capturing and using information over long contexts. Later, scaling laws helped for- malize how performance improves in systematic ways when model size, data, and compute are increased together.xi
FOREWORDxii One of the strengths of this book is that it focuses on the motivations and not just the results. That’s because once we understand the original motivation behind a method, we are in a much better position to judge whether its core idea still matters and which lessons continue to general- ize. This is also why reading older papers remains so valuable for practi- tioners and researchers getting started in a field. Even when specific implementation details become outdated, the problem formulations and design choices often remain highly relevant, and they help us learn to for- mulate our own ideas and approaches to design the next generation of methods. For readers who like papers, this book offers a particularly rewarding kind of experience, and it helps restore context to works that are often reduced to citations, memes, or benchmark references. And for readers who are newer to this literature, it provides an accessible path into a body of work that can otherwise feel scattered across subfields and eras. In that sense, this book is both historical and practical. It looks back at the papers that shaped modern AI, but it also gives readers a better frame- work for understanding the present. If you want to learn where today’s methods came from, what motivated them, and how a relatively small set of ideas helped define the trajectory of the field, this is an excellent place to begin. —SEBASTIAN RASCHKA AUTHOR OF BUILD A LARGE LANGUAGE MODEL (FROM SCRATCH) FOUNDER AND PRINCIPAL AI/LLM RESEARCH ENGINEER, RAIR LAB
preface In late 2023, a question ricocheted online: “What did Ilya see?” The back- drop was a dramatic power struggle at OpenAI, where Ilya Sutskever had reportedly been alarmed enough to support ousting the CEO. Elon Musk’s public musings on the matter fueled speculation. Amid the com- motion, one intriguing clue to Sutskever’s mindset resurfaced: a personal reading list of research papers he had shared with John Carmack, infor- mally known as “Sutskever’s List.” Sutskever allegedly claimed it captured “90% of what matters today” in AI. The notion that a single collection of papers could hold the keys to modern AI lent the list an aura of mystery and importance. I first learned of Sutskever’s List through these whispers and online treasure hunts. Enthusiasts pieced together clues, and eventually a recon- structed version of the list appeared, quickly gaining almost a million views. It became a cultural touchstone. To ask, “Have you read Sutskever’s List?” was shorthand for claiming a handle on the fundamentals of con- temporary AI. And yet, this question was often more signaling, as many knew of the list without truly engaging with the papers on it. In a field rac- ing forward with speed, the idea of a stable canon of ideas felt refreshing and necessary. I realized that exploring these works and the context in which they arose could provide invaluable insight into how AI evolved into what we see today. This book grew out of that realization. xiii
acknowledgments My deepest thanks go to the researchers and engineers whose work ani- mates Sutskever’s List. Their hard-won victories have moved the field for- ward. I am grateful to the pioneers of deep learning who challenged conventional wisdom and proved skeptics wrong, to the problem-solvers who engineered systems that made AI work in the real world, and to the educators and communicators who clarified these complex ideas for a broader audience. I am especially grateful to those whose skepticism rose above paper doubt into a living demand for better evidence—those who did not merely declare the limits of AI but sharpened benchmarks, exposed brit- tle claims, and compelled the field to earn confidence through results. Their doubts strengthened the field more than paper hopes and doubts could ever. I am also deeply grateful to Manning Publications for giving this proj- ect a home and for guiding it with care from a rough proposal to a fin- ished book. Jonathan Gennick, my acquisitions editor, believed in the project early and helped bring it into being. Doug Rudder, my develop- mental editor, helped sharpen its structure, pacing, world-building, and voice. My thanks also go to Tulika Bhatt, whose technical review improved the accuracy of the book. Tulika is a senior software engineer at Netflix, where she works on real-time data systems and personalization infrastruc- ture at scale. Lastly, I am honored that Sebastian Raschka contributed the foreword; his work as a researcher, educator, and communicator inspires me and has helped make modern machine learning more accessible to countless readers.xiv
ACKNOWLEDGMENTS xv To all the reviewers: Adi Shavit, Anirban Majumder, Arslan Gabdulkha- kov, Arturo Geigel, Bhavin Thaker, Bin Hu, Bryan Cardillo, Casey Robin- son, Jérémie Clos, Thomas Briegel, Emin Tahirovic, Frances Buontempo, Francisco Perez-Sorrosal, Georgerobert Freeman, Giovanni Alzetta, Har- dev Ranglani, Henry Beveridge, Ian Yang, Ioannis Atsonios, Iurii Iurch- enko, Jay Kelkar, Jean-François Morin, Jerry Kuch, Joaquin Gracia, John Williams, Jonathan Thoms, Joseph Catanzarite, Julian Squires, Kostas Pas- sadis, Krishna Kandi, Krzysztof Kamyczek, Leo Huovinen, Lucas Roberts, M. Milan Loseke, Manoj Agarwal, Marielle Dado, Mengyue (J.P.) Li, Milo- rad Imbra, Natasha Chong, Nico Smuts, Nicolas Bievre, Payam Pourashraf, Praveen Nair, Prof. Hyung-Jong (John) Kim, Ravi Kiran Bam- idi, Recep Erol, Ruben Gonzalez-Rubio, Sanny Mulyono, Sidharth Maho- tra, Stefano Lottini, Stephen Wolff, Surbhi Madan, Vitosh K. Doynov, Wendy Langer, and William Springer: your suggestions helped made this a better book.
about this book The goal of Sutskever’s List is to illuminate the connections between key breakthroughs, placing each into a narrative that spans the deep learning revolution of the 2010s and early 2020s. Rather than treating each paper as an isolated achievement, the chapters weave them together to tell a larger story that connects technological innovation with shifting para- digms, cultural milestones, and evolving philosophies within AI. In these pages, you will discover what problems each solved, what doors they opened, and even what controversies or questions they raised. Examining the landmark research through the eyes of Ilya Sutskever, the book offers a cohesive framework for understanding how we arrived at today’s state of AI and where we might be heading. You will meet the researchers who carried the torch of deep learning, the engineers who scaled algorithms, and even the critics and skeptics who challenged the hype. We also highlight how ideas build on one another: how a trick in one experiment enabled a breakthrough in the next and how a theoretical insight found its way into practical systems. In doing so, Sutskever’s List serves both as a map of modern AI’s intellectual terrain and as a commentary on its journey. Whether you are here to learn the back- story of famous algorithms or to grasp the mindset of an AI pioneer, the book’s guiding purpose is the same: to shed light on how foundational AI research became the world-transforming force we now witness. Intended audience Sutskever’s List is the kind of book people might actually enjoy reading— and finish—which is more than can be said for most books in machinexvi
ABOUT THIS BOOK xviilearning and AI, which are generally written to be referenced, not read. Specifically, this book is written for a broad but technically curious audi- ence. Its primary audience is software engineers, data scientists, and machine learning practitioners. If you build or work with AI models, this book will enrich your understanding of why those models exist in their current form and the key engineering patterns that enabled them. However, Sutskever’s List is also meant to be accessible to motivated gen- eral readers. We assume only a basic familiarity with computing and math. If you have ever read popular science accounts of technology or enjoyed learning the stories behind scientific breakthroughs, you will find a simi- lar narrative approach here, albeit one that doesn’t shy away from the technical heart of each idea. In short, this book welcomes anyone eager to learn how AI evolved, offering different layers of insight for different readers. An engineer might appreciate the finer technical points or his- torical references, while a non-specialist will come away with a solid con- ceptual understanding. Code, software, and resources One of the distinguishing features of this book is that no software or cod- ing is required to engage with it. This is not a hands-on programming guide but rather a conceptual journey through ideas and history. No code downloads, libraries, or sandbox environments are required. All of the important algorithms and experiments are described and discussed within the text, so you won’t need to run anything on your own machine. The “output” is understanding, not software. Guidance for readers The chapters are organized in a logical order, and reading them sequen- tially is recommended, especially if you’re new to AI. If you start from chapter 1 and move forward, you’ll follow a storyline that gradually builds context, including concepts and terminology that later chapters will build upon or reference. For example, understanding how convolutional net- works triumphed in vision (chapter 2) will add depth to your appreciation of the challenges in training deeper networks (chapter 3); knowing what RNNs and LSTMs are (chapter 4) sets the stage for the Transformer’s innovation (chapter 5); and so on. The narrative is designed to be cumu- lative, revealing recurring themes and evolving perspectives as you prog- ress. Reading straight through will give you a cohesive understanding of Sutskever’s worldview and how the field of AI transformed over time. That said, the book is also structured to allow selective reading. Each chapter focuses on a distinct theme and a specific set of papers, so if you
ABOUT THIS BOOKxviiihave a particular interest in one topic—say, you’re most curious about how Transformers work, or you want to jump directly into the discussion of AI safety—you can head directly to the relevant chapter. We have made an effort to ensure that each chapter provides enough background to stand on its own. Key concepts are reintroduced when necessary, and whenever we refer to an idea from an earlier chapter, we provide a brief recap or a pointer to where it was first explained. So, a motivated reader can certainly dip into chapter 5 on attention or chapter 9 on superintelli- gence and still follow the argument. If you choose this approach, you might occasionally skip some cross-references, but you will still gain insight into that slice of the story. Whether you read it cover-to-cover or jump around, consider using the references as a tool for deeper exploration. If a particular idea fascinates you, the references can direct you to the primary source for more detailed treatment. In addition, the appendix includes key takeaways and engi- neering principles for each chapter, which can help reinforce what you’ve learned. Finally, the field of AI is full of debates. As you read, you’ll encounter these threads woven through the narrative. They include the balance between empiricism and theory, ethics and risks, safety and progress, and what “intelligence” even means. Engaging with them actively will enrich your experience. By the end, you will have not only a clearer picture of where AI’s core ideas came from but also a framework to think about where AI is going and how we, as readers and practitioners, might navi- gate that future. liveBook discussion forum Purchase of Sutskever’s List includes free access to liveBook, Manning’s online reading platform. Using liveBook’s exclusive discussion features, you can attach comments to the book globally or to specific sections or para- graphs. It’s a snap to make notes for yourself, ask and answer technical ques- tions, and receive help from the author and other users. To access the forum, go to https://livebook.manning.com/book/sutskevers-list/discussion. Manning’s commitment to our readers is to provide a venue where a meaningful dialogue between individual readers and between readers and the author can take place. It is not a commitment to any specific amount of participation on the part of the author, whose contribution to the forum remains voluntary (and unpaid). We suggest you try asking the author some challenging questions lest his interest stray! The forum and the archives of previous discussions will be accessible from the publisher’s website as long as the book is in print.
about the author RICH HEIMANN is a researcher, practitioner, educator, and communicator. With over a decade of experience with machine learning, he has witnessed firsthand the field’s rapid transformation. The author has been involved in projects spanning from fundamental algo- rithm development to applied AI systems. This hands- on experience is complemented by a passion for peda- gogy and public engagement in science and technology. In academia, the author has taught machine learning courses and mentored students, earning a reputation for demystifying complex concepts without dumbing them down. xix