An in-depth history of Large Language Models—and what their ubiquity, disruption, and creativity mean from a wider sociopolitical perspective.
In November 2022, ChatGPT swept the globe with a mixed frenzy of excitement and anxiety. Was this a step closer to reaching singularity or just another marvel in machine learning? Author Stephan Raaijmakers provides a comprehensive introduction to Large Language Models (LLMs), describing what exactly they are capable of from a technical and creative standpoint. This concise volume covers everything from the architecture of LLM neural networks to the limitations of LLMs to how our governments can regulate this technology. In explaining how exactly LLMs learn from data sets, Raaijmakers defangs the more sensational arguments we may be familiar with. Instead, he offers a more grounded approach to how this groundbreaking—and increasingly ubiquitous—form of artificial intelligence will shape our society for years to come.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A grounded, accessible tour of Large Language Models—how they work, where they came from, and what their rise means for society—perfect for curious non-specialists, students, and professionals who want to understand ChatGPT and its peers beyond the hype.
【Book Arc】
- **Opening (~0%–9%)**: Sets the stage with the ChatGPT launch in November 2022 and frames the book’s mission: demystify LLMs while addressing their societal impact. It traces early chatbots from ELIZA (1960s) to Q&A (1990s) and positions LLMs as the culmination of decades of language-model research.
- **Early (~9%–25%)**: Introduces the core concept of a language model—predicting words from context, like a Cloze test—and distinguishes statistical language models from neural ones. It defines what makes a model “large” (data size, parameter count, computing power) and explains perplexity as a measure of surprise.
- **Early (~25%–34%)**: Explores the training curriculum of LLMs, from massive reading to instruction-following, and introduces key controversies: bias (selection, algorithm, training), hallucination, and the phenomenon of emergence—unexpected capabilities that appear at scale. It also raises questions about creativity and whether LLMs merely recombine existing knowledge.
- **Middle (~38%–47%)**: Shifts to governance and AI sovereignty, highlighting how only big tech has the resources to build these models, often with opaque data, design choices, and ethical rules. It then pivots to the intellectual history of computational linguistics, covering generative grammar (Chomsky), categorial grammar (Lambek), and statistical methods (Markov, Viterbi).
- **Middle (~47%–53%)**: Continues the historical journey, showing how these four pathways—generative, logical, statistical, and machine-learning—converged to make modern LLMs possible. It discusses how linguistic principles may manifest as statistical regularities in neural networks, bridging old theory and new practice.
【Key Takeaways】
- **Language models are probability machines** (Early): At their core, they predict the next word given context, like filling in a Cloze test. This simple idea underlies everything from autocomplete to ChatGPT, and understanding it defuses sensational claims about AI sentience.
- **“Large” means three things** (Early): Size is defined by data volume, parameter count, and computing power (GPU FLOPS). Bigger isn’t always better—many models are undertrained because parameter growth outpaced data, and better-balanced smaller models can rival larger ones.
- **LLMs are trained through a curriculum** (Early): They start with massive reading, then learn to follow instructions, and some become AI assistants like ChatGPT, Copilot, and Gemini. But all are still LLMs under the hood—an umbrella term covering many variants.
- **Bias is built in, not accidental** (Early): LLMs inherit bias from their data (selection bias), architecture (algorithm bias), and human trainers (training bias). These are hard to detect and prevent, so users need healthy skepticism when models are embedded in everyday software.
- **Emergence is real but mysterious** (Early): LLMs show capabilities they weren’t explicitly trained for, which can be triggered by a few examples. Whether these are truly emergent or gradual “mirages” is an open question—and a key risk for real-world deployment.
- **Hallucination and staleness are structural** (Early): Models are built from fixed data snapshots (ChatGPT only knew up to 2022 in early 2023), and forcing accurate facts into generated text is technically hard. Digital watermarks to detect synthetic text are still unproven.
- **Governance is a power problem** (Middle): Only big tech has the data, money, and compute to build LLMs, and their processes are opaque—undisclosed data, hidden design choices, and private ethical rules. This raises urgent questions about AI sovereignty and public oversight.
- **LLMs have deep intellectual roots** (Middle): Four pathways—generative linguistics (Chomsky), categorial grammar (Lambek), statistical methods (Markov, Viterbi), and machine learning—converged to make modern LLMs possible. Old linguistic theories may now live on as statistical regularities in neural networks.
【Reading Tips】
- **Skim the early history** (~0%–9%): The ELIZA and Q&A anecdotes are charming but not essential; the key takeaway is that rule-based chatbots failed because they couldn’t learn from data. Move quickly to the technical core.
- **Deep-read the “what is a language model” section** (~16%–25%): The Cloze test analogy and the three dimensions of size (data, parameters, compute) are the foundation for everything else. Master these before proceeding.
- **Pay attention to the controversy chapters** (~28%–38%): Bias, hallucination, and emergence are where the book earns its keep. These sections are dense with implications for real-world use—read slowly and take notes on the distinctions between bias types.
- **Treat the linguistics history as context, not core** (~44%–53%): The four pathways are fascinating but heavy on theory. If you’re not a linguistics buff, skim the Chomsky and Lambek details and focus on how they connect to modern LLMs.
- **Take away the governance questions** (~38%–44%): The AI sovereignty discussion is the book’s most distinctive contribution. Even if you skip technical details, read this section to understand the societal stakes.
【Coverage Limits】
This guide is based on excerpts covering roughly the first half of the book (through ~53%). Later chapters on advanced architectures, practical applications, and regulatory proposals are not covered here.
Excerpt 1
retrieval) without permission in writing from the publisher. The MIT Press would like to thank the anonymous peer reviewers who provided comments on drafts o...
painting were puffy”). Fill in: sul , “on,” and gli , “the.” When you perform this test, you mimic a language model. Based on your—still rudimentary—knowledg...
Roose had an intense question-and-answer session with Bing. After some interactions, Bing—through ChatGPT—revealed that its real name was Sydney (which in fa...
hat can be subjected to computation did not occur overnight. How exactly, then, did language meet up with mathematics, statistics, and computation? It appear...
se, words in context) to learn to generate words in context. Cognitive scientist Steven Piantadosi discusses at length the ongoing debate between generative...
evant for LLMs: probabilistic models of language (figure 6). The origins of this framework can be traced back to the eighteenth century, and we need to talk...
rty expresses: Long-term history plays no role for player 3. Table 2 Word Drawing Without Replacement (“Reinsertion”) Bag before draw Draw Bag after draw (n...
icate measured values of these objects like mass and weight. A well-known logical problem, so-called exclusive-or (XOR), defines a nonlinearly separable prob...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Large Language Models (Stephan Raaijmakers)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Large Language Models (Stephan Raaijmakers)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment