High Performance Spark, 2nd Edition (Holden Karau, Adi Polak, Rachel Warren)(Z-Library)
big data
AApache Spark is amazing when everything clicks. But if you haven't seen the performance improvements you expected or still don't feel confident enough to use Spark in production, this practical book is for you. Authors Holden Karau, Adi Polak, and Rachel Warren walk you through the secrets of the Spark code base and demonstrate performance optimizations that will help your data pipelines run faster, scale to larger datasets, and avoid costly antipatterns.
2
Views
0
Downloads
0.00
Total Donations
Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Page
1
(This page has no text content)
Page
2
Praise for the Second Edition of High Performance Spark The best Spark book has been updated with coverage of the cutting-edge of Apache Spark. With guidance on GenAI, Spark Connect, adaptive query execution, and almost everything added to the project since the first edition, it’s amazing it all fits in a single volume. I thought I’d never remove my first edition copy from my desk, but it looks like that time has come! —Russell Spitzer, PMC member of the Apache Iceberg and Apache Polaris Projects, principal engineer at Snowflake High Performance Spark bridges the gap between writing Spark code and understanding how it actually executes, giving engineers the intuition needed to debug and optimize real-world Spark workloads at scale. Backed by years of hands-on experience and contributions to the Spark community, the authors bring credibility to every concept in this book. —Dipankar Mazumdar, director of Developer Relations at Cloudera, author of Engineering Lakehouses Scaling Spark is easy until it isn’t. Also, Spark has changed a lot since they published the first edition of this great book. Holden, Adi, and Rachel cut through the noise to deliver a practical, no-nonsense masterclass on maximizing Spark’s potential. What I like is that they write in a personable way, drawing from deep first hand experience. If you’re serious about engineering with Spark, this book is required reading. —Joe Reis, coauthor of Fundamentals of Data Engineering
Page
3
You can’t build an AI-first business without the data layer, and Spark remains one of the most consequential data technologies. This book does a great job laying out the fundamentals of Spark and how to achieve scalable data transformations. It also covers how to make Spark work on GPUs and its integration with DeepSpeed, TensorFlow, and PyTorch, making Spark ready for the AI era. — Chip Huyen, author of Designing Machine Learning Systems and AI Engineering I make my living telling people their data isn’t as big as they think. But when it actually is, I want them holding this book. Holden, Adi, and Rachel have turned years of hard-won Spark scars into the kind of practical, internals-deep guidance that can turn a 3 AM job failure into a one-line config change. —Ryan Boyd, cofounder, MotherDuck
Page
4
High Performance Spark 2ND EDITION Best Practices for Scaling and Optimizing Apache Spark Holden Karau, Adi Polak, and Rachel Warren
Page
5
High Performance Spark by Holden Karau, Adi Polak, and Rachel Warren Copyright © 2026 Pigs Can Fly Labs LLC and Adi Polak. All rights reserved. Published by O’Reilly Media, Inc., 141 Stony Circle, Suite 195, Santa Rosa, CA 95401. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Aaron Black Development Editor: Shira Evans Production Editor: Christopher Faucher Copyeditor: nSight, Inc. Proofreader: Kim Cofer Indexer: WordCo Indexing Services, Inc. Cover Designer: Susan Thompson Cover Illustrator: Karen Montgomery Interior Designer: David Futato Interior Illustrator: Kate Dullea May 2017: First Edition
Page
6
June 2026: Second Edition Revision History for the Second Edition 2026-05-29: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781098145859 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. High Performance Spark, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights. 978-1-098-14585-9 [LSI]
Page
7
Preface We wrote this book for data engineers, data scientists, and ML practitioners who are looking to get the most out of Spark. If you’ve been working with Spark and invested in Spark but your experience so far has been mired by memory errors and mysterious, intermittent failures, this book is for you. If you have been using Spark for some exploratory work or experimenting with it on the side but have not felt confident enough to put it into production, this book may help. If you are enthusiastic about Spark but have not seen the performance improvements from it that you expected, we hope this book can help. This book is intended for those who have some working knowledge of Spark and may be difficult to understand for those with little or no experience with Spark or distributed computing. For recommendations of more introductory literature, see “Supporting Books and Materials”. We expect this text will be most useful to those who care about optimizing repeated queries in production, rather than to those who are primarily doing exploratory work. While writing highly performant queries is perhaps more important to the data engineer, writing those queries with Spark, in contrast to other frameworks, requires a good knowledge of the data, which is usually more intuitive to the data scientist. Thus, it may be more useful to a data engineer who may be less experienced with thinking critically about the statistical nature, distribution, and layout of data when considering performance. We hope that this book will help data engineers think more critically about their data as they put pipelines into production. Similarly for data scientists we hope to provide more understanding of how Spark works so they can use their knowledge of the data for high performance queries. We want to help our readers ask questions such as “How is my data distributed?” “Is it skewed?” “What is the range of values in a column?” and “How do we expect a given value to group?” and then apply the answers to those questions to the logic of their Spark queries.
Page
8
However, even for data scientists using Spark mostly for exploratory purposes, this book should cultivate some important intuition about writing performant Spark queries, so that as the scale of the exploratory analysis inevitably grows, you may have a better shot of getting something to run the first time. We hope to guide data <span class="keep- together">scientists</span>, even those who are already comfortable thinking about data in a distributed way, to think critically about how their programs are evaluated, empowering them to explore their data more fully and more quickly and to communicate effectively with anyone helping them put their algorithms into production. Regardless of your job title, it is likely that the amount of data with which you are working is growing quickly. Your original solutions may need to be scaled, and your old techniques for solving new problems may need to be updated. We hope this book will help you leverage Apache Spark to tackle new problems more easily and old problems more efficiently. Second Edition Notes You are reading the second edition of High Performance Spark, and for that, we thank you! If you find errors or mistakes or have ideas for ways to improve this book, please reach out to us at high-performance- spark@googlegroups.com. If you wish to be included in a “thanks” section in future editions of the book, please include your preferred display name. Supporting Books and Materials For data scientists and developers new to Spark, Learning Spark, 2nd edition, by Jules S. Damji, Brooke Wenig, Tathagata Das, and Denny Lee is an excellent introduction.1 If you’re interested in a deeper dive into ML pipelines, check out Scaling Machine Learning with Spark by Adi Polak and Advanced Analytics with Spark by Sandy Ryza, Uri Laserson, Sean Owen, and Josh Wills. Both are great books for interested data scientists and ML practitioners. For individuals more interested in streaming, Stream
Page
9
Processing with Apache Spark by Gerard Maas and François Garillot may also be of use. Beyond books, there is also a collection of intro-level Spark training material available. For individuals who prefer video. Commercially, Databricks as well as Cloudera and other Hadoop/Spark vendors offer Spark training. Previous recordings of Data + AI Summit, Spark Summit, and Spark camps, as well as many other great resources, have been posted on the Apache Spark documentation page. If you don’t have experience with Scala, we do our best to convince you to pick up Scala in Chapter 1, and if you are interested in learning, Programming Scala, 3rd Edition, by Dean Wampler is a good introduction.2 For people wanting to dive into the internals of Spark, Jacek Laskowski’s “The Internals of Spark Core” remains one of the best online resources. If you want to introduce your children, or executives, to the world of Apache Spark, Holden would like to recommend her own soon-to-be- published “Distributed Computing 4 Kids (and Executives)”. Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, and keywords.
Page
10
TIP This element signifies a tip or suggestion. NOTE This element signifies a general note. WARNING This element indicates a warning or caution. Using Code Examples Supplemental material (code examples, exercises, etc.) is available for download from the High Performance Spark GitHub repository and some of the testing code is available at the spark-testing-base GitHub repository and the Spark Validator repo. Structured Streaming machine learning examples are available at https://github.com/holdenk/spark-structured-streaming-ml. If you have a technical question or a problem using the code examples, please send email to support@oreilly.com. This book is here to help you get your job done. In general, if example code is offered with this book, you may use it in your programs and documentation. You do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission.
Page
11
We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “High Performance Spark, 2nd edition by Holden Karau, Adi Polack, and Rachel Warren (O’Reilly). Copyright 2026 Pigs Can Fly Labs LLC and Adi Polak, 978-1-098-14585-9.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning NOTE For more than 40 years, O’Reilly Media has provided technology and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit https://oreilly.com. How to Contact Us For feedback, email the authors at high-performance- spark@googlegroups.com. For random ramblings, occasionally about Spark, follow us on “social media”: Holden Karau: https://linkedin.com/in/holdenkarau, https://twitch.tv/holdenkarau, https://bsky.app/profile/holdenkarau.com, https://tech.lgbt/@holden Adi Polak: https://linkedin.com/in/polak-adi
Page
12
Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 141 Stony Circle, Suite 195 Santa Rosa, CA 95401 800-889-8969 (in the United States or Canada) 707-827-7019 (international or local) 707-829-0104 (fax) support@oreilly.com https://oreilly.com/about/contact.html We have a web page for this book, where we list errata and any additional information. You can access this page at https://oreil.ly/high-performance- spark-2e. For news and information about our books and courses, visit https://oreilly.com. Find us on LinkedIn: https://linkedin.com/company/oreilly. Watch us on YouTube: https://youtube.com/oreillymedia. General Acknowledgments Second Edition In the 2nd edition we’ve had some amazing technical reviewers, including Tristen Wentling, Scott Haines, Anya Bida, and Zubair Muhammad. In
Page
13
addition we also had a copy review from Kat Tory on some of our new content. Two appendices were contributed externally. Appendix E was written in a large part by Tomasz Magdanski with Scott Haines contributing an update for the new “real-time” mode. Appendix F was contributed by Daniel Aronovich. All mistakes remain our fault of course. This project would not have been possible without the amazing O’Reilly staff we’ve worked with throughout the process (we’re sorry it took so long) including Shira Evans, Sara Hunter, Christopher Faucher, Kristen Brown, and more. None of this would have been possible without the community that makes Apache Spark, from the core developers and PMC to the users; thank you. The 2nd edition would not have been possible without the 1st edition, and all who contributed to it. First Edition The authors would like to acknowledge everyone who has helped with comments and suggestions on early drafts of our work. Special thanks to Anya Bida, Jakob Odersky, and Katharine Kearnan for reviewing early drafts and diagrams. We’d like to thank Mahmoud Hanafy for reviewing and improving the sample code as well as early drafts. We’d also like to thank Michael Armbrust for reviewing and providing feedback on early drafts of the SQL chapter. Justin Pihony has been one of the most active early readers, suggesting fixes in every respect (language, formatting, etc.). Thanks to all of the readers of our O’Reilly early release who have provided feedback on various errata, including Kanak Kshetri and Rubén Berenguel. We’d also like to thank our dedicated (official) technical reviewers from the first edition Neelesh Srinivas Salian and Denny Lee, who read through every page, providing detailed feedback, and helped us decide what content belonged where. Finally, thank you to our respective employers for being understanding as we’ve worked on this book. Especially Lawrence Spracklen, who insisted
Page
14
we mention him here :p. Personal Acknowledgments Holden Karau I would also like to thank my wife and partners for putting up with my long in-the-bathtub writing sessions. A special thank you to Timbit for guarding the house and generally giving me a reason to get out of bed (albeit often a bit too early for my taste). An extra thank you to Paco Nathan and Ann Spencer whose advice have helped me navigate this book and so many more things. Adi Polak This book is dedicated to my husband and Arya, my daughter. Thank you for being my grounding force, my source of joy, and the reason behind so
Page
15
much of what I do. To my husband, thank you for enduring the long writing sessions, the endless technical tangents, and the many moments when I was physically present but mentally buried somewhere inside a Spark execution plan. Your patience, support, and ability to pull me back to earth made this book possible. To Arya, thank you for shining so brightly and for bringing light, laughter, and perspective into every day. Even during the most intense deadlines, you reminded me what truly matters. I also want to thank the people who challenged and sharpened my thinking over the years through deep technical conversations, difficult questions, and honest feedback. A special thank you to Holden for the opportunity to coauthor, and to the mentors, peers, and collaborators whose guidance shaped not only this book, but many parts of my career. And finally, to my cat, Chupa, for consistently reminding me that no matter how critical Spark performance may seem, an empty food bowl is always the highest-priority production incident.
Page
16
(This page has no text content)
Page
17
1 Though we may be biased, Holden was one of the coauthors for the first edition. 2 Although it’s important to note that some of the practices suggested in this book are not common practice in Spark code.
Page
18
Chapter 1. Introduction to High Performance Spark This chapter provides an overview of what we hope you will be able to learn from this book and does its best to convince you to learn to read some Scala. Feel free to skip ahead to Chapter 2 if you already know what you’re looking for. What Is Spark and Why Performance Matters ASF (currently) stands for Apache Software Foundation, although there are calls to rename the foundation.1 Spark is a high performance, general- purpose data-parallel distributed computing system that has become the most active ASF open source project, with more than 2,000 active contributors. Spark enables us to process large quantities of data, beyond what can fit on a single machine, with a high-level, relatively easy-to-use API. Spark’s design and interface are unique, and it is one of the fastest systems of its kind. Uniquely, Spark allows us to write the logic of data transformations and machine learning algorithms in a parallelizable way while being relatively system agnostic.2 However, despite its many advantages and the excitement around Spark, the simplest implementation of many common data science routines in Spark can be much slower and less robust than the best version. Since the computations we are concerned with may involve data at a very large scale, the time and resource gains from tuning code for performance are enormous. Performance does not just mean running faster; often, at this scale, it means getting something to run at all. It is possible to construct a Spark query that succeeds on megabytes while failing on gigabytes, but when refactored and adjusted with an eye toward the structure of the data and the requirements of the cluster, it succeeds on the same system with
Page
19
terabytes of data. In the authors’ experience writing production Spark code, we have seen the same tasks, run on the same clusters, run 100× faster using some of the optimizations discussed in this book. In terms of data processing, time is money, and we hope this book pays for itself through a reduction in data infrastructure costs and developer hours. Not all of these techniques apply to every use case. In fact, some techniques, when applied to the wrong datasets, can make your job slower. Especially because Spark is highly configurable and is exposed at a higher level than other computational frameworks of comparable power, we can reap tremendous benefits just by becoming more attuned to the shape and structure of our data. Some techniques can work well on certain data sizes or even certain key distributions, but not all. As a simple example, using groupByKey in Spark can very easily cause the dreaded out-of-memory (OOM) exceptions, but for data with few duplicates, this operation can be just as quick as the alternatives that we will present. Learning to understand your particular use case and system and how Spark will interact with it is a must to solve the most complex data science problems with Spark. What You Can Expect to Get from This Book Our hope is that this book will help you take your Spark queries and make them faster, able to handle larger data sizes, while using fewer resources. This book covers a broad range of tools and scenarios. You will likely pick up some techniques that might not apply to the problems you are working with but might apply to a problem in the future and may help shape your understanding of Spark more generally. Most of the chapters in this book are written with enough context to allow the book to be used as a reference; however, the structure of this book is intentional, and reading the sections in order should give you not only a few scattered tips but also a comprehensive understanding of Spark and how to make it sing. It’s equally important to point out what you will likely not get from this book. This book is not intended to introduce Spark, Scala, or Python; several other books and video series are available to get you started. The
Page
20
authors may be a little biased in this regard,3 but we think Learning Spark, 2nd edition, by Jules Damji, Brooke Wenig, Tathagata Das, and Denny Lee (O’Reilly, 2020), is an excellent option for Spark beginners, and there may be a third edition by the time you are reading this. While this book focuses on performance, it is not an operations book, so topics such as setting up a cluster and multitenancy are not covered. We assume you already have a way to use Spark in your system, so we won’t provide much assistance in making higher-level architecture decisions. Spark Versions Spark attempts to follow semantic versioning with the standard [MAJOR]. [MINOR].[MAINTENANCE] with API stability for public non- experimental nondeveloper APIs within minor and maintenance releases. Many of these experimental components are some of the more exciting ones from a performance standpoint, including things like custom aggregations, and expressions to accelerate Spark’s Datasets evaluation. Spark aims for binary API compatibility between releases, using MiMa;4 so if you are using the stable API theoretically, you generally should not need to recompile to run a job against a new version of Spark unless the major version has changed. In practice, we recommend recompiling all Spark jobs against the latest MINOR version, as mistakes in binary compatibility have been known to happen. TIP This book was created using the Spark 4.x APIs, but much of the code will work in earlier versions of Spark as well. In places where this is not the case, we have attempted to call that out. Why the Focus on Scala and Python?
The above is a preview of the first 20 pages. Register to read the complete e-book.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
# High Performance Spark, 2nd Edition
## 【One-Line Pitch】
A practical, code-first guide to making Apache Spark pipelines run faster, scale further, and avoid costly antipatterns—essential reading for data engineers and developers who already know Spark basics but struggle with production performance.
## 【Book Arc】
- **Opening (~0%–10%)**: Sets expectations for the book's scope—performance, not operations—and establishes the Spark execution model foundation: SparkContext, executors, partitions, tasks, and stages. This stage solves the problem of understanding *where* performance bottlenecks actually occur in Spark's distributed architecture.
- **Early (~10%–23%)**: Covers Spark 3.x and 4.x feature improvements, including predicate pushdown, column statistics, bloom filter joins, and shuffle hash joins. This stage helps readers modernize existing applications to leverage newer Spark capabilities.
- **Early (~23%–32%)**: Dives deep into DataFrames, Datasets, and Spark SQL types—complex types, schemas, mathematical expressions, and the Dataset API's functional transformations. This stage addresses how data representation choices affect performance.
- **Middle (~32%–48%)**: Explores advanced customization: custom encoders, user-defined table functions, RDD/DataFrame conversion trade-offs, and the critical topic of joins—including manual broadcast hash joins, handling skewed keys, and understanding SQL join execution operators. This stage solves real-world join performance problems.
- **Middle (~48%–end)**: Covers aggregation strategies, including memory-efficient aggregation objects (using arrays instead of case classes), wide versus narrow transformations, and the performance implications of shuffle-heavy operations. The book concludes with practical guidance on structuring code to support optimal join and aggregation techniques.
## 【Key Takeaways】
- **Understand the execution model before optimizing** (Early): A Spark job is a set of transformations computing one result; stages are segments computable without driver involvement, and tasks are per-partition work units. Knowing this hierarchy helps you identify where shuffles and driver bottlenecks actually occur.
- **Newer Spark versions bring significant join improvements** (Early): Bloom filter joins suit "medium-sized" joins where the smaller table is too large to broadcast, and shuffle hash joins reduce sorting overhead. The query optimizer applies these automatically, but join hints let you guide it when needed.
- **Column statistics enable smarter file skipping** (Early): With modern catalogs like Iceberg, metadata about nulls and value distributions lets Spark skip files without partition pruning—though partition-based pushdown remains more efficient when it applies.
- **Custom encoders are powerful but costly** (Early): For sparse data or nonstandard objects, custom encoders (a special case of UnaryExpression) can dramatically improve storage efficiency, but the API is unstable across minor versions and requires substantial implementation work.
- **Manual broadcast joins solve specific skew problems** (Middle): When a smaller RDD fits in memory, collecting it locally and broadcasting via `sc.broadcast` avoids shuffle overhead entirely. For partially fitting data, use `countByKeyApprox` to identify overrepresented keys and broadcast only those.
- **Join execution operators have distinct trade-offs** (Middle): Broadcast hash joins, shuffle hash joins, and sort-merge joins each suit different data sizes, skew patterns, and join conditions. Understanding these trade-offs lets you provide correct join hints and structure code to support the desired technique.
- **Aggregation object design impacts memory and speed** (Middle): Using arrays of primitive integers instead of case classes for aggregation state reduces object overhead and garbage collection pressure while maintaining readable code through Scala objects with well-named index constants.
- **Wide transformations require shuffles; narrow ones don't** (Middle): Sorting requires wide dependencies because order must be defined across all records, not just within partitions. Recognizing which operations force shuffles helps you minimize them in pipeline design.
## 【Reading Tips】
- **Skim the opening chapters** (~0%–10%) if you already understand Spark's execution model; they're foundational but review-level for experienced users. Focus instead on the Spark 3.x/4.x feature discussions that follow.
- **Deep-read the joins chapter** (Middle, ~39%–48%): This is where the most actionable performance wins live. The manual broadcast join examples are worth studying line-by-line, and the join execution operator comparison directly informs real optimization decisions.
- **Pay attention to code examples over prose**: The book's strength is its concrete Scala and Python examples—the aggregation object patterns and join implementations are directly reusable in production code.
- **Watch for version-specific notes**: The authors flag where APIs differ between Spark versions and where binary compatibility issues have occurred. Recompile against the latest minor version even if the API is theoretically stable.
- **Skip the operations content**: The book explicitly excludes cluster setup and multitenancy—don't expect deployment guidance here.
## 【Coverage Limits】
This guide covers the book's core performance topics—execution model, joins, aggregations, and Spark SQL internals—but the excerpts do not cover the book's later chapters on testing, debugging, or Spark ML performance in detail.
##
Passage locations
Excerpt 1
bsky.app/profile/holdenkarau.com, https://tech.lgbt/@holden Adi Polak: https://linkedin.com/in/polak-adi authors may be a little biased in this regard,3 but...
View in text
Excerpt 2
ample 3-3. Example 3-3. Adding Scalafix rules (./build.sbt) // Note: post upgrade you'll want to switch to modern rules // or comment out because of Scala ve...
View in text
Excerpt 3
lace of its internal encoders. For example, to use Holden’s compressed binary AltEncoder on arrays, you would add implicit def encode[Array[Dataset]] = AltEn...
View in text
Excerpt 4
park SQL join techniques are broadcast hash join, broadcast nested loop, shuffle-hash, shuffle-sort-merge, and shuffle and replicate nested loop. Review Tabl...
View in text
Recommended for You
{{#thumbnailUrl}}
{{/thumbnailUrl}}
{{^thumbnailUrl}}
{{/thumbnailUrl}}
Loading recommended books...
Failed to load, please try again later
Tip the Site
Scan the WeChat Pay or Alipay code to tip. No login required.
WeChat Pay
Alipay