Share E-Book

High Performance Spark, 2nd Edition (Holden Karau, Adi Polak, Rachel Warren)(Z-Library)

Author Holden Karau, Adi Polak, Rachel Warren

big data
Language English

AApache Spark is amazing when everything clicks. But if you haven't seen the performance improvements you expected or still don't feel confident enough to use Spark in production, this practical book is for you. Authors Holden Karau, Adi Polak, and Rachel Warren walk you through the secrets of the Spark code base and demonstrate performance optimizations that will help your data pipelines run faster, scale to larger datasets, and avoid costly antipatterns.

Format PDF
Size 7.1 MB
2
Views
0
Downloads
0.00
Total Donations
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
(This page has no text content)
Page 2
Praise for the Second Edition of High Performance Spark The best Spark book has been updated with coverage of the cutting-edge of Apache Spark. With guidance on GenAI, Spark Connect, adaptive query execution, and almost everything added to the project since the first edition, it’s amazing it all fits in a single volume. I thought I’d never remove my first edition copy from my desk, but it looks like that time has come! —Russell Spitzer, PMC member of the Apache Iceberg and Apache Polaris Projects, principal engineer at Snowflake High Performance Spark bridges the gap between writing Spark code and understanding how it actually executes, giving engineers the intuition needed to debug and optimize real-world Spark workloads at scale. Backed by years of hands-on experience and contributions to the Spark community, the authors bring credibility to every concept in this book. —Dipankar Mazumdar, director of Developer Relations at Cloudera, author of Engineering Lakehouses Scaling Spark is easy until it isn’t. Also, Spark has changed a lot since they published the first edition of this great book. Holden, Adi, and Rachel cut through the noise to deliver a practical, no-nonsense masterclass on maximizing Spark’s potential. What I like is that they write in a personable way, drawing from deep first hand experience. If you’re serious about engineering with Spark, this book is required reading. —Joe Reis, coauthor of Fundamentals of Data Engineering
Page 3
You can’t build an AI-first business without the data layer, and Spark remains one of the most consequential data technologies. This book does a great job laying out the fundamentals of Spark and how to achieve scalable data transformations. It also covers how to make Spark work on GPUs and its integration with DeepSpeed, TensorFlow, and PyTorch, making Spark ready for the AI era. — Chip Huyen, author of Designing Machine Learning Systems and AI Engineering I make my living telling people their data isn’t as big as they think. But when it actually is, I want them holding this book. Holden, Adi, and Rachel have turned years of hard-won Spark scars into the kind of practical, internals-deep guidance that can turn a 3 AM job failure into a one-line config change. —Ryan Boyd, cofounder, MotherDuck
Page 4
High Performance Spark 2ND EDITION Best Practices for Scaling and Optimizing Apache Spark Holden Karau, Adi Polak, and Rachel Warren
Page 5
High Performance Spark by Holden Karau, Adi Polak, and Rachel Warren Copyright © 2026 Pigs Can Fly Labs LLC and Adi Polak. All rights reserved. Published by O’Reilly Media, Inc., 141 Stony Circle, Suite 195, Santa Rosa, CA 95401. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Aaron Black Development Editor: Shira Evans Production Editor: Christopher Faucher Copyeditor: nSight, Inc. Proofreader: Kim Cofer Indexer: WordCo Indexing Services, Inc. Cover Designer: Susan Thompson Cover Illustrator: Karen Montgomery Interior Designer: David Futato Interior Illustrator: Kate Dullea May 2017: First Edition
Page 6
June 2026: Second Edition Revision History for the Second Edition 2026-05-29: First Release See http://oreilly.com/catalog/errata.csp?isbn=9781098145859 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. High Performance Spark, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights. 978-1-098-14585-9 [LSI]
Page 7
Preface We wrote this book for data engineers, data scientists, and ML practitioners who are looking to get the most out of Spark. If you’ve been working with Spark and invested in Spark but your experience so far has been mired by memory errors and mysterious, intermittent failures, this book is for you. If you have been using Spark for some exploratory work or experimenting with it on the side but have not felt confident enough to put it into production, this book may help. If you are enthusiastic about Spark but have not seen the performance improvements from it that you expected, we hope this book can help. This book is intended for those who have some working knowledge of Spark and may be difficult to understand for those with little or no experience with Spark or distributed computing. For recommendations of more introductory literature, see “Supporting Books and Materials”. We expect this text will be most useful to those who care about optimizing repeated queries in production, rather than to those who are primarily doing exploratory work. While writing highly performant queries is perhaps more important to the data engineer, writing those queries with Spark, in contrast to other frameworks, requires a good knowledge of the data, which is usually more intuitive to the data scientist. Thus, it may be more useful to a data engineer who may be less experienced with thinking critically about the statistical nature, distribution, and layout of data when considering performance. We hope that this book will help data engineers think more critically about their data as they put pipelines into production. Similarly for data scientists we hope to provide more understanding of how Spark works so they can use their knowledge of the data for high performance queries. We want to help our readers ask questions such as “How is my data distributed?” “Is it skewed?” “What is the range of values in a column?” and “How do we expect a given value to group?” and then apply the answers to those questions to the logic of their Spark queries.
Page 8
However, even for data scientists using Spark mostly for exploratory purposes, this book should cultivate some important intuition about writing performant Spark queries, so that as the scale of the exploratory analysis inevitably grows, you may have a better shot of getting something to run the first time. We hope to guide data <span class="keep- together">scientists</span>, even those who are already comfortable thinking about data in a distributed way, to think critically about how their programs are evaluated, empowering them to explore their data more fully and more quickly and to communicate effectively with anyone helping them put their algorithms into production. Regardless of your job title, it is likely that the amount of data with which you are working is growing quickly. Your original solutions may need to be scaled, and your old techniques for solving new problems may need to be updated. We hope this book will help you leverage Apache Spark to tackle new problems more easily and old problems more efficiently. Second Edition Notes You are reading the second edition of High Performance Spark, and for that, we thank you! If you find errors or mistakes or have ideas for ways to improve this book, please reach out to us at high-performance- spark@googlegroups.com. If you wish to be included in a “thanks” section in future editions of the book, please include your preferred display name. Supporting Books and Materials For data scientists and developers new to Spark, Learning Spark, 2nd edition, by Jules S. Damji, Brooke Wenig, Tathagata Das, and Denny Lee is an excellent introduction.1 If you’re interested in a deeper dive into ML pipelines, check out Scaling Machine Learning with Spark by Adi Polak and Advanced Analytics with Spark by Sandy Ryza, Uri Laserson, Sean Owen, and Josh Wills. Both are great books for interested data scientists and ML practitioners. For individuals more interested in streaming, Stream
Page 9
Processing with Apache Spark by Gerard Maas and François Garillot may also be of use. Beyond books, there is also a collection of intro-level Spark training material available. For individuals who prefer video. Commercially, Databricks as well as Cloudera and other Hadoop/Spark vendors offer Spark training. Previous recordings of Data + AI Summit, Spark Summit, and Spark camps, as well as many other great resources, have been posted on the Apache Spark documentation page. If you don’t have experience with Scala, we do our best to convince you to pick up Scala in Chapter 1, and if you are interested in learning, Programming Scala, 3rd Edition, by Dean Wampler is a good introduction.2 For people wanting to dive into the internals of Spark, Jacek Laskowski’s “The Internals of Spark Core” remains one of the best online resources. If you want to introduce your children, or executives, to the world of Apache Spark, Holden would like to recommend her own soon-to-be- published “Distributed Computing 4 Kids (and Executives)”. Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, and keywords.
Page 10
TIP This element signifies a tip or suggestion. NOTE This element signifies a general note. WARNING This element indicates a warning or caution. Using Code Examples Supplemental material (code examples, exercises, etc.) is available for download from the High Performance Spark GitHub repository and some of the testing code is available at the spark-testing-base GitHub repository and the Spark Validator repo. Structured Streaming machine learning examples are available at https://github.com/holdenk/spark-structured-streaming-ml. If you have a technical question or a problem using the code examples, please send email to support@oreilly.com. This book is here to help you get your job done. In general, if example code is offered with this book, you may use it in your programs and documentation. You do not need to contact us for permission unless you’re reproducing a significant portion of the code. For example, writing a program that uses several chunks of code from this book does not require permission. Selling or distributing examples from O’Reilly books does require permission. Answering a question by citing this book and quoting example code does not require permission. Incorporating a significant amount of example code from this book into your product’s documentation does require permission.
Page 11
We appreciate, but generally do not require, attribution. An attribution usually includes the title, author, publisher, and ISBN. For example: “High Performance Spark, 2nd edition by Holden Karau, Adi Polack, and Rachel Warren (O’Reilly). Copyright 2026 Pigs Can Fly Labs LLC and Adi Polak, 978-1-098-14585-9.” If you feel your use of code examples falls outside fair use or the permission given above, feel free to contact us at permissions@oreilly.com. O’Reilly Online Learning NOTE For more than 40 years, O’Reilly Media has provided technology and business training, knowledge, and insight to help companies succeed. Our unique network of experts and innovators share their knowledge and expertise through books, articles, and our online learning platform. O’Reilly’s online learning platform gives you on-demand access to live training courses, in-depth learning paths, interactive coding environments, and a vast collection of text and video from O’Reilly and 200+ other publishers. For more information, visit https://oreilly.com. How to Contact Us For feedback, email the authors at high-performance- spark@googlegroups.com. For random ramblings, occasionally about Spark, follow us on “social media”: Holden Karau: https://linkedin.com/in/holdenkarau, https://twitch.tv/holdenkarau, https://bsky.app/profile/holdenkarau.com, https://tech.lgbt/@holden Adi Polak: https://linkedin.com/in/polak-adi
Page 12
Please address comments and questions concerning this book to the publisher: O’Reilly Media, Inc. 141 Stony Circle, Suite 195 Santa Rosa, CA 95401 800-889-8969 (in the United States or Canada) 707-827-7019 (international or local) 707-829-0104 (fax) support@oreilly.com https://oreilly.com/about/contact.html We have a web page for this book, where we list errata and any additional information. You can access this page at https://oreil.ly/high-performance- spark-2e. For news and information about our books and courses, visit https://oreilly.com. Find us on LinkedIn: https://linkedin.com/company/oreilly. Watch us on YouTube: https://youtube.com/oreillymedia. General Acknowledgments Second Edition In the 2nd edition we’ve had some amazing technical reviewers, including Tristen Wentling, Scott Haines, Anya Bida, and Zubair Muhammad. In
Page 13
addition we also had a copy review from Kat Tory on some of our new content. Two appendices were contributed externally. Appendix E was written in a large part by Tomasz Magdanski with Scott Haines contributing an update for the new “real-time” mode. Appendix F was contributed by Daniel Aronovich. All mistakes remain our fault of course. This project would not have been possible without the amazing O’Reilly staff we’ve worked with throughout the process (we’re sorry it took so long) including Shira Evans, Sara Hunter, Christopher Faucher, Kristen Brown, and more. None of this would have been possible without the community that makes Apache Spark, from the core developers and PMC to the users; thank you. The 2nd edition would not have been possible without the 1st edition, and all who contributed to it. First Edition The authors would like to acknowledge everyone who has helped with comments and suggestions on early drafts of our work. Special thanks to Anya Bida, Jakob Odersky, and Katharine Kearnan for reviewing early drafts and diagrams. We’d like to thank Mahmoud Hanafy for reviewing and improving the sample code as well as early drafts. We’d also like to thank Michael Armbrust for reviewing and providing feedback on early drafts of the SQL chapter. Justin Pihony has been one of the most active early readers, suggesting fixes in every respect (language, formatting, etc.). Thanks to all of the readers of our O’Reilly early release who have provided feedback on various errata, including Kanak Kshetri and Rubén Berenguel. We’d also like to thank our dedicated (official) technical reviewers from the first edition Neelesh Srinivas Salian and Denny Lee, who read through every page, providing detailed feedback, and helped us decide what content belonged where. Finally, thank you to our respective employers for being understanding as we’ve worked on this book. Especially Lawrence Spracklen, who insisted
Page 14
we mention him here :p. Personal Acknowledgments Holden Karau I would also like to thank my wife and partners for putting up with my long in-the-bathtub writing sessions. A special thank you to Timbit for guarding the house and generally giving me a reason to get out of bed (albeit often a bit too early for my taste). An extra thank you to Paco Nathan and Ann Spencer whose advice have helped me navigate this book and so many more things. Adi Polak This book is dedicated to my husband and Arya, my daughter. Thank you for being my grounding force, my source of joy, and the reason behind so
Page 15
much of what I do. To my husband, thank you for enduring the long writing sessions, the endless technical tangents, and the many moments when I was physically present but mentally buried somewhere inside a Spark execution plan. Your patience, support, and ability to pull me back to earth made this book possible. To Arya, thank you for shining so brightly and for bringing light, laughter, and perspective into every day. Even during the most intense deadlines, you reminded me what truly matters. I also want to thank the people who challenged and sharpened my thinking over the years through deep technical conversations, difficult questions, and honest feedback. A special thank you to Holden for the opportunity to coauthor, and to the mentors, peers, and collaborators whose guidance shaped not only this book, but many parts of my career. And finally, to my cat, Chupa, for consistently reminding me that no matter how critical Spark performance may seem, an empty food bowl is always the highest-priority production incident.
Page 16
(This page has no text content)
Page 17
1 Though we may be biased, Holden was one of the coauthors for the first edition. 2 Although it’s important to note that some of the practices suggested in this book are not common practice in Spark code.
Page 18
Chapter 1. Introduction to High Performance Spark This chapter provides an overview of what we hope you will be able to learn from this book and does its best to convince you to learn to read some Scala. Feel free to skip ahead to Chapter 2 if you already know what you’re looking for. What Is Spark and Why Performance Matters ASF (currently) stands for Apache Software Foundation, although there are calls to rename the foundation.1 Spark is a high performance, general- purpose data-parallel distributed computing system that has become the most active ASF open source project, with more than 2,000 active contributors. Spark enables us to process large quantities of data, beyond what can fit on a single machine, with a high-level, relatively easy-to-use API. Spark’s design and interface are unique, and it is one of the fastest systems of its kind. Uniquely, Spark allows us to write the logic of data transformations and machine learning algorithms in a parallelizable way while being relatively system agnostic.2 However, despite its many advantages and the excitement around Spark, the simplest implementation of many common data science routines in Spark can be much slower and less robust than the best version. Since the computations we are concerned with may involve data at a very large scale, the time and resource gains from tuning code for performance are enormous. Performance does not just mean running faster; often, at this scale, it means getting something to run at all. It is possible to construct a Spark query that succeeds on megabytes while failing on gigabytes, but when refactored and adjusted with an eye toward the structure of the data and the requirements of the cluster, it succeeds on the same system with
Page 19
terabytes of data. In the authors’ experience writing production Spark code, we have seen the same tasks, run on the same clusters, run 100× faster using some of the optimizations discussed in this book. In terms of data processing, time is money, and we hope this book pays for itself through a reduction in data infrastructure costs and developer hours. Not all of these techniques apply to every use case. In fact, some techniques, when applied to the wrong datasets, can make your job slower. Especially because Spark is highly configurable and is exposed at a higher level than other computational frameworks of comparable power, we can reap tremendous benefits just by becoming more attuned to the shape and structure of our data. Some techniques can work well on certain data sizes or even certain key distributions, but not all. As a simple example, using groupByKey in Spark can very easily cause the dreaded out-of-memory (OOM) exceptions, but for data with few duplicates, this operation can be just as quick as the alternatives that we will present. Learning to understand your particular use case and system and how Spark will interact with it is a must to solve the most complex data science problems with Spark. What You Can Expect to Get from This Book Our hope is that this book will help you take your Spark queries and make them faster, able to handle larger data sizes, while using fewer resources. This book covers a broad range of tools and scenarios. You will likely pick up some techniques that might not apply to the problems you are working with but might apply to a problem in the future and may help shape your understanding of Spark more generally. Most of the chapters in this book are written with enough context to allow the book to be used as a reference; however, the structure of this book is intentional, and reading the sections in order should give you not only a few scattered tips but also a comprehensive understanding of Spark and how to make it sing. It’s equally important to point out what you will likely not get from this book. This book is not intended to introduce Spark, Scala, or Python; several other books and video series are available to get you started. The
Page 20
authors may be a little biased in this regard,3 but we think Learning Spark, 2nd edition, by Jules Damji, Brooke Wenig, Tathagata Das, and Denny Lee (O’Reilly, 2020), is an excellent option for Spark beginners, and there may be a third edition by the time you are reading this. While this book focuses on performance, it is not an operations book, so topics such as setting up a cluster and multitenancy are not covered. We assume you already have a way to use Spark in your system, so we won’t provide much assistance in making higher-level architecture decisions. Spark Versions Spark attempts to follow semantic versioning with the standard [MAJOR]. [MINOR].[MAINTENANCE] with API stability for public non- experimental nondeveloper APIs within minor and maintenance releases. Many of these experimental components are some of the more exciting ones from a performance standpoint, including things like custom aggregations, and expressions to accelerate Spark’s Datasets evaluation. Spark aims for binary API compatibility between releases, using MiMa;4 so if you are using the stable API theoretically, you generally should not need to recompile to run a job against a new version of Spark unless the major version has changed. In practice, we recommend recompiling all Spark jobs against the latest MINOR version, as mistakes in binary compatibility have been known to happen. TIP This book was created using the Spark 4.x APIs, but much of the code will work in earlier versions of Spark as well. In places where this is not the case, we have attempted to call that out. Why the Focus on Scala and Python?
The above is a preview of the first 20 pages. Register to read the complete e-book.

Recommended for You

Loading recommended books...
Failed to load, please try again later

Tip the Site

Scan the WeChat Pay or Alipay code to tip. No login required.

WeChat Pay
Alipay
Back to List