Page
1
(This page has no text content)
Page
2
PySpark Cookbook Over 60 recipes for implementing big data processing and analytics using Apache Spark and Python Denny Lee Tomasz Drabas BIRMINGHAM - MUMBAI
Page
3
PySpark Cookbook Copyright © 2018 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the authors, nor Packt Publishing or its dealers and distributors, will be held liable for any damages caused or alleged to have been caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. Commissioning Editor: Amey Varangaonkar Acquisition Editor: Aman Singh Content Development Editor: Mayur Pawanikar Technical Editor: Dinesh Pawar Copy Editor: Safis Editing Project Coordinator: Nidhi Joshi Proofreader: Safis Editing Indexer: Mariammal Chettiyar Graphics: Tania Dutta Production Coordinator: Shantanu Zagade First published: June 2018 Production reference: 1280618 Published by Packt Publishing Ltd. Livery Place 35 Livery Street Birmingham B3 2PB, UK. ISBN 978-1-78883-536-7 www.packtpub.com
Page
4
mapt.io Mapt is an online digital library that gives you full access to over 5,000 books and videos, as well as industry leading tools to help you plan your personal development and advance your career. For more information, please visit our website. Why subscribe? Spend less time learning and more time coding with practical eBooks and Videos from over 4,000 industry professionals Improve your learning with Skill Plans built especially for you Get a free eBook or video every month Mapt is fully searchable Copy and paste, print, and bookmark content PacktPub.com Did you know that Packt offers eBook versions of every book published, with PDF and ePub files available? You can upgrade to the eBook version at www.PacktPub.com and as a print book customer, you are entitled to a discount on the eBook copy. Get in touch with us at service@packtpub.com for more details. At www.PacktPub.com, you can also read a collection of free technical articles, sign up for a range of free newsletters, and receive exclusive discounts and offers on Packt books and eBooks.
Page
5
Contributors About the authors Denny Lee is a technology evangelist at Databricks. He is a hands-on data science engineer with 15+ years of experience. His key focuses are solving complex large-scale data problems—providing not only architectural direction but hands-on implementation of such systems. He has extensive experience of building greenfield teams as well as being a turnaround/change catalyst. Prior to joining Databricks, he was a senior director of data science engineering at Concur and was part of the incubation team that built Hadoop on Windows and Azure (currently known as HDInsight). Tomasz Drabas is a data scientist specializing in data mining, deep learning, machine learning, choice modeling, natural language processing, and operations research. He is the author of Learning PySpark and Practical Data Analysis Cookbook. He has a PhD from University of New South Wales, School of Aviation. His research areas are machine learning and choice modeling for airline revenue management.
Page
6
About the reviewer Sridhar Alla is a big data practitioner helping companies solve complex problems in distributed computing and implement large-scale data science and analytics practice. He presents regularly at several prestigious conferences and provides training and consulting to companies. He loves writing code in Python, Scala, and Java. He has extensive hands-on knowledge of several Hadoop-based technologies, Spark, machine learning, deep learning and blockchain. Packt is searching for authors like you If you're interested in becoming an author for Packt, please visit authors.packtpub.com and apply today. We have worked with thousands of developers and tech professionals, just like you, to help them share their insight with the global tech community. You can make a general application, apply for a specific hot topic that we are recruiting an author for, or submit your own idea.
Page
7
Table of Contents Preface 1 Chapter 1: Installing and Configuring Spark 6 Introduction 6 Installing Spark requirements 8 Getting ready 8 How to do it... 8 How it works... 10 There's more... 13 Installing Java 13 Installing Python 14 Installing R 14 Installing Scala 14 Installing Maven 15 Updating PATH 15 Installing Spark from sources 16 Getting ready 16 How to do it... 16 How it works... 17 There's more... 22 See also 22 Installing Spark from binaries 23 Getting ready 23 How to do it... 23 How it works... 24 There's more... 25 Configuring a local instance of Spark 26 Getting ready 26 How to do it... 26 How it works... 26 See also 28 Configuring a multi-node instance of Spark 28 Getting ready 28 How to do it... 30 How it works... 31 See also 39 Installing Jupyter 40 Getting ready 40 How to do it... 40 How it works... 41
Page
8
Table of Contents [ ii ] There's more... 41 See also 42 Configuring a session in Jupyter 42 Getting ready 43 How to do it... 43 How it works... 44 There's more... 47 See also 50 Working with Cloudera Spark images 50 Getting ready 51 How to do it... 51 How it works... 55 Chapter 2: Abstracting Data with RDDs 56 Introduction 56 Creating RDDs 57 Getting ready 57 How to do it... 57 How it works... 58 Spark context parallelize method 58 .take(...) method 59 Reading data from files 59 Getting ready 60 How to do it... 60 How it works... 61 .textFile(...) method 61 .map(...) method 62 Partitions and performance 63 Overview of RDD transformations 65 Getting ready 65 How to do it... 66 .map(...) transformation 67 .filter(...) transformation 67 .flatMap(...) transformation 68 .distinct() transformation 68 .sample(...) transformation 69 .join(...) transformation 69 .repartition(...) transformation 70 .zipWithIndex() transformation 70 .reduceByKey(...) transformation 71 .sortByKey(...) transformation 72 .union(...) transformation 73 .mapPartitionsWithIndex(...) transformation 74 How it works... 75 Overview of RDD actions 79 Getting ready 79
Page
9
Table of Contents [ iii ] How to do it... 80 .take(...) action 80 .collect() action 81 .reduce(...) action 81 .count() action 82 .saveAsTextFile(...) action 82 How it works... 83 Pitfalls of using RDDs 87 Getting ready 89 How to do it... 89 How it works... 91 Chapter 3: Abstracting Data with DataFrames 95 Introduction 95 Creating DataFrames 96 Getting ready 96 How to do it... 96 How it works... 97 There's more... 99 From JSON 99 From CSV 100 See also 101 Accessing underlying RDDs 101 Getting ready 101 How to do it... 101 How it works... 102 Performance optimizations 104 Getting ready 105 How to do it... 105 How it works... 105 There's more... 107 See also 108 Inferring the schema using reflection 108 Getting ready 109 How to do it... 109 How it works... 109 See also 110 Specifying the schema programmatically 111 Getting ready 111 How to do it... 111 How it works... 112 See also 113 Creating a temporary table 113 Getting ready 113 How to do it... 113
Page
10
Table of Contents [ iv ] How it works... 114 There's more... 114 Using SQL to interact with DataFrames 115 Getting ready 115 How to do it... 115 How it works... 116 There's more... 116 Overview of DataFrame transformations 117 Getting ready 117 How to do it... 118 The .select(...) transformation 118 The .filter(...) transformation 118 The .groupBy(...) transformation 119 The .orderBy(...) transformation 120 The .withColumn(...) transformation 121 The .join(...) transformation 122 The .unionAll(...) transformation 126 The .distinct(...) transformation 126 The .repartition(...) transformation 127 The .fillna(...) transformation 128 The .dropna(...) transformation 129 The .dropDuplicates(...) transformation 130 The .summary() and .describe() transformations 131 The .freqItems(...) transformation 131 See also 132 Overview of DataFrame actions 132 Getting ready 133 How to do it... 133 The .show(...) action 133 The .collect() action 134 The .take(...) action 134 The .toPandas() action 134 See also 135 Chapter 4: Preparing Data for Modeling 136 Introduction 136 Handling duplicates 138 Getting ready 138 How to do it... 138 How it works... 139 There's more... 139 Only IDs differ 140 ID collisions 141 Handling missing observations 144 Getting ready 144 How to do it... 145 How it works... 145
Page
11
Table of Contents [ v ] Missing observations per row 146 Missing observations per column 148 There's more... 149 See also 152 Handling outliers 152 Getting ready 152 How to do it... 152 How it works... 153 See also 155 Exploring descriptive statistics 155 Getting ready 156 How to do it... 156 How it works... 156 There's more... 157 Descriptive statistics for aggregated columns 157 See also 159 Computing correlations 159 Getting ready 159 How to do it... 159 How it works... 159 There's more... 160 Drawing histograms 161 Getting ready 161 How to do it... 161 How it works... 162 There's more... 165 See also 166 Visualizing interactions between features 166 Getting ready 166 How to do it... 167 How it works... 167 There's more... 169 Chapter 5: Machine Learning with MLlib 171 Loading the data 171 Getting ready 171 How to do it... 172 How it works... 172 There's more... 173 Exploring the data 174 Getting ready 174 How to do it... 174 How it works... 175 Numerical features 175 Categorical features 177
Page
12
Table of Contents [ vi ] There's more... 178 See also 179 Testing the data 180 Getting ready 181 How to do it... 181 How it works... 182 See also... 184 Transforming the data 185 Getting ready 185 How to do it... 185 How it works... 186 There's more... 187 See also... 188 Standardizing the data 188 Getting ready 188 How to do it... 189 How it works... 189 Creating an RDD for training 190 Getting ready 191 How to do it... 191 Classification 191 Regression 191 How it works... 191 There's more... 193 See also 193 Predicting hours of work for census respondents 194 Getting ready 194 How to do it... 194 How it works... 194 Forecasting the income levels of census respondents 196 Getting ready 196 How to do it... 196 How it works... 196 There's more... 198 Building a clustering models 199 Getting ready 199 How to do it... 199 How it works... 200 There's more... 200 See also 201 Computing performance statistics 201 Getting ready 201 How to do it... 202 How it works... 202 Regression metrics 202
Page
13
Table of Contents [ vii ] Classification metrics 203 See also 205 Chapter 6: Machine Learning with the ML Module 206 Introducing Transformers 207 Getting ready 209 How to do it... 209 How it works... 210 There's more... 211 See also 212 Introducing Estimators 213 Getting ready 215 How to do it... 215 How it works... 216 There's more... 217 Introducing Pipelines 219 Getting ready 219 How to do it... 219 How it works... 220 See also 222 Selecting the most predictable features 222 Getting ready 222 How to do it... 222 How it works... 223 There's more... 224 See also 226 Predicting forest coverage types 227 Getting ready 227 How to do it... 227 How it works... 228 There's more... 229 Estimating forest elevation 231 Getting ready 231 How to do it... 231 How it works... 232 There's more... 233 Clustering forest cover types 235 Getting ready 235 How to do it... 235 How it works... 235 See also 237 Tuning hyperparameters 237 Getting ready 237 How to do it... 238 How it works... 239
Page
14
Table of Contents [ viii ] There's more... 241 Extracting features from text 242 Getting ready 242 How to do it... 242 How it works... 243 There's more... 245 See also 245 Discretizing continuous variables 245 Getting ready 246 How to do it... 246 How it works... 246 Standardizing continuous variables 247 Getting ready 248 How to do it... 248 How it works... 248 Topic mining 249 Getting ready 249 How to do it... 250 How it works... 251 Chapter 7: Structured Streaming with PySpark 253 Introduction 253 Understanding Spark Streaming 254 Understanding DStreams 256 Getting ready 256 How to do it... 257 Terminal 1 – Netcat window 257 Terminal 2 – Spark Streaming window 257 How it works... 259 There's more... 260 Understanding global aggregations 261 Getting ready 262 How to do it... 262 Terminal 1 – Netcat window 262 Terminal 2 – Spark Streaming window 262 How it works... 265 Continuous aggregation with structured streaming 266 Getting ready 266 How to do it... 266 Terminal 1 – Netcat window 267 Terminal 2 – Spark Streaming window 267 How it works... 270 Chapter 8: GraphFrames – Graph Theory with PySpark 271 Introduction 272 Installing GraphFrames 273
Page
15
Table of Contents [ ix ] Getting ready 274 How to do it... 274 How it works... 275 Preparing the data 276 Getting ready 276 How to do it... 277 How it works... 278 There's more... 278 Building the graph 280 How to do it... 280 How it works... 282 Running queries against the graph 282 Getting ready 282 How to do it... 282 How it works... 284 Understanding the graph 285 Getting ready 285 How to do it... 285 How it works... 286 Using PageRank to determine airport ranking 287 Getting ready 288 How to do it... 288 How it works... 289 Finding the fewest number of connections 289 Getting ready 290 How to do it... 290 How it works... 290 There's more... 291 See also 292 Visualizing the graph 293 Getting ready 293 How to do it... 293 How it works... 299 Index 300
Page
16
Preface Apache Spark is an open source framework for efficient cluster computing with a strong interface for data parallelism and fault tolerance. This book presents effective and time- saving recipes for leveraging the power of Python and putting it to use in the Spark ecosystem. You'll start by learning about the Apache Spark architecture and seeing how to set up a Python environment for Spark. You'll then get familiar with the modules available in PySpark and start using them effortlessly. In addition to this, you'll discover how to abstract data with RDDs and DataFrames, and understand the streaming capabilities of PySpark. You'll then move on to using ML and MLlib in order to solve any problems related to the machine learning capabilities of PySpark, and you'll use GraphFrames to solve graph-processing problems. Finally, you will explore how to deploy your applications to the cloud using the spark-submit command. By the end of this book, you will be able to use the Python API for Apache Spark to solve any problems associated with building data-intensive applications. Who this book is for This book is for you if you are a Python developer looking for hands-on recipes for using the Apache Spark 2.x ecosystem in the best possible way. A thorough understanding of Python (and some familiarity with Spark) will help you get the best out of the book. What this book covers Chapter 1, Installing and Configuring Spark, shows us how to install and configure Spark, either as a local instance, as a multi-node cluster, or in a virtual environment. Chapter 2, Abstracting Data with RDDs, covers how to work with Apache Spark Resilient Distributed Datasets (RDDs).
Page
17
Preface [ 2 ] Chapter 3, Abstracting Data with DataFrames, explores the current fundamental data structure—DataFrames. Chapter 4, Preparing Data for Modeling, covers how to clean up your data and prepare it for modeling. Chapter 5, Machine Learning with MLlib, shows how to build machine learning models with PySpark's MLlib module. Chapter 6, Machine Learning with the ML Module, moves on to the currently supported machine learning module of PySpark—the ML module. Chapter 7, Structured Streaming with PySpark, covers how to work with Apache Spark structured streaming within PySpark. Chapter 8, GraphFrames – Graph Theory with PySpark, shows how to work with GraphFrames for Apache Spark. To get the most out of this book You need the following to smoothly work through the chapters: Apache Spark (downloadable from http:/ /spark. apache. org/ downloads. html) Python Download the example code files You can download the example code files for this book from your account at www.packtpub.com. If you purchased this book elsewhere, you can visit www.packtpub.com/support and register to have the files emailed directly to you. You can download the code files by following these steps: Log in or register at www.packtpub.com.1. Select the SUPPORT tab.2. Click on Code Downloads & Errata.3. Enter the name of the book in the Search box and follow the onscreen4. instructions.
Page
18
Preface [ 3 ] Once the file is downloaded, please make sure that you unzip or extract the folder using the latest version of: WinRAR/7-Zip for Windows Zipeg/iZip/UnRarX for Mac 7-Zip/PeaZip for Linux The code bundle for the book is also hosted on GitHub at https:/ / github. com/ PacktPublishing/PySpark- Cookbook. In case there's an update to the code, it will be updated on the existing GitHub repository. We also have other code bundles from our rich catalog of books and videos available at https://github. com/ PacktPublishing/ . Check them out! Download the color images We also provide a PDF file that has color images of the screenshots/diagrams used in this book. You can download it here: https:/ /www. packtpub. com/ sites/ default/ files/ downloads/PySparkCookbook_ ColorImages. pdf. Conventions used There are a number of text conventions used throughout this book. CodeInText: Indicates code words in text, database table names, folder names, filenames, file extensions, pathnames, dummy URLs, user input, and Twitter handles. Here is an example: "Next, we call three functions: printHeader, checkJava, and checkPython." A block of code is set as follows: if [ "${_check_R_req}" = "true" ]; then checkR fi When we wish to draw your attention to a particular part of a code block, the relevant lines or items are set in bold: if [ "$_machine" = "Mac" ]; then curl -O $_spark_source elif [ "$_machine" = "Linux"]; then wget $_spark_source
Page
19
Preface [ 4 ] Any command-line input or output is written as follows: tar -xvf sbt-1.0.4.tgz sudo mv sbt-1.0.4/ /opt/scala/ Bold: Indicates a new term, an important word, or words that you see onscreen. For example, words in menus or dialog boxes appear in the text like this. Here is an example: "Go to File | Import appliance; click on the button next to the path selection." Warnings or important notes appear like this. Tips and tricks appear like this. Sections In this book, you will find several headings that appear frequently (Getting ready, How to do it..., How it works..., There's more..., and See also). To give clear instructions on how to complete a recipe, use these sections as follows: Getting ready This section tells you what to expect in the recipe and describes how to set up any software or any preliminary settings required for the recipe. How to do it... This section contains the steps required to follow the recipe. How it works... This section usually consists of a detailed explanation of what happened in the previous section.
Page
20
Preface [ 5 ] There's more... This section consists of additional information about the recipe in order to make you more knowledgeable about the recipe. See also This section provides helpful links to other useful information for the recipe. Get in touch Feedback from our readers is always welcome. General feedback: Email feedback@packtpub.com and mention the book title in the subject of your message. If you have questions about any aspect of this book, please email us at questions@packtpub.com. Errata: Although we have taken every care to ensure the accuracy of our content, mistakes do happen. If you have found a mistake in this book, we would be grateful if you would report this to us. Please visit www.packtpub.com/submit-errata, selecting your book, clicking on the Errata Submission Form link, and entering the details. Piracy: If you come across any illegal copies of our works in any form on the internet, we would be grateful if you would provide us with the location address or website name. Please contact us at copyright@packtpub.com with a link to the material. If you are interested in becoming an author: If there is a topic that you have expertise in and you are interested in either writing or contributing to a book, please visit authors.packtpub.com. Reviews Please leave a review. Once you have read and used this book, why not leave a review on the site that you purchased it from? Potential readers can then see and use your unbiased opinion to make purchase decisions, we at Packt can understand what you think about our products, and our authors can see your feedback on their book. Thank you! For more information about Packt, please visit packtpub.com.