Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Tomasz Drabas, Denny Lee

Rating No ratings yet

Build data-intensive applications locally and deploy at scale using the combined powers of Python and Spark 2.0 About This Book Learn why and how you can efficiently use Python to process data and build machine learning models in Apache Spark 2.0 Develop and deploy efficient, scalable real-time Spark solutions Take your understanding of using Spark with Python to the next level with this jump start guide Who This Book Is For If you are a Python developer who wants to learn about the Apache Spark 2.0 ecosystem, this book is for you. A firm understanding of Python is expected to get the best out of the book. Familiarity with Spark would be useful, but is not mandatory. What You Will Learn Learn about Apache Spark and the Spark 2.0 architecture Build and interact with Spark DataFrames using Spark SQL Learn how to solve graph and deep learning problems using GraphFrames and TensorFrames respectively Read, transform, and understand data and use it to train machine learning models Build machine learning models with MLlib and ML Learn how to submit your applications programmatically using spark-submit Deploy locally built applications to a cluster In Detail Apache Spark is an open source framework for efficient cluster computing with a strong interface for data parallelism and fault tolerance. This book will show you how to leverage the power of Python and put it to use in the Spark ecosystem. You will start by getting a firm understanding of the Spark 2.0 architecture and how to set up a Python environment for Spark. You will get familiar with the modules available in PySpark. You will learn how to abstract data with RDDs and DataFrames and understand the streaming capabilities of PySpark. Also, you will get a thorough overview of machine learning capabilities of PySpark using ML and MLlib, graph processing using GraphFrames, and polyglot persistence using B

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide for Python developers to master Apache Spark 2.0, covering everything from core architecture and data abstractions to machine learning, graph processing, and streaming—ideal for anyone wanting to build scalable data applications without leaving Python. 【Book Arc】 - **Opening (~0%–9%)**: Introduces Spark 2.0's architecture, the unified SparkSession, and the core abstractions (RDDs, DataFrames, Datasets), explaining how Catalyst Optimizer and Project Tungsten deliver performance parity across languages. - **Early (~9%–28%)**: Dives into Resilient Distributed Datasets (RDDs)—the schema-less foundation—covering creation from various sources, transformations like `.map()`, and actions like `.take()` and `.collect()`, with practical warnings about partitioning and executor behavior. - **Early (~28%–38%)**: Shifts to DataFrames, showing how to create them from JSON, infer schemas via reflection, or specify schemas programmatically, plus SQL-like querying and RDD interoperation. - **Middle (~38%–47%)**: Focuses on data preparation for modeling—handling duplicates, missing values, and outliers with aggregation functions, and emphasizes the importance of understanding your data before building models. - **Late (~47%–100%)**: Covers the MLlib and ML packages for machine learning (classification, clustering, regression), GraphFrames for graph analytics, TensorFrames for deep learning, Blaze for polyglot persistence, and Structured Streaming for real-time processing. 【Key Takeaways】 - **SparkSession unifies all contexts** (Early): In Spark 2.0, SparkConf, SparkContext, SQLContext, and HiveContext merge into a single entry point, simplifying configuration and resource management—use it for reading data, metadata, and cluster control. - **DataFrames bring performance parity** (Early): Unlike RDDs, which can be slow in Python, DataFrames leverage Catalyst Optimizer and Project Tungsten for fast in-memory encoding, making Python competitive with Scala/Java—critical for production workloads. - **RDDs are schema-less but flexible** (Early): Transformations like `.map()` and `.filter()` shape data lazily, while actions like `.take(n)` execute jobs—prefer `.take()` over `.collect()` for large datasets to avoid moving everything to the driver. - **Schema inference vs. programmatic schemas** (Middle): Use reflection for concise code when schema is known; specify schemas programmatically when columns and types are only known at runtime—essential for dynamic data sources. - **Data cleaning is a multi-step process** (Middle): Check for duplicates (exact vs. ID mismatches), count missing values per row and column, and identify outliers using aggregation functions like `skewness()` and `stddev()`—this groundwork determines model quality. - **MLlib and ML serve different needs** (Late): MLlib works on RDDs for legacy models, while ML is the modern DataFrame-based API with feature extraction, hyper-tuning (grid search, train-validation split), and models for classification, clustering, and regression. - **GraphFrames simplify graph analytics** (Late): With a flights dataset, you can run queries (delays, top transfer airports), compute vertex degrees, use PageRank for ranking, and find motifs—making graph problems accessible without specialized tools. 【Reading Tips】 - **Skim Chapter 1** if you're familiar with Spark basics; focus on the SparkSession unification and DataFrame vs. Dataset distinctions, as these are foundational for later chapters. - **Deep-read Chapters 2–3** on RDDs and DataFrames—they're the core of PySpark; practice the code examples locally to internalize transformations vs. actions and schema handling. - **Pay attention to Chapter 4** on data preparation; the missing-value and outlier techniques are directly reusable in real projects, even if you skip the ML chapters. - **Treat Chapters 5–6 as a reference** for MLlib and ML; skim model overviews and focus on the pipeline examples (feature extraction, grid search) rather than memorizing every algorithm. - **Watch for Spark version differences**: The book targets Spark 2.0, so note that Datasets are only in Scala/Java—check current PySpark docs for updates on this and streaming APIs. 【Coverage Limits】 Excerpts cover roughly the first half of the book (through data preparation and early ML), with chapter titles for later topics (GraphFrames, TensorFrames, Blaze, Structured Streaming) but limited detail—specific code and examples for those sections are not fully represented here.
Excerpt 1
essing using GraphFrames, and polyglot persistence using B Foreword Thank you for choosing this book to start your PySpark adventures, I hope you are as exci...
View in text
Excerpt 2
and features to Spark SQL and to allow external developers to extend the optimizer (for example, adding data source specific rules, support for new data type...
View in text
Excerpt 3
specifies the number of records to return, and the third is a seed to the pseudo-random numbers generator: data_take_sampled = data_from_file_conv.takeSample...
View in text
Excerpt 4
(1 - (fn.count(c) / fn.count('*'))).alias(c + '_missing') for c in df_miss.columns ]).show() [ 61 ] Prepare Data for Modeling The preceding code produces the...
View in text
Excerpt 5
omplex one. MLlib allows us to select the most predictable features using a Chi-Square selector. Here's how you do it: selector = ft.ChiSqSelector(4).fit(bir...
View in text
Excerpt 6
et looks like (abbreviated for brevity): [ 116 ] Chapter 6 First, we need to create a vector representation of our continuous variable (as it is only a singl...
View in text
Excerpt 7
ons: # display list of one-stop flights between SFO and BUF filteredPaths = tripGraph.bfs( fromExpr = "id = 'SFO'", toExpr = "id = 'BUF'", maxPathLength = 2)...
View in text
Excerpt 8
to create a library. Installing TensorFlow on your cluster In a notebook, run one of the following commands to install TensorFlow. This has been tested with...
View in text
Tags
AI categories
DataBig DataProgramming
ISBN: 1786463709
Publisher: Packt Publishing
Publish Year: 2017
Language: English
Pages: 274
File Format: PDF
File Size: 7.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…