Build data-intensive applications locally and deploy at scale using the combined powers of Python and Spark 2.0 About This Book Learn why and how you can efficiently use Python to process data and build machine learning models in Apache Spark 2.0 Develop and deploy efficient, scalable real-time Spark solutions Take your understanding of using Spark with Python to the next level with this jump start guide Who This Book Is For If you are a Python developer who wants to learn about the Apache Spark 2.0 ecosystem, this book is for you. A firm understanding of Python is expected to get the best out of the book. Familiarity with Spark would be useful, but is not mandatory. What You Will Learn Learn about Apache Spark and the Spark 2.0 architecture Build and interact with Spark DataFrames using Spark SQL Learn how to solve graph and deep learning problems using GraphFrames and TensorFrames respectively Read, transform, and understand data and use it to train machine learning models Build machine learning models with MLlib and ML Learn how to submit your applications programmatically using spark-submit Deploy locally built applications to a cluster In Detail Apache Spark is an open source framework for efficient cluster computing with a strong interface for data parallelism and fault tolerance. This book will show you how to leverage the power of Python and put it to use in the Spark ecosystem. You will start by getting a firm understanding of the Spark 2.0 architecture and how to set up a Python environment for Spark. You will get familiar with the modules available in PySpark. You will learn how to abstract data with RDDs and DataFrames and understand the streaming capabilities of PySpark. Also, you will get a thorough overview of machine learning capabilities of PySpark using ML and MLlib, graph processing using GraphFrames, and polyglot persistence using B
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A hands-on guide for Python developers to master Apache Spark 2.0, covering everything from core architecture and data abstractions to machine learning, graph processing, and streaming—ideal for anyone wanting to build scalable data applications without leaving Python.
【Book Arc】
- **Opening (~0%–9%)**: Introduces Spark 2.0's architecture, the unified SparkSession, and the core abstractions (RDDs, DataFrames, Datasets), explaining how Catalyst Optimizer and Project Tungsten deliver performance parity across languages.
- **Early (~9%–28%)**: Dives into Resilient Distributed Datasets (RDDs)—the schema-less foundation—covering creation from various sources, transformations like `.map()`, and actions like `.take()` and `.collect()`, with practical warnings about partitioning and executor behavior.
- **Early (~28%–38%)**: Shifts to DataFrames, showing how to create them from JSON, infer schemas via reflection, or specify schemas programmatically, plus SQL-like querying and RDD interoperation.
- **Middle (~38%–47%)**: Focuses on data preparation for modeling—handling duplicates, missing values, and outliers with aggregation functions, and emphasizes the importance of understanding your data before building models.
- **Late (~47%–100%)**: Covers the MLlib and ML packages for machine learning (classification, clustering, regression), GraphFrames for graph analytics, TensorFrames for deep learning, Blaze for polyglot persistence, and Structured Streaming for real-time processing.
【Key Takeaways】
- **SparkSession unifies all contexts** (Early): In Spark 2.0, SparkConf, SparkContext, SQLContext, and HiveContext merge into a single entry point, simplifying configuration and resource management—use it for reading data, metadata, and cluster control.
- **DataFrames bring performance parity** (Early): Unlike RDDs, which can be slow in Python, DataFrames leverage Catalyst Optimizer and Project Tungsten for fast in-memory encoding, making Python competitive with Scala/Java—critical for production workloads.
- **RDDs are schema-less but flexible** (Early): Transformations like `.map()` and `.filter()` shape data lazily, while actions like `.take(n)` execute jobs—prefer `.take()` over `.collect()` for large datasets to avoid moving everything to the driver.
- **Schema inference vs. programmatic schemas** (Middle): Use reflection for concise code when schema is known; specify schemas programmatically when columns and types are only known at runtime—essential for dynamic data sources.
- **Data cleaning is a multi-step process** (Middle): Check for duplicates (exact vs. ID mismatches), count missing values per row and column, and identify outliers using aggregation functions like `skewness()` and `stddev()`—this groundwork determines model quality.
- **MLlib and ML serve different needs** (Late): MLlib works on RDDs for legacy models, while ML is the modern DataFrame-based API with feature extraction, hyper-tuning (grid search, train-validation split), and models for classification, clustering, and regression.
- **GraphFrames simplify graph analytics** (Late): With a flights dataset, you can run queries (delays, top transfer airports), compute vertex degrees, use PageRank for ranking, and find motifs—making graph problems accessible without specialized tools.
【Reading Tips】
- **Skim Chapter 1** if you're familiar with Spark basics; focus on the SparkSession unification and DataFrame vs. Dataset distinctions, as these are foundational for later chapters.
- **Deep-read Chapters 2–3** on RDDs and DataFrames—they're the core of PySpark; practice the code examples locally to internalize transformations vs. actions and schema handling.
- **Pay attention to Chapter 4** on data preparation; the missing-value and outlier techniques are directly reusable in real projects, even if you skip the ML chapters.
- **Treat Chapters 5–6 as a reference** for MLlib and ML; skim model overviews and focus on the pipeline examples (feature extraction, grid search) rather than memorizing every algorithm.
- **Watch for Spark version differences**: The book targets Spark 2.0, so note that Datasets are only in Scala/Java—check current PySpark docs for updates on this and streaming APIs.
【Coverage Limits】
Excerpts cover roughly the first half of the book (through data preparation and early ML), with chapter titles for later topics (GraphFrames, TensorFrames, Blaze, Structured Streaming) but limited detail—specific code and examples for those sections are not fully represented here.
Excerpt 1
essing using GraphFrames, and polyglot persistence using B Foreword Thank you for choosing this book to start your PySpark adventures, I hope you are as exci...
and features to Spark SQL and to allow external developers to extend the optimizer (for example, adding data source specific rules, support for new data type...
specifies the number of records to return, and the third is a seed to the pseudo-random numbers generator: data_take_sampled = data_from_file_conv.takeSample...
(1 - (fn.count(c) / fn.count('*'))).alias(c + '_missing') for c in df_miss.columns ]).show() [ 61 ] Prepare Data for Modeling The preceding code produces the...
omplex one. MLlib allows us to select the most predictable features using a Chi-Square selector. Here's how you do it: selector = ft.ChiSqSelector(4).fit(bir...
et looks like (abbreviated for brevity): [ 116 ] Chapter 6 First, we need to create a vector representation of our continuous variable (as it is only a singl...
to create a library. Installing TensorFlow on your cluster In a notebook, run one of the following commands to install TensorFlow. This has been tested with...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Loading comments...
Reply to Comment
Edit Comment