Centralized data warehouses, the long-time defacto standard for housing data for analytics, are rapidly giving way to multi-faceted cloud data platforms. Companies that embrace modern cloud data platforms benefit from an integrated view of their business using all of their data and can take advantage of advanced analytic practices to drive predictions and as yet unimagined data services. Designing Cloud Data Platforms is a hands-on guide to envisioning and designing a modern scalable data platform that takes full advantage of the flexibility of the cloud. As you read, you’ll learn the core components of a cloud data platform design, along with the role of key technologies like Spark and Kafka Streams. You’ll also explore setting up processes to manage cloud-based data, keep it secure, and using advanced analytic and BI tools to analyze it.
About the Technology
Well-designed pipelines, storage systems, and APIs eliminate the complicated scaling and maintenance required with on-prem data centers. Once you learn the patterns for designing cloud data platforms, you’ll maximize performance no matter which cloud vendor you use.
About the book
In Designing Cloud Data Platforms, Danil Zburivsky and Lynda Partner reveal a six-layer approach that increases flexibility and reduces costs. Discover patterns for ingesting data from a variety of sources, then learn to harness pre-built services provided by cloud vendors.
What's inside
• Best practices for structured and unstructured data sets
• Cloud-ready machine learning tools
• Metadata and real-time analytics
• Defensive architecture, access, and security
About the reader
For data professionals familiar with the basics of cloud computing, and Hadoop or Spark.
About the authors
Danil Zburivsky has over 10 years of experience designing and supporting large-scale data infrastructure for enterprises across the globe.
Lynda Partner is the VP of Analytics-as-a-Service at Pythian, and has been on the business side of data for over 20
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical, vendor-neutral guide for data professionals who want to design scalable, cost-effective cloud data platforms that handle batch and streaming data, integrate with data warehouses, and support advanced analytics and machine learning.
【Book Arc】
- **Opening (~0%–9%)**: Introduces the shift from traditional data warehouses and Hadoop-based data lakes to cloud-native data platforms, defining the core concept and the "three V's" (variety, volume, velocity) that drive the need for a new architecture.
- **Early (~9%–28%)**: Compares cloud data platforms with cloud data warehouses in depth, using a hands-on example (Azure) to show how ingestion, schema handling, and processing differ—highlighting the flexibility of data lakes and the limitations of warehouse-only designs.
- **Middle (~28%–47%)**: Expands the architecture into a six-layer model (ingestion, storage, processing, serving, metadata, and security), mapping each layer to tools from AWS, Azure, and Google Cloud, and discusses batch vs. streaming pipelines, including the lambda architecture and modern alternatives like Apache Beam and Kafka Streams.
- **Late (~47%–75%)**: Dives into real-time processing, technical metadata management, schema evolution, and data serving—covering how to organize data for different consumers (data warehouse, applications, ML, BI tools) and how to handle schema changes gracefully.
- **Ending (~75%–100%)**: Wraps up with organizational challenges and best practices for driving business value from a data platform, including governance, team structures, and project success factors.
【Key Takeaways】
- **Cloud data platforms solve the "three V's"** (variety, volume, velocity) that plague traditional warehouses and Hadoop-based lakes, by separating storage from compute and using cloud-native services (Early).
- **Data warehouses alone are insufficient** for modern analytics; they struggle with schema changes and semistructured data like JSON, while data platforms with a lake layer offer far more flexibility (Early).
- **Schema-on-read is a game-changer**: unlike warehouses that require strict schemas upfront, data platforms let you ingest raw data first and define structure later, making it much easier to handle evolving data sources (Early).
- **Batch and streaming are both first-class citizens**: a well-designed platform supports both paths—batch via storage and processing layers, streaming via direct ingestion to processing—without forcing a lambda architecture's dual-pipeline complexity (Middle).
- **The processing layer must scale and be flexible**: frameworks like Spark and Apache Beam handle both batch and real-time workloads, support multiple languages (Python, Java, Scala), and ideally offer a SQL interface for analyst productivity (Middle).
- **Technical metadata is the backbone of a data platform**: tracking schema info, pipeline status, row counts, and lineage is essential for operational visibility and debugging, and it should be a central, cross-cutting layer (Middle).
- **Schema evolution is manageable with the right design**: by avoiding strict schemas at ingestion and using tools that support schema-on-read, you can handle source changes without breaking pipelines (Late).
- **Serving data to diverse consumers requires multiple access points**: a data warehouse for SQL-based BI tools, direct lake access for data scientists via Spark, and APIs for applications—each with its own trade-offs (Late).
【Reading Tips】
- **Skim the Azure-specific example in Chapter 2** if you're not using Azure; the key lesson is the conceptual difference between warehouse-only and platform designs, not the specific tooling.
- **Deep-read the six-layer architecture discussion** (Chapter 3) as it's the core mental model of the book; map each layer to your preferred cloud vendor (AWS, Azure, GCP) for practical reference.
- **Pay close attention to the batch vs. streaming distinction** (Chapters 3 and 6); this is where many real-world designs go wrong, and the book's clear separation of pipelines is a valuable pattern.
- **The metadata and schema chapters (7–8) are dense but crucial**; if you're short on time, focus on the principles (why metadata matters, schema-on-read) rather than the specific implementation options.
- **Skip the code-heavy SQL/JSON examples** if you're not a hands-on engineer; the conceptual takeaways are summarized clearly in the chapter summaries and exercises.
【Coverage Limits】
This guide covers the book's core architecture and design principles, but does not include detailed vendor-specific configuration steps, code walkthroughs, or the final chapter's organizational case studies in depth.
Excerpt 1
of cloud computing, and Hadoop or Spark. About the authors Danil Zburivsky has over 10 years of experience designing and supporting large-scale data infrastr...
swer came along—one that had the benefits of Hadoop, elimi- nated its shortcomings, and brought even more flexibility to designers of data systems. Along cam...
he schema of the source and destination tables. This infor- mation is required up front, meaning the schema must be available to the pipeline before the pipe...
source frameworks and cloud services that allow you to pro- cess data from both fast and slow storage at the same time. A good example is an open source Apac...
to- matically destroy the cluster once the job is completed. This type of elastic resource usage is one of the primary methods of cloud cost management. REAL...
last couple of days worth of data. This list goes on and on. This lack of standardization for data access and the resulting variety of API access methods mak...
purposes. MS SQL Server has a built-in capability that can extract change events from this log and populate a special “change table.” A change table is a reg...
can read messages from Azure Event Hubs and write them into Azure SQL Warehouse. On Google Cloud Platform, you can use Cloud Dataflow to read messages from C...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Designing Cloud Data Platforms (Danil Zburivsky, Lynda Partner)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Designing Cloud Data Platforms (Danil Zburivsky, Lynda Partner)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment