Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Danil Zburivsky, Lynda Partner

Rating No ratings yet

Centralized data warehouses, the long-time defacto standard for housing data for analytics, are rapidly giving way to multi-faceted cloud data platforms. Companies that embrace modern cloud data platforms benefit from an integrated view of their business using all of their data and can take advantage of advanced analytic practices to drive predictions and as yet unimagined data services. Designing Cloud Data Platforms is a hands-on guide to envisioning and designing a modern scalable data platform that takes full advantage of the flexibility of the cloud. As you read, you’ll learn the core components of a cloud data platform design, along with the role of key technologies like Spark and Kafka Streams. You’ll also explore setting up processes to manage cloud-based data, keep it secure, and using advanced analytic and BI tools to analyze it. About the Technology Well-designed pipelines, storage systems, and APIs eliminate the complicated scaling and maintenance required with on-prem data centers. Once you learn the patterns for designing cloud data platforms, you’ll maximize performance no matter which cloud vendor you use. About the book In Designing Cloud Data Platforms, Danil Zburivsky and Lynda Partner reveal a six-layer approach that increases flexibility and reduces costs. Discover patterns for ingesting data from a variety of sources, then learn to harness pre-built services provided by cloud vendors. What's inside • Best practices for structured and unstructured data sets • Cloud-ready machine learning tools • Metadata and real-time analytics • Defensive architecture, access, and security About the reader For data professionals familiar with the basics of cloud computing, and Hadoop or Spark. About the authors Danil Zburivsky has over 10 years of experience designing and supporting large-scale data infrastructure for enterprises across the globe. Lynda Partner is the VP of Analytics-as-a-Service at Pythian, and has been on the business side of data for over 20

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical, vendor-neutral guide for data professionals who want to design scalable, cost-effective cloud data platforms that handle batch and streaming data, integrate with data warehouses, and support advanced analytics and machine learning. 【Book Arc】 - **Opening (~0%–9%)**: Introduces the shift from traditional data warehouses and Hadoop-based data lakes to cloud-native data platforms, defining the core concept and the "three V's" (variety, volume, velocity) that drive the need for a new architecture. - **Early (~9%–28%)**: Compares cloud data platforms with cloud data warehouses in depth, using a hands-on example (Azure) to show how ingestion, schema handling, and processing differ—highlighting the flexibility of data lakes and the limitations of warehouse-only designs. - **Middle (~28%–47%)**: Expands the architecture into a six-layer model (ingestion, storage, processing, serving, metadata, and security), mapping each layer to tools from AWS, Azure, and Google Cloud, and discusses batch vs. streaming pipelines, including the lambda architecture and modern alternatives like Apache Beam and Kafka Streams. - **Late (~47%–75%)**: Dives into real-time processing, technical metadata management, schema evolution, and data serving—covering how to organize data for different consumers (data warehouse, applications, ML, BI tools) and how to handle schema changes gracefully. - **Ending (~75%–100%)**: Wraps up with organizational challenges and best practices for driving business value from a data platform, including governance, team structures, and project success factors. 【Key Takeaways】 - **Cloud data platforms solve the "three V's"** (variety, volume, velocity) that plague traditional warehouses and Hadoop-based lakes, by separating storage from compute and using cloud-native services (Early). - **Data warehouses alone are insufficient** for modern analytics; they struggle with schema changes and semistructured data like JSON, while data platforms with a lake layer offer far more flexibility (Early). - **Schema-on-read is a game-changer**: unlike warehouses that require strict schemas upfront, data platforms let you ingest raw data first and define structure later, making it much easier to handle evolving data sources (Early). - **Batch and streaming are both first-class citizens**: a well-designed platform supports both paths—batch via storage and processing layers, streaming via direct ingestion to processing—without forcing a lambda architecture's dual-pipeline complexity (Middle). - **The processing layer must scale and be flexible**: frameworks like Spark and Apache Beam handle both batch and real-time workloads, support multiple languages (Python, Java, Scala), and ideally offer a SQL interface for analyst productivity (Middle). - **Technical metadata is the backbone of a data platform**: tracking schema info, pipeline status, row counts, and lineage is essential for operational visibility and debugging, and it should be a central, cross-cutting layer (Middle). - **Schema evolution is manageable with the right design**: by avoiding strict schemas at ingestion and using tools that support schema-on-read, you can handle source changes without breaking pipelines (Late). - **Serving data to diverse consumers requires multiple access points**: a data warehouse for SQL-based BI tools, direct lake access for data scientists via Spark, and APIs for applications—each with its own trade-offs (Late). 【Reading Tips】 - **Skim the Azure-specific example in Chapter 2** if you're not using Azure; the key lesson is the conceptual difference between warehouse-only and platform designs, not the specific tooling. - **Deep-read the six-layer architecture discussion** (Chapter 3) as it's the core mental model of the book; map each layer to your preferred cloud vendor (AWS, Azure, GCP) for practical reference. - **Pay close attention to the batch vs. streaming distinction** (Chapters 3 and 6); this is where many real-world designs go wrong, and the book's clear separation of pipelines is a valuable pattern. - **The metadata and schema chapters (7–8) are dense but crucial**; if you're short on time, focus on the principles (why metadata matters, schema-on-read) rather than the specific implementation options. - **Skip the code-heavy SQL/JSON examples** if you're not a hands-on engineer; the conceptual takeaways are summarized clearly in the chapter summaries and exercises. 【Coverage Limits】 This guide covers the book's core architecture and design principles, but does not include detailed vendor-specific configuration steps, code walkthroughs, or the final chapter's organizational case studies in depth.
Excerpt 1
of cloud computing, and Hadoop or Spark. About the authors Danil Zburivsky has over 10 years of experience designing and supporting large-scale data infrastr...
View in text
Excerpt 2
swer came along—one that had the benefits of Hadoop, elimi- nated its shortcomings, and brought even more flexibility to designers of data systems. Along cam...
View in text
Excerpt 3
he schema of the source and destination tables. This infor- mation is required up front, meaning the schema must be available to the pipeline before the pipe...
View in text
Excerpt 4
source frameworks and cloud services that allow you to pro- cess data from both fast and slow storage at the same time. A good example is an open source Apac...
View in text
Excerpt 5
to- matically destroy the cluster once the job is completed. This type of elastic resource usage is one of the primary methods of cloud cost management. REAL...
View in text
Excerpt 6
last couple of days worth of data. This list goes on and on. This lack of standardization for data access and the resulting variety of API access methods mak...
View in text
Excerpt 7
purposes. MS SQL Server has a built-in capability that can extract change events from this log and populate a special “change table.” A change table is a reg...
View in text
Excerpt 8
can read messages from Azure Event Hubs and write them into Azure SQL Warehouse. On Google Cloud Platform, you can use Cloud Dataflow to read messages from C...
View in text
Tags
AI categories
Cloud NativeDataBackend
ISBN: 1617296449
Publish Year: 2021
Language: English
Pages: 336
File Format: PDF
File Size: 15.6 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…