The Enterprise Big Data Lake Delivering the Promise of Big Data and Data Science (Alex Gorelik) (Z-Library)
Other
No description
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
# The Enterprise Big Data Lake: Delivering the Promise of Big Data and Data Science
## 【One-Line Pitch】
A practical field guide for enterprise leaders and data professionals who want to build a data lake that actually works—not a data swamp—by covering architecture, governance, and self-service strategies. If you're a CDO, data architect, or analytics lead wrestling with big data platforms, this book gives you the playbook to deliver on big data's promises.
## 【Book Arc】
- **Opening (~0%–10%)**: Introduces the core problem—big data technologies deliver scalable storage but often fail to provide usable, governed access, resulting in "data swamps." Sets up the book's promise: how to deliver on all three promises of big data (cost-effective storage, tiered access, and self-service).
- **Early (~10%–23%)**: Defines the maturity model from data puddles to ponds to lakes to oceans, explaining how each stage serves different user communities. Covers the four stages of analysis (find, understand, provision, prepare) and introduces the zone-based architecture (raw, gold, work, sensitive).
- **Early–Middle (~23%–39%)**: Provides historical context on how we got here—from file systems to relational databases to data warehouses—explaining why traditional approaches fail for modern analytics. Introduces ETL vs. ELT and the data warehouse ecosystem's metadata flow.
- **Middle (~39%–48%)**: Dives into the data warehouse ecosystem's tooling: data quality rules (scalar, field-level), data governance tools, and data stewardship roles. Explains how these disciplines carry forward into data lake design.
- **Middle (~48%–52%+)**: Covers big data fundamentals—MapReduce, schema-on-read vs. schema-on-write—and how these paradigms shift the analytics workflow. The excerpts begin exploring how these technologies enable new approaches to data preparation and analysis.
## 【Key Takeaways】
- **Data lakes fail without governance** (Opening): The technology delivers storage and compute, but without careful management, you get a "data swamp"—unusable, unnavigable, and dangerous for decisions. Governance isn't optional; it's the difference between a lake and a swamp.
- **Maturity is a spectrum, not a destination** (Early): Data puddles (analytical sandboxes) and ponds (big data warehouses) are legitimate starting points. The key distinction is focus: ponds run routine production queries, while lakes enable ad hoc experimentation with new data types and tools.
- **Zones are the organizing principle** (Early): Raw, gold, work, and sensitive zones serve different communities—business analysts live in gold, data engineers work in raw, data scientists experiment in work. Governance intensity varies by zone, with stewards focusing on sensitive and gold areas.
- **Self-service requires a four-step workflow** (Early): Analysts need to find, understand, provision, and prepare data. Each step needs tooling and processes—from catalogs with facets and ranking to deidentification for sensitive data—or self-service collapses.
- **Data preparation is sophisticated work** (Early): Shaping, cleaning, and blending are non-trivial operations. Excel doesn't scale; automation is crucial to avoid repeating tedious steps across thousands of tables. This is where much of the real value (and pain) lives.
- **The data warehouse legacy shapes everything** (Middle): Understanding ETL vs. ELT, normalization trade-offs, and the metadata flow from warehouse ecosystems is essential. Data lakes don't replace this history; they build on it—and inherit its quality and governance challenges.
- **Data quality rules are layered** (Middle): Scalar rules (value-level), field-level rules (uniqueness, density), and cross-field rules form a hierarchy. Data stewardship is complex and cross-functional—identifying who owns what is the governance tool's most important job.
- **MapReduce teaches scalability lessons** (Middle): Parallel processing works only if work is distributed properly—one slow mapper can bottleneck the entire job. Schema-on-read flips the relational model, allowing data to land first and be interpreted later, which is liberating but requires new discipline.
## 【Reading Tips】
- **Skim the historical chapters (Chunk #10–#11)** if you already know relational database basics; the normalization and primary/foreign key review is standard material. But don't skip the ETL vs. ELT discussion—it's crucial context for data lake architecture decisions.
- **Deep-read the zone architecture and self-service sections (Chunks #7–#8)**; these are the book's practical core. The four stages of analysis (find, understand, provision, prepare) are a framework you'll want to internalize and apply to your own organization.
- **Pay attention to the data swamp warning signs** (Chunk #7): a pond that grows to lake size but fails to attract analysts is the failure mode to avoid. The book's advice on catalogs, facets, and contextual search is directly actionable.
- **The data quality and governance sections (Chunks #13–#14)** are dense but worth the effort—they explain the tooling landscape you'll need to evaluate. The scalar/field-level rule taxonomy is a useful checklist for assessing your own data quality program.
- **If you're a business leader rather than a technologist**, focus on the maturity model (Chunk #5) and the zone governance expectations (Chunk #7); you can skim the MapReduce and schema-on-read details (Chunk #16) without losing the plot.
## 【Coverage Limits】
The excerpts cover roughly the first half of the book (through ~52%), including introduction, historical perspective, and early big data fundamentals. Later chapters on architecting the data lake in detail, self-service optimization, and implementation roadmaps are only partially represented—the full book likely contains deeper practical guidance on those topics.
##
Excerpt 1
112 Data Wrangling in the Data Lake 113 Situating Data Preparation in Hadoop 113 Common Use Cases for Data Preparation 114 Analyzing and Visualizing 116 The...
View in text
Excerpt 2
enterprises today is thrown away. Some small percentage is aggregated and kept in a data warehouse for a few years, but most detailed opera‐ tional data, mac...
View in text
Excerpt 3
still designed to support applications. Because writing or reading data to and from disk was orders of magnitude slower than processing it in memory, a schem...
View in text
Excerpt 4
consulted and can authorize access and other data policies. Once ownership has been documented, the next step in rolling out a data governance program is to...
View in text
Excerpt 5
accurately train the model, we would need data from a rep‐ resentative mix of high- and low-performing school districts. Feature engineering is one of the mo...
View in text
Excerpt 6
opportunity of using Hadoop as a cost-effective platform to provide centralized governance and compliance for the enterprise. Traditionally, governance has r...
View in text
Excerpt 7
a malfunction occurs, it will need to be handled right away. In addition, the behav‐ ior leading up to it—sometimes over many days, months, or even years—sho...
View in text
Excerpt 8
f customers for an email campaign by doing customer segmen‐ tation, or producing reports to estimate house prices. Figure 5-10. Results of processing in the...
View in text
Tags
AI categories
Big DataDataBackend
Text Preview (First 20 pages)
Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Generating text preview…
Loading comments...
Reply to Comment
Edit Comment