Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Barr Moses, Lior Gavish, Molly Vorwerck

Rating No ratings yet

Do your product dashboards look funky? Are your quarterly reports stale? Is the data set you're using broken or just plain wrong? These problems affect almost every team, yet they're usually addressed on an ad hoc basis and in a reactive manner. If you answered yes to these questions, this book is for you. Many data engineering teams today face the "good pipelines, bad data" problem. It doesn't matter how advanced your data infrastructure is if the data you're piping is bad. In this book, Barr Moses, Lior Gavish, and Molly Vorwerck, from the data observability company Monte Carlo, explain how to tackle data quality and trust at scale by leveraging best practices and technologies used by some of the world's most innovative companies. • Build more trustworthy and reliable data pipelines • Write scripts to make data checks and identify broken pipelines with data observability • Learn how to set and maintain data SLAs, SLIs, and SLOs • Develop and lead data quality initiatives at your company • Learn how to treat data services and systems with the diligence of production software • Automate data lineage graphs across your data ecosystem • Build anomaly detectors for your critical data assets

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Data Quality Fundamentals: A Practitioner's Guide to Building Trustworthy Data Pipelines ## 【One-Line Pitch】 A practical playbook for data engineers and analytics leaders who are tired of "good pipelines, bad data" — this book shows you how to systematically prevent, detect, and resolve data quality issues by treating data with the same rigor as production software. If stale dashboards, broken reports, or silent data errors are costing your team time and trust, this is your field guide. ## 【Book Arc】 - **Opening (~0%–9%)**: Why data quality deserves urgent attention — the "data downtime" problem, its financial and reputational costs, and the organizational landscape of modern data teams. Sets up the core thesis: data quality is a systems problem, not a one-off fix. - **Early (~16%–28%)**: Building blocks of a reliable data system — walks through the modern data stack (warehouses, lakes, catalogs) and explains how each component contributes to or detracts from data quality. Introduces the three pillars: process, technology, and people. - **Early–Middle (~28%–38%)**: Hands-on monitoring and metrics collection — concrete SQL examples for tracking freshness, volume, schema, and health in Snowflake, plus guidance on what to monitor across streaming data, ML models, and dashboards. - **Middle (~44%–47%)**: The data pipeline lifecycle — collection, cleaning, transformation, and testing. Covers entrypoint design, outlier removal, feature assessment, normalization, and data reconstruction techniques with practical code-level advice. - **Late (~47%–end)**: Advanced topics and organizational change — data lineage automation, democratizing data quality, treating data as a product, and assigning ownership across roles (CDO, analytics engineer, data product manager). Includes case studies from Fox, Uber, and Convoy. ## 【Key Takeaways】 - **Data downtime is a measurable business cost** (Early): Poor data quality consumes up to 40% of data team time and has caused companies to lose customers — one 2019 study found one in five companies lost a customer due to a data quality issue. This frames data quality as a financial imperative, not just an engineering nicety. - **Prevention beats reaction** (Early): Most data downtime can be prevented with the right systems and processes — a single schema change or code push can break downstream reports, so designing for quality at every pipeline stage is essential. The book's framework splits solutions into process, technology, and people. - **Know your storage trade-offs** (Early): Data warehouses like Snowflake offer fast SQL querying and flexible pricing (separate compute/storage fees) but impose schema-on-write constraints, limited semi-structured data support, and SQL-only access — all of which create specific data quality risks. Match your storage choice to your actual workflow needs. - **Metadata is your quality radar** (Early–Middle): Data catalogs — whether built manually or via SQL parsers like Sqlparse, ANTLR, or Apache Calcite — provide the context needed to track location, ownership, and usage. Start by aligning with downstream stakeholders on which data matters most, then assign ownership by source, schema, or domain. - **Monitor five dimensions of data health** (Middle): Freshness, volume, schema, and completeness/distinctness checks — implemented via SQL queries against historical records — catch anomalies early. The book shows concrete query patterns for tracking row counts, null rates, string patterns, and value distributions over time. - **The entrypoint is where quality is won or lost** (Middle): The most upstream point where data enters your pipeline deserves the most attention — if bad data gets in, everything downstream inherits the problem. Cleaning techniques include outlier removal (statistical scoring, isolation forests), feature pruning, normalization (L1/L2 norms, demeaning), and reconstruction (interpolation, labeling). - **Data quality is an organizational design problem** (Late): Lasting solutions require assigning clear ownership — Chief Data Officer, analytics engineers, data product managers, and governance leads each have distinct roles — and balancing data accessibility with trust. Case studies from Fox and Uber show how decentralized teams and "controlled freedom" for stakeholders drive adoption. ## 【Reading Tips】 - **Skim the opening chapters (0–16%)** if you're already convinced data quality matters — the historical anecdotes (Mars Climate Orbiter, Antarctic explorers) are engaging but not actionable. Focus instead on the framework definitions and the three-pillar model. - **Deep-read the middle sections (28–47%)** for the hands-on SQL examples and pipeline techniques — these are the most immediately applicable parts. The Snowflake monitoring queries and cleaning techniques (outlier removal, normalization, reconstruction) translate directly to real work. - **Pay special attention to the data catalog and lineage discussions** — these are often underinvested areas in practice, and the book provides concrete starting points (spreadsheets for alignment, open-source SQL parsers for automation). - **Read the case studies (Fox, Uber, Convoy) for organizational patterns** — they're less about technology and more about how to navigate stakeholder buy-in, team structure, and the build-versus-buy decision. Useful if you're leading a data quality initiative. - **Watch for the recurring "treat data like production software" theme** — this mental model (SLAs, SLIs, SLOs, monitoring, incident response) is the book's core contribution and worth internalizing even if you skip the technical details. ## 【Coverage Limits】 Excerpts cover roughly the first half of the book in depth (through pipeline testing and cleaning techniques) plus table-of-contents-level detail on later chapters. The guide does not cover the full content of chapters on anomaly detection algorithms, field-level lineage implementation, or the detailed case study walkthroughs — these appear only as summaries in the source material. ##
Page 6
al, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/...
View in text
Excerpt 2
ntal, issues with their data. There had to be a better way. Poor data quality and unreliable data have been problems for organizations for decades, whether i...
View in text
Excerpt 3
dards as to the structure of your data. That structure will be constantly changing, and constant schema change is not something a data warehouse happily supp...
View in text
Excerpt 4
lked about with HTTP status codes, sometimes whole sections of your data are irrelevant for a downstream task. Throw them out! Granted, the cost of cloud sto...
View in text
Excerpt 5
lt atop Apache Spark, so it has a lot of format flexibility. Anything that can fit into a Spark DataFrame—CSV data, JSON, warehouse table data, application 6...
View in text
Excerpt 6
e distribution detectors we built earlier in the chapter to get the first date of appreciable zero rates in the habitability field, as depicted in Example 4-...
View in text
Excerpt 7
o? What determines the ‘sweet spot’ for my model parameters?” Choosing an Fβ score to optimize will implicitly decide how you weigh these occurrences, and th...
View in text
Excerpt 8
data became essential—not just for determining advertising spend, but also for understanding the current state of how users were interacting with the Blinkis...
View in text
Tags
AI categories
DataTechnologyBackend
ISBN: 1098112040
Publisher: O'Reilly Media
Publish Year: 2022
Language: English
Pages: 311
File Format: PDF
File Size: 9.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…