Do your product dashboards look funky? Are your quarterly reports stale? Is the data set you're using broken or just plain wrong? These problems affect almost every team, yet they're usually addressed on an ad hoc basis and in a reactive manner. If you answered yes to these questions, this book is for you.
Many data engineering teams today face the "good pipelines, bad data" problem. It doesn't matter how advanced your data infrastructure is if the data you're piping is bad. In this book, Barr Moses, Lior Gavish, and Molly Vorwerck, from the data observability company Monte Carlo, explain how to tackle data quality and trust at scale by leveraging best practices and technologies used by some of the world's most innovative companies.
• Build more trustworthy and reliable data pipelines
• Write scripts to make data checks and identify broken pipelines with data observability
• Learn how to set and maintain data SLAs, SLIs, and SLOs
• Develop and lead data quality initiatives at your company
• Learn how to treat data services and systems with the diligence of production software
• Automate data lineage graphs across your data ecosystem
• Build anomaly detectors for your critical data assets
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Data Quality Fundamentals: A Practitioner's Guide to Building Trustworthy Data Pipelines
## 【One-Line Pitch】
A practical playbook for data engineers and analytics leaders who are tired of "good pipelines, bad data" — this book shows you how to systematically prevent, detect, and resolve data quality issues by treating data with the same rigor as production software. If stale dashboards, broken reports, or silent data errors are costing your team time and trust, this is your field guide.
## 【Book Arc】
- **Opening (~0%–9%)**: Why data quality deserves urgent attention — the "data downtime" problem, its financial and reputational costs, and the organizational landscape of modern data teams. Sets up the core thesis: data quality is a systems problem, not a one-off fix.
- **Early (~16%–28%)**: Building blocks of a reliable data system — walks through the modern data stack (warehouses, lakes, catalogs) and explains how each component contributes to or detracts from data quality. Introduces the three pillars: process, technology, and people.
- **Early–Middle (~28%–38%)**: Hands-on monitoring and metrics collection — concrete SQL examples for tracking freshness, volume, schema, and health in Snowflake, plus guidance on what to monitor across streaming data, ML models, and dashboards.
- **Middle (~44%–47%)**: The data pipeline lifecycle — collection, cleaning, transformation, and testing. Covers entrypoint design, outlier removal, feature assessment, normalization, and data reconstruction techniques with practical code-level advice.
- **Late (~47%–end)**: Advanced topics and organizational change — data lineage automation, democratizing data quality, treating data as a product, and assigning ownership across roles (CDO, analytics engineer, data product manager). Includes case studies from Fox, Uber, and Convoy.
## 【Key Takeaways】
- **Data downtime is a measurable business cost** (Early): Poor data quality consumes up to 40% of data team time and has caused companies to lose customers — one 2019 study found one in five companies lost a customer due to a data quality issue. This frames data quality as a financial imperative, not just an engineering nicety.
- **Prevention beats reaction** (Early): Most data downtime can be prevented with the right systems and processes — a single schema change or code push can break downstream reports, so designing for quality at every pipeline stage is essential. The book's framework splits solutions into process, technology, and people.
- **Know your storage trade-offs** (Early): Data warehouses like Snowflake offer fast SQL querying and flexible pricing (separate compute/storage fees) but impose schema-on-write constraints, limited semi-structured data support, and SQL-only access — all of which create specific data quality risks. Match your storage choice to your actual workflow needs.
- **Metadata is your quality radar** (Early–Middle): Data catalogs — whether built manually or via SQL parsers like Sqlparse, ANTLR, or Apache Calcite — provide the context needed to track location, ownership, and usage. Start by aligning with downstream stakeholders on which data matters most, then assign ownership by source, schema, or domain.
- **Monitor five dimensions of data health** (Middle): Freshness, volume, schema, and completeness/distinctness checks — implemented via SQL queries against historical records — catch anomalies early. The book shows concrete query patterns for tracking row counts, null rates, string patterns, and value distributions over time.
- **The entrypoint is where quality is won or lost** (Middle): The most upstream point where data enters your pipeline deserves the most attention — if bad data gets in, everything downstream inherits the problem. Cleaning techniques include outlier removal (statistical scoring, isolation forests), feature pruning, normalization (L1/L2 norms, demeaning), and reconstruction (interpolation, labeling).
- **Data quality is an organizational design problem** (Late): Lasting solutions require assigning clear ownership — Chief Data Officer, analytics engineers, data product managers, and governance leads each have distinct roles — and balancing data accessibility with trust. Case studies from Fox and Uber show how decentralized teams and "controlled freedom" for stakeholders drive adoption.
## 【Reading Tips】
- **Skim the opening chapters (0–16%)** if you're already convinced data quality matters — the historical anecdotes (Mars Climate Orbiter, Antarctic explorers) are engaging but not actionable. Focus instead on the framework definitions and the three-pillar model.
- **Deep-read the middle sections (28–47%)** for the hands-on SQL examples and pipeline techniques — these are the most immediately applicable parts. The Snowflake monitoring queries and cleaning techniques (outlier removal, normalization, reconstruction) translate directly to real work.
- **Pay special attention to the data catalog and lineage discussions** — these are often underinvested areas in practice, and the book provides concrete starting points (spreadsheets for alignment, open-source SQL parsers for automation).
- **Read the case studies (Fox, Uber, Convoy) for organizational patterns** — they're less about technology and more about how to navigate stakeholder buy-in, team structure, and the build-versus-buy decision. Useful if you're leading a data quality initiative.
- **Watch for the recurring "treat data like production software" theme** — this mental model (SLAs, SLIs, SLOs, monitoring, incident response) is the book's core contribution and worth internalizing even if you skip the technical details.
## 【Coverage Limits】
Excerpts cover roughly the first half of the book in depth (through pipeline testing and cleaning techniques) plus table-of-contents-level detail on later chapters. The guide does not cover the full content of chapters on anomaly detection algorithms, field-level lineage implementation, or the detailed case study walkthroughs — these appear only as summaries in the source material.
##
Page 6
al, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/...
ntal, issues with their data. There had to be a better way. Poor data quality and unreliable data have been problems for organizations for decades, whether i...
dards as to the structure of your data. That structure will be constantly changing, and constant schema change is not something a data warehouse happily supp...
lked about with HTTP status codes, sometimes whole sections of your data are irrelevant for a downstream task. Throw them out! Granted, the cost of cloud sto...
lt atop Apache Spark, so it has a lot of format flexibility. Anything that can fit into a Spark DataFrame—CSV data, JSON, warehouse table data, application 6...
e distribution detectors we built earlier in the chapter to get the first date of appreciable zero rates in the habitability field, as depicted in Example 4-...
o? What determines the ‘sweet spot’ for my model parameters?” Choosing an Fβ score to optimize will implicitly decide how you weigh these occurrences, and th...
data became essential—not just for determining advertising spend, but also for understanding the current state of how users were interacting with the Blinkis...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Data Quality Fundamentals A Practitioners Guide to Building Trustworthy Data Pipelines (Barr Moses, Lior Gavish, Molly Vorwerck) (Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Data Quality Fundamentals A Practitioners Guide to Building Trustworthy Data Pipelines (Barr Moses, Lior Gavish, Molly Vorwerck) (Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment