Share E-Book
Scan to open this page

Scan with your phone to open this page

AuthorDaniel Vaughan

This practical guide provides a collection of techniques and best practices that are generally overlooked in most data engineering and data science pedagogy. A common misconception is that great data scientists are experts in the "big themes" of the discipline—machine learning and programming. But most of the time, these tools can only take us so far. In practice, the smaller tools and skills really separate a great data scientist from a not-so-great one. Taken as a whole, the lessons in this book make the difference between an average data scientist candidate and a qualified data scientist working in the field. Author Daniel Vaughan has collected, extended, and used these skills to create value and train data scientists from different companies and industries. With this book, you will: Understand how data science creates value Deliver compelling narratives to sell your data science project Build a business case using unit economics principles Create new features for a ML model using storytelling Learn how to decompose KPIs Perform growth decompositions to find root causes for changes in a metric Daniel Vaughan is head of data at Clip, the leading paytech company in Mexico. He's the author of Analytical Skills for AI and Data Science (O'Reilly).

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Data Science: The Hard Parts — Reading Guide ## 【One-Line Pitch】 A practical field manual for data scientists who want to move beyond ML algorithms and programming to master the "hard parts" of the discipline—business value creation, stakeholder communication, metric design, and causal reasoning. Read this if you're transitioning from academic or Kaggle-style data science into a role where your work must actually move business metrics. ## 【Book Arc】 - **Opening (~0%–10%)**: The book opens by challenging the common misconception that great data scientists are defined by mastery of machine learning and programming. Vaughan argues that the "big themes" only take you so far—the smaller, often-overlooked skills around business impact and communication are what separate average candidates from exceptional practitioners. - **Early (~10%–35%)**: The focus shifts to value creation and workflow design. Vaughan emphasizes ensuring your data science workflow actually creates value, designing actionable/timely/relevant metrics, and delivering compelling narratives to gain stakeholder buy-in. This stage is about reframing data science as a business function, not just a technical one. - **Middle (~35%–65%)**: The book moves into more technical but still under-taught territory: using simulation to validate whether an ML algorithm is the right tool for a problem, identifying/correcting/preventing data leakage, and understanding incrementality through causal effect estimation. This is where the "hard parts" of modeling rigor come in. - **Late (~65%–85%)**: Vaughan connects the earlier material to real-world decision-making processes. The book covers how data insights actually drive decisions in practice—covering domains from economics to advertising to epidemiology—and how to build business cases using unit economics principles. - **Ending (~85%–100%)**: The final stretch ties together the full arc: from creating features for ML models using storytelling, to decomposing KPIs, to performing growth decompositions to find root causes for metric changes. The book closes with the message that these collected skills, used together, make the difference between an average and an exceptional data scientist. ## 【Key Takeaways】 - **The "big themes" aren't enough** (Early): Mastery of ML and programming is table stakes; the differentiator is the ability to impact the business. This reframing should change how you prioritize skill development. - **Value creation is the north star** (Early): Every data science workflow should be designed with value creation in mind—not just model accuracy. This means asking "so what?" at every step. - **Metrics must be actionable, timely, and relevant** (Early): Designing good metrics is a skill in itself. A metric that isn't actionable or timely won't drive decisions, no matter how well-measured it is. - **Simulation validates tool choice** (Middle): Before committing to an ML algorithm, use simulation to verify it's the right tool for the problem. This prevents wasted effort on sophisticated solutions to problems that don't need them. - **Data leakage is a silent killer** (Middle): Identifying, correcting, and preventing data leakage is one of the most under-taught but critical skills in applied ML. Leakage inflates performance metrics and leads to models that fail in production. - **Causal thinking is essential** (Middle): Understanding incrementality by estimating causal effects—not just correlations—is what separates models that inform decisions from models that mislead them. - **Storytelling is a technical skill** (Late): Creating new features for ML models using storytelling, and delivering compelling narratives to stakeholders, are as important as the modeling itself. Communication is part of the craft. - **KPI decomposition finds root causes** (Ending): Decomposing KPIs and performing growth decompositions lets you trace metric changes to their root causes—a core skill for diagnosing business problems. ## 【Reading Tips】 - **Skim the front matter** (~0%–10%): The opening repeats the book's thesis multiple times. Read it once, absorb the message, and move on—the real content starts with the value-creation chapters. - **Deep-read the simulation and leakage chapters** (Middle): These are the most technical and most likely to be new material for many readers. Take notes and consider working through the examples. - **Pay special attention to the causal inference material** (Middle): This is where the book earns its "hard parts" title. If you're unfamiliar with causal effects and incrementality, budget extra time here. - **Treat the KPI decomposition and growth decomposition sections as reference material** (Ending): These are practical techniques you'll want to return to when facing real business problems, not just read once. - **Read with your own projects in mind**: The book is most valuable when you apply each technique to a current work problem. Consider keeping a running list of how each chapter's lessons apply to your situation. ## 【Coverage Limits】 The excerpts cover the book's thesis, table of contents, endorsements, and front/back matter, but do not include detailed chapter content. Specific techniques, code examples, and case studies are referenced but not fully visible in the source material. ##
Excerpt 1
书名: Data Science The Hard Parts Techniques for Excelling at Data Science (Daniel Vaughan) (Z-Library) 作者: Daniel Vaughan This practical guide provides a coll...
View in text
Page 2
. But most of the time, these tools can only take us so far. In reality, it’s the nuances within these larger themes, and the ability to impact the business,...
View in text
Excerpt 3
tist’s bookshelf.” —Brett Holleman Freelance data scientist Daniel Vaughan Data Science: The Hard Parts Techniques for Excelling at Data Science Boston Farnh...
View in text
Page 4
and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open sourc...
View in text
Tags
AI categories
DataTechnologyBackend
ISBN: 1098146476
Publisher: O'Reilly Media
Publish Year: 2023
Language: English
Pages: 257
File Format: PDF
File Size: 2.9 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…