Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: David Tan, Ada Leung, David Colls

Rating No ratings yet

Gain the valuable skills and techniques you need to accelerate the delivery of machine learning solutions. With this practical guide, data scientists and ML engineers will learn how to bridge the gap between data science and Lean software delivery in a practical and simple way. David Tan and Ada Leung from Thoughtworks show you how to apply time-tested software engineering skills and Lean delivery practices that will improve your effectiveness in ML projects. Based on the authors’ experience across multiple real-world data and ML projects, the proven techniques in this book will help teams avoid common traps in the ML world, so you can iterate more quickly and reliably. With these techniques, data scientists and ML engineers can overcome friction and experience flow when delivering machine learning solutions. This book shows you how to: • Apply engineering practices such as writing automated tests, containerizing development environments, and refactoring problematic code bases • Apply MLOps and CI/CD practices to accelerate experimentation cycles and improve reliability of ML solutions • Design maintainable and evolvable ML solutions that allow you to respond to changes in an agile fashion • Apply delivery and product practices to iteratively improve your odds of building the right product for your users • Use intelligent code editor features to code more effectively

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical field guide for data scientists and ML engineers who want to stop fighting their tooling and start shipping: it applies proven software engineering, Lean delivery, and product discovery practices to the messy reality of machine learning projects. Read it if you work on an ML team and suspect your bottlenecks are organizational and engineering, not algorithmic. 【Book Arc】 - **Opening (~0%–10%)**: Sets the thesis — ML projects fail for systemic reasons (waste, handoffs, weak feedback loops), not just modeling skill — and previews the book's five disciplines: product, delivery, software engineering, data, and ML. - **Early (~10%–35%)**: Diagnoses everyday impediments through the story of Dana, an ML engineer drowning in alerts, manual testing, and slow retraining loops; introduces Lean thinking, cross-functional "vertically sliced" teams, and the Inverse Conway Maneuver. - **Middle (~35%–60%)**: Shifts to product and delivery: Double Diamond and Design Thinking discovery, hypotheses and experiments, user stories as the unit of shippable value, spikes, and inception practices that set teams up before coding. - **Late (~60%–90%)**: The engineering core — reproducible dependency management, automated testing for software, models, and data (including LLM testing), editor/IDE techniques, refactoring and technical debt, then MLOps and CD4ML for continuous delivery. - **Ending (~90%–100%)**: Zooms out to team effectiveness and organizations: feedback loops, cognitive load, flow state, Team Topologies for ML, portfolio management, and leadership tactics like psychological safety and embracing failure. 【Key Takeaways】 - **ML delivery is a systems problem, not a modeling problem** (Early): the book's central claim is that most friction comes from waste, handoffs, and missing feedback loops, which Lean thinking exposes and reduces. - **Cross-functional, vertically sliced teams beat functional silos** (Early): splitting data science, data engineering, and product engineering creates backlog coupling and Conway's Law artifacts; cohering capabilities into one team speeds decisions and cuts waiting. - **Product discovery must run alongside delivery** (Middle): because nobody knows a priori what an ML product can do, dual-track delivery, hypotheses, and spikes let teams test desirability and feasibility continuously rather than betting everything upfront. - **User stories are the quantum of shippable value** (Middle): breaking large ideas into small, independently shippable stories — including timeboxed spikes — links product progress to project progress and keeps scope honest. - **Reproducible environments are a prerequisite, not a luxury** (Late): consistent, production-like, secure dependency management lets teammates "check out and go" instead of sinking into dependency hell. - **Testing must extend beyond software to models and data** (Late): automated tests shorten feedback cycles, but software testing paradigms have limits on ML models — fitness functions and behavioral tests scale model validation, and LLM applications need their own techniques. - **Refactoring and technical debt management are ongoing disciplines** (Late): readable, testable, evolvable code is what lets ML solutions respond to change; the book works through a messy notebook as a concrete case. - **MLOps and CD4ML close the loop from experiment to production** (Late): continuous delivery practices accelerate experimentation cycles and improve reliability, though the excerpts note MLOps has strengths and missing puzzle pieces. - **Team health is an engineering concern** (Ending): feedback loops, cognitive load, and flow state are treated as first-class design considerations, supported by Team Topologies and intentional leadership. 【Reading Tips】 - **Skim the diagnosis, deep-read the practices**: the early Dana narrative and Lean framing are motivating but familiar; the engineering chapters (dependency management, testing, refactoring) are where the concrete, reusable techniques live. - **Treat it as a reference, not a novel**: jump to the chapter matching your current pain — dependency hell, flaky manual testing, or product uncertainty — rather than reading linearly. - **Do the hands-on examples**: the book explicitly includes code-along material for environments and refactoring a problematic notebook; the value is in applying it to your own repo. - **Pair it with a systems-design text**: the authors deliberately defer ML systems design concepts to Chip Huyen's book, so keep that nearby for architecture questions. - **Bring the organizational chapters to your manager**: the Team Topologies and leadership material is aimed at structural change, not individual productivity hacks. 【Coverage Limits】 These excerpts are heavily front-loaded with table-of-contents fragments, preface material, and product-discovery chapters; the engineering, MLOps, and organizational chapters are represented mainly by headings and summaries, so this guide cannot detail their specific techniques, examples, or figures.
Excerpt 1
269 MLOps 101 270 Smells: Hints That We Missed Something 276 Continuous Delivery for Machine Learning 280 Benefits of CD4ML 280 A Crash Course on Continuous...
View in text
Excerpt 2
rnal documentation on how to resolve such customer queries. • Dana sends a reminder on the team chat to ask for a volunteer to review a pull request she crea...
View in text
Excerpt 3
scope. While they recognize that Responsible AI is crucial for addressing AI risks, such as safety, bias, fairness, and privacy issues, they admit to neglect...
View in text
Excerpt 4
it of value that can be independently shipped to production. A user story is the quantum, or building block, of the functionality of the product. Under agile...
View in text
Excerpt 5
ntime to automatically select the image that matches the OS and architecture of the respective development and deployment environments. For that, teams can u...
View in text
Excerpt 6
and merged. Huzzah! A big leap toward more secure software. Secure Dependency Management | 129 that we used in the previous chapter and you can find the same...
View in text
Excerpt 7
est_convert_keys_to_snake_case_replaces_title_cased_keys(): result = convert_keys_to_snake_case({"Job Description": "Wizard", "Work_Address": "Hogwarts Castl...
View in text
Excerpt 8
not working with tabular data (e.g., images, text, audio), as long as you can associate the segment with the data (e.g., images with correspond‐ ing metadata...
View in text
Tags
AI categories
Artificial IntelligenceMLOpsSoftware
ISBN: 1098144635
Publish Year: 2024
Language: English
Pages: 402
File Format: PDF
File Size: 15.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…