Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Jim Dowling

Get up to speed on a new unified approach to building machine learning (ML) systems with a feature store. Using this practical book, data scientists and ML engineers will learn in detail how to develop and operate batch, real-time, and agentic ML systems. Author Jim Dowling introduces fundamental principles and practices for developing, testing, and operating ML and AI systems at scale. You'll see how any AI system can be decomposed into independent feature, training, and inference pipelines connected by a shared data layer. Through example ML systems, you'll tackle the hardest part of ML systems–the data, learning how to transform data into features and embeddings, and how to design a data model for AI. - Develop batch ML systems at any scale - Develop real-time ML systems by shifting left or shifting right feature computation - Develop agentic ML systems that use LLMs, tools, and retrieval-augmented generation - Understand and apply MLOps principles when developing and operating ML systems

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Building Machine Learning Systems with a Feature Store ## 【One-Line Pitch】 A practical guide for data scientists and ML engineers who want to move beyond model training to build complete, production-ready ML systems—batch, real-time, and LLM-powered—using the Feature, Training, and Inference (FTI) architecture unified by a feature store. ## 【Book Arc】 - **Opening (~0%–9%)**: Introduces the FTI architecture—feature, training, and inference pipelines connected by a shared data layer (feature store and model registry)—and positions the book as focused on ML systems rather than traditional MLOps concerns like Docker or experiment tracking. - **Early (~9%–28%)**: Covers ML fundamentals (supervised, unsupervised, self-supervised, reinforcement, and in-context learning) and the anatomy of ML systems, including how RAG works for LLM applications and why modular feature engineering matters for testability and maintainability. - **Early (~28%–38%)**: Dives into feature pipelines—programs that orchestrate data transformations—with concrete examples including an air quality forecasting service using Open-Meteo APIs, plus practical scheduling options like GitHub Actions and Modal. - **Middle (~38%–47%)**: Explores feature stores in depth: online vs. offline stores, feature freshness, TTL policies, and data modeling approaches (star schema and snowflake schema) for systems like credit card fraud detection. - **Middle (~47%–end)**: Covers batch and streaming feature pipelines, data sources (batch, streaming, object stores, APIs), backfilling and incremental updates, job orchestration (Airflow, Hopsworks Jobs), and advanced topics like model-dependent transformations, encoding categorical variables, and real-time ML systems. ## 【Key Takeaways】 - **FTI architecture is the unifying pattern** (Early): Any AI system decomposes into independent feature, training, and inference pipelines connected by a shared data layer—the feature store and model registry. This decomposition maps cleanly onto team responsibilities and enables scalable, maintainable systems. - **Feature stores solve the data problem in ML** (Middle): A feature store acts as the shared data layer connecting pipelines, with online stores for low-latency inference and offline stores for training. Feature freshness—the time from event ingestion to feature availability—is the critical metric for real-time systems. - **Modular feature engineering is non-negotiable** (Early): Monolithic functions that compute multiple features (like a single `compute_features` function) are hard to test, document, and debug. Store feature functions in Python modules with independent unit tests, not notebooks. - **Streaming pipelines enable real-time ML** (Middle): For systems like TikTok-style recommenders, features must be available within seconds of user actions. Streaming feature pipelines update the feature store via a stream API, making fresh features available for inference. - **RAG solved the context problem for LLMs** (Early): Stateless LLM applications needed relevant, timely context—both post-training-cutoff events and private data. RAG retrieves context at request time via vector databases and ANN search, with vector embedding pipelines keeping data current. - **Data modeling matters for feature groups** (Middle): Star schema (labels as facts, feature groups as dimensions) and snowflake schema are both viable for feature stores. Label feature groups are just normal feature groups—you identify features vs. labels only when selecting data for training. - **Scheduling pipelines is a solved problem with options** (Middle): GitHub Actions (free tier: 2,000 compute minutes/month), Modal, Airflow, Dagster, and Mage AI all work for orchestrating feature and inference pipelines. Choose based on whether you need notebook scheduling, credit card access, or managed platforms. ## 【Reading Tips】 - **Skim the ML fundamentals chapter** (~9%–16%) if you already know supervised vs. unsupervised learning; focus instead on the FTI architecture diagrams and the RAG explanation, which are the book's core contributions. - **Deep-read the feature engineering chapters** (~28%–38%): The refactoring example (monolithic `compute_features` → modular functions) is the single most practical lesson for writing maintainable ML code. The air quality example is a complete walkthrough you can replicate. - **Pay attention to the data modeling section** (~44%–47%): The star schema vs. snowflake schema discussion for feature groups is subtle but essential for designing your own feature store schemas. The credit card fraud example clarifies how labels fit into feature groups. - **The book is not a traditional MLOps text** (per the author's own admission): Skip it if you need Docker, Terraform, or experiment tracking. It assumes automatic containerization and focuses on data pipelines, not deployment infrastructure. - **Exercises at chapter ends** are worth doing if you're building skills; the book is structured so each chapter stands alone, so you can jump to the topic you need (batch, real-time, or LLM systems). ## 【Coverage Limits】 This guide covers the book's core architecture, feature engineering principles, and data modeling approaches as revealed in the opening ~47% of the book. The excerpts do not cover the later chapters on real-time ML systems (shifting left/right feature computation), agentic LLM systems with tools, or the TikTok recommender implementation in detail—these are mentioned but not fully explored in the available material. ##
Page 10
175 Transforming Numerical Variables 178 Storing Transformed Feature Data in a Feature ...
View in text
Excerpt 2
like a lawyer. The Anatomy of a Machine Learning System | 5 also used to create training data for training models. Feature pipelines can be batch programs th...
View in text
Excerpt 3
mbeddings that are stored in a vector index (in the feature store), and the feature data validation pipeline, which is an asynchronous program that runs data...
View in text
Excerpt 4
der system, where features are created in streaming feature pipelines using information about your viewing activity. Within a second of a user action, featur...
View in text
Excerpt 5
-disk columns. However, in-memory tables require enough RAM to store the data, and when you have feature groups that will store many TBs of online data, it m...
View in text
Excerpt 6
used to create those features. Figure 6-7 shows the feature pipeline that uses the tables (and event-streaming platform) in our data mart as the data sources...
View in text
Excerpt 7
eatures, we can look at how we write pipelines to run them. The following exercises will help you learn how to design your own MDTs and ODTs: • I have a feat...
View in text
Excerpt 8
SQL is that the online feature store must support a SQL API. For example, not all online feature stores support pushdown aggregations, as many o
View in text
Tags
AI categories
AIProgramming LanguageData
ISBN: 1098165233
Publish Year: 2025
Language: English
Pages: 509
File Format: PDF
File Size: 14.7 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…