Machine Learning for Business Using Amazon SageMaker and Jupyter (Doug Hudgeon, Richard Nichol)(Z-Library)
Artificial Intelligence
No description
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
# Machine Learning for Business Using Amazon SageMaker and Jupyter
## 【One-Line Pitch】
A practical, hands-on guide for business professionals and aspiring practitioners who want to apply machine learning to real business problems using Amazon SageMaker and Jupyter notebooks, with no prior ML experience required.
## 【Book Arc】
- **Opening (~0%–10%)**: Introduces the business case for machine learning, contrasting pattern-based decision making with traditional rules-based systems, and explains why ML is becoming essential for companies to stay competitive.
- **Early (~10%–23%)**: Covers the fundamentals of how machines learn, including supervised vs. unsupervised learning, the concept of training a mathematical function to separate target variables, and introduces XGBoost as the primary algorithm used throughout the book.
- **Early (~23%–32%)**: Walks through the first complete business scenario—predicting whether purchase orders need technical approval—including setting up AWS, S3 storage, and SageMaker notebook instances.
- **Middle (~32%–42%)**: Details the full machine learning workflow: loading and examining data, transforming it into the right shape, splitting into training/validation/test datasets, training the model, hosting it on an endpoint, and making predictions.
- **Middle (~42%–48%)**: Continues with practical implementation details, including deploying models to servers, testing predictions, and cleaning up resources to avoid unnecessary AWS costs.
- **Late (~48%+)**: Expands into additional business scenarios, including customer retention prediction and time series forecasting for power consumption, demonstrating how to apply the same principles to different problems.
## 【Key Takeaways】
- **Pattern-based decision making is the core of ML** (Early): Unlike traditional software that follows explicit rules, machine learning identifies patterns from historical data to make decisions, making systems more resilient to novel inputs and edge cases.
- **XGBoost is the go-to algorithm for beginners** (Early): It's forgiving, works well across diverse problems without heavy tuning, requires relatively little data, produces explainable predictions, and performs strongly in competitions—making it ideal for learning.
- **The ML workflow follows a consistent six-part pattern** (Middle): Load and examine data, reshape it appropriately, create train/validate/test splits, train the model, host it on an endpoint, then test and use it for predictions—this structure repeats across all business scenarios.
- **Data splitting is critical for honest evaluation** (Middle): Using 70% of data for training, 20% for validation, and 10% for testing ensures the algorithm learns properly while you can objectively assess how well it generalizes to unseen data.
- **Hosting models requires different resources than training** (Middle): Training uses more powerful servers (like m5.large), while prediction endpoints can run on smaller instances (like t2.medium) since inference is less computationally intensive.
- **Data transformation matters for model performance** (Middle): Normalizing data (e.g., converting dollar amounts to percentages relative to averages) and calculating changes between periods helps algorithms detect meaningful patterns that raw numbers might obscure.
- **Cost management is built into the workflow** (Middle): Every chapter includes explicit steps for deleting endpoints and shutting down notebook instances, emphasizing that cloud resources must be cleaned up to avoid ongoing charges.
## 【Reading Tips】
- **Skim Part 1 for concepts, deep-read Part 2 for practice**: The early chapters explain ML theory and business context, but the real value comes from working through the Jupyter notebooks in the scenario chapters.
- **Follow the appendix setup before starting**: Appendixes A–C walk through AWS account creation, S3 bucket setup, and SageMaker configuration—complete these before attempting any chapter exercises.
- **Run the code yourself rather than just reading**: The authors emphasize that understanding comes from executing the notebooks and observing what happens at each step, not from passive reading.
- **Pay attention to the train/validate/test split logic**: This is a concept that trips up many beginners—understanding why three separate datasets are needed will help you evaluate models properly in your own projects.
- **Note the cost-control practices early**: Deleting endpoints and shutting down instances after each exercise is essential for avoiding surprise AWS bills, so internalize this habit from the first chapter.
## 【Coverage Limits】
The excerpts cover the book's introduction, ML fundamentals, and the first two business scenarios (technical approval prediction and customer retention). Later chapters on time series forecasting with DeepAR and serverless API deployment are mentioned but not detailed in the available material.
##
Page 15
your business processes faster and more resilient to change. This book is for people begin- ning their journey in machine learning or for those who are more...
View in text
Excerpt 2
development these days, which is rules-based decision mak- ing--where programmers write code that employs a series of rules to perform a task. When your mark...
View in text
Excerpt 3
gh this chapter and the book, visit appendixes A, B, and C, and follow the instructions you find there. After working your way through the appendixes (to the...
View in text
Excerpt 4
diction in listing 2.14 takes every column in the test data (except for the first column, because that is the value you are trying to predict), sends it to t...
View in text
Excerpt 5
ne of the most confusingly named terms in machine learning. But, because it is also one of the most helpful tools in understanding the performance of a model...
View in text
Excerpt 6
e label is then followed by the tokenized text of the tweet. Tokenizing is the process of taking text and breaking it into parts that are linguistically mean...
View in text
Excerpt 7
oth sections contain dark dots, level 2 of the tree diagram is shown as dark. Next, the section containing the target data point is split again. Figure 5.13...
View in text
Excerpt 8
dentify half the erroneous lines submitted by the law firms (201 lines). When you set the cutoff at this point, you misidentify 67 correct invoice lines as b...
View in text
Tags
AI categories
Cloud NativeArtificial IntelligenceProgramming
Text Preview (First 20 pages)
Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Generating text preview…
Loading comments...
Reply to Comment
Edit Comment