Summary
Machine Learning in Action is unique book that blends the foundational theories of machine learning with the practical realities of building tools for everyday data analysis. You'll use the flexible Python programming language to build programs that implement algorithms for data classification, forecasting, recommendations, and higher-level features like summarization and simplification.
About the Book
A machine is said to learn when its performance improves with experience. Learning requires algorithms and programs that capture data and ferret out the interesting or useful patterns. Once the specialized domain of analysts and mathematicians, machine learning is becoming a skill needed by many.
Machine Learning in Action
is a clearly written tutorial for developers. It avoids academic language and takes you straight to the techniques you'll use in your day-to-day work. Many (Python) examples present the core algorithms of statistical data processing, data analysis, and data visualization in code you can reuse. You'll understand the concepts and how they fit in with tactical tasks like classification, forecasting, recommendations, and higher-level features like summarization and simplification.
Readers need no prior experience with machine learning or statistical processing. Familiarity with Python is helpful.
What's Inside
A no-nonsense introduction
Examples showing common ML tasks
Everyday data analysis
Implementing classic algorithms like Apriori and Adaboos
===================================
Table of Contents
PART 1 CLASSIFICATION
Machine learning basics
Classifying with k-Nearest Neighbors
Splitting datasets one feature at a time: decision trees
Classifying with probability theory: naïve Bayes
Logistic regression
Support vector machines
Improving classification with the AdaBoost meta algorithm
PART 2 FORECASTING NUMERIC VALUES WITH REGRESSION
Predicting numeric values: regression
Tree-based regression
PART 3 UNSUPERVISED LEARNING
Grouping unlabele
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A hands-on tutorial that teaches classic machine learning algorithms by having you build them from scratch in Python, rather than calling a library. Best for developers and self-taught programmers who know a little Python and want to understand how classification, regression, and clustering actually work under the hood.
【Book Arc】
- **Opening (~0%–15%)**: Frames what machine learning is, why deterministic modeling fails for messy human problems, and walks the standard workflow from data collection through training, testing, and deployment. Establishes Python (with NumPy and Matplotlib) as the implementation language and explains why.
- **Early (~15%–35%)**: Begins Part 1 (Classification) with the k-Nearest Neighbors algorithm, using a movie "kicks vs. kisses" toy example and a dating-profile dataset to show data parsing, normalization, and a working classifier.
- **Middle (~35%–55%)**: Continues classification with decision trees — information theory, Shannon entropy, recursive tree construction, and Matplotlib-based tree plotting — then applies it to predicting contact lens prescriptions.
- **Late (~55%–80%)**: Covers the remaining classifiers (naïve Bayes, logistic regression, support vector machines) and the AdaBoost meta-algorithm, then moves into Part 2 on regression for predicting numeric values, including tree-based regression.
- **Ending (~80%–100%)**: Part 3 turns to unsupervised learning, starting with grouping unlabeled data (clustering); the excerpts do not cover the closing chapters in detail.
【Key Takeaways】
- **Machine learning is defined by improvement through experience** (Opening): the book opens by grounding the field in the idea that a program learns when its performance improves with data, and that many real problems resist deterministic modeling. This framing matters because it justifies why statistical tools are needed at all.
- **Python is chosen for readability, not performance** (Early): clear syntax, built-in high-level data types, and NumPy/SciPy (compiled in C and Fortran) let you write code that reads like linear algebra while still running fast. Matplotlib handles visualization.
- **kNN is the gentlest entry point to classification** (Early): the movie and dating examples show that a usable classifier needs only distance computation, a k value, and normalized features — normalization is emphasized because raw feature scales distort distance.
- **Decision trees trade interpretability for overfitting risk** (Middle): the book explicitly lists pros (cheap, human-readable, tolerant of missing/irrelevant features) and cons (prone to overfitting), then builds the tree recursively using Shannon entropy to pick the best split.
- **Entropy drives the split decision** (Middle): rather than guessing which feature to branch on, the algorithm tries each feature and measures which split best organizes the data — a concrete, testable criterion you can verify with the book's own entropy function.
- **The workflow is iterative, not linear** (Early): the book stresses that data collection or preparation problems often send you back to step 1, and new data can force revisiting earlier steps — a realistic counterweight to textbook pipelines.
- **Algorithms are taught as reusable code, not theory** (Throughout): each chapter pairs a classic algorithm (kNN, decision trees, naïve Bayes, logistic regression, SVM, AdaBoost, regression, Apriori) with runnable Python, so the takeaway is working implementations you can adapt.
- **Unsupervised learning closes the arc** (Ending): Part 3 shifts from labeled classification and numeric forecasting to grouping unlabeled data, completing the progression from supervised to unsupervised techniques.
【Reading Tips】
- **Deep-read Part 1's first two chapters** (kNN and decision trees): they establish the data-parsing, normalization, and evaluation habits reused throughout the book; skimming them will make later chapters harder than they need to be.
- **Type the code rather than reading it**: the book is explicitly a tutorial, and the value comes from running the examples, reloading modules, and inspecting arrays in the Python shell as the excerpts demonstrate.
- **Treat the pros/cons boxes as your map**: each algorithm chapter states when it works and its weaknesses; use these to decide which technique fits a problem before diving into implementation.
- **Skim the Matplotlib plotting sections if visualization isn't your goal**: tree plotting and annotation code is useful but can be treated as a utility you return to later.
- **Don't skip the workflow chapter**: the iterative collect→prepare→analyze→train→test→use loop is the most transferable idea in the book, more than any single algorithm.
【Coverage Limits】
This guide is based on stratified excerpts covering roughly the first half of the book (through decision trees and into later classification/regression material); the excerpts do not cover the full details of naïve Bayes, SVM, AdaBoost, regression, or the unsupervised chapters, so those sections are described only at the level of the table of contents.
Excerpt 1
algorithm PART 2 FORECASTING NUMERIC VALUES WITH REGRESSION Predicting numeric values: regression Tree-based regression PART 3 UNSUPERVISED LEARNING Grouping...
thon is the best choice for these reasons. 1.6 Why Python? Python is a great language for machine learning for a large number of reasons. First, Python has c...
eload kNN and then type kNN.datingClassTest() at the Python prompt. You should get results similar to the following example: >>> kNN.datingClassTest() the cl...
key in secondDict.keys(): if type(secondDict[key]).__name__=='dict':Download from Wow! eBook <www.wowebook.com> 58 CHAPTER 3 Splitting datasets one feature a...
The answer is no. Figure 4.4 plots two functions, f(x) and ln(f(x)). If you examine both of these plots, you’ll see that they increase andDownload from Wow!...
and one feature was useless? Do you throw out all the data? What about the 19 other features; do they have anything useful to tell you? Yes, they do. Sometim...
ee this in action, type the following in your Python shell: >>> dataArr,labelArr = svmMLiA.loadDataSet('testSet.txt') >>> b,alphas = svmMLiA.smoP(dataArr, la...
ked at this dataset with logistic regression. At that time, the average error rate was 0.35. With AdaBoost we never have an error rate that high, and with on...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Machine Learning in Action (Peter Harrington)(Z-Library) (1)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Machine Learning in Action (Peter Harrington)(Z-Library) (1)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment