Page
1
(This page has no text content)
Page
2
Copyright © 2024 © 2025 © 2026 Carl McBride Ellis Depósito Legal: M-000763/2025 email: carl.mcbride@protonmail.ch ORCiD: 0000-0003-2966-7530 LinkedIn: www.linkedin.com/in/carl-mcbride-ellis ISBN 978-1-80808-131-6 Book GitHub: github.com/Carl-McBride-Ellis/TOBoML Green edition, version 2.5.3 (3 June 2026) Disclaimers: No generative AI whatsoever has been used in the writing of this book; any hallucinations, errors, or conceptual mistakes are entirely of my own making. I would be delighted if you could bring such issues to my attention either using the dedicated Discussions section hosted on GitHub, or directly via email. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing or its dealers and distributors, will be held liable for any damages caused or alleged to have been caused directly or indirectly by this book.
Page
3
Contents 1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 1.1 The β̂ and the ŷ . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 1.2 Estimators and approximators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.3 Transductive vs inductive models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.4 Prediction vs forecast . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.5 Errors and residuals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.6 Sources of uncertainty: aleatoric and epistemic . . . . . . . . . . . . . . . . . . . . . . . . 3 1.7 Confidence and prediction intervals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 1.8 Explainability and interpretability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 1.9 Correlation and causation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2 Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.1 Centrality: mean, median, and mode . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.1.1 Law of large numbers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 2.2 Dispersion: range, variance, MAD, and quartiles . . . . . . . . . . . . . . . . . . . . . . . . 9 2.2.1 Quantiles, quartiles and the interquartile range (IQR) . . . . . . . . . . . . . . . . . . . . . . . . 10 2.3 Statistical (and mechanical) moments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.4 Gaussian distribution: additive . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.4.1 Statistical tests for (non)-normality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.5 Chebyshev’s inequality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.6 Galton distribution: multiplicative . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14
Page
4
2.7 Power-law and exponential distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.8 Bernoulli distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.9 The exponential family of probability distributions . . . . . . . . . . . . . . . . . . . . . . 17 2.10 Skewness and kurtosis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 3 Exploratory data analysis (EDA) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.1 Data quality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 3.1.1 Synthetic and syncretic data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 3.2 Getting to know your dataframe . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 3.2.1 The curse of dimensionality . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 3.2.2 Descriptive statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.2.3 Pivot and contingency tables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.2.4 pandas groupby . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 3.3 Outliers, inliers and extreme values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 3.4 Correlation coefficients . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 3.4.1 Uncertainty coefficient: Theil’s U (category-category) . . . . . . . . . . . . . . . . . . . . . . . 30 3.4.2 ΦK (universal) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 3.4.3 Mutual information (MI) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 3.5 Why visualize data? Anscombe’s quintet . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 3.6 Univariate visualisation: histograms and the eCDF . . . . . . . . . . . . . . . . . . . . . . 33 3.6.1 Box, violin, and raincloud plots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 3.6.2 Discrete (categorical) data: countplots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 3.7 Bivariate visualisation: scatter plot . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 3.7.1 Pairplots (or perhaps not) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 4 Data cleaning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 4.1 Missing values: NULL or NaN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 4.1.1 Visualization of NaN with missingno . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 4.1.2 Types of missing: MCAR, MAR, and MNAR . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 4.1.3 Global delete . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 4.1.4 Global fill . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 4.1.5 Univariate imputation with an average value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 4.1.6 Multiple imputation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 4.1.7 Do nothing! . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 4.1.8 Binary indicator column . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 4.2 Outlier and inlier removal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 4.2.1 Univariate outliers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 4.2.2 Cook’s distance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48 4.2.3 Isolation forest . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48 4.2.4 Removal using PCA . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49
Page
5
4.3 Duplicated rows . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.4 Duplicated columns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.5 Boolean columns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.6 Zero variance columns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.7 Feature scaling: standardization and normalization . . . . . . . . . . . . . . . . . . . . . 51 4.8 Encoding categorical features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 4.8.1 Ordinal features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 4.8.2 Nominal features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 4.8.3 Label encoding of classification targets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 4.9 Removing personally identifiable information . . . . . . . . . . . . . . . . . . . . . . . . . . 56 5 Cross-validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 5.1 Why out-of-sample evaluation? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 5.2 Three-way split: train, validation, and test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 5.3 Permutation invariance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 5.4 Cross-validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 5.4.1 E(CV): ...now change the random_state . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63 5.4.2 How many folds? Pessimistic bias . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 64 5.5 Nested cross-validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 5.6 Data leakage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 5.7 Covariate shift and Concept drift . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 5.7.1 Comparing univariate distributions: two-sample tests . . . . . . . . . . . . . . . . . . . . . . . . 68 5.7.2 Comparing multivariate distributions: C2ST . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72 6 Interpolation and smoothing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 6.1 Interpolation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 6.1.1 k-nearest neighbors (k-NN) regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 6.1.2 Splines: piece-wise polynomials . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 6.1.3 B-spline basis functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77 6.2 Smoothing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 6.3 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 7 Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 81 7.1 Regression baseline model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82 7.2 Linear regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 82 7.3 Calculating β1 and β0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 7.3.1 Ordinary least squares . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 7.3.2 Normal equation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 84 7.3.3 Scikit-learn LinearRegression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 7.3.4 statsmodels OLS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86
Page
6
7.4 Assumptions of linear regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 7.4.1 Heteroscedasticity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87 7.4.2 Simpson’s paradox . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88 7.5 Polynomial regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89 7.6 Extrapolation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 7.6.1 Convex hull . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91 7.7 Interpretability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91 7.8 The loss function and cost functional . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 7.9 Gradient descent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 7.9.1 Stochastic gradient descent regressor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 7.10 Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 7.10.1 Root mean squared error (RMSE) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 7.10.2 Mean absolute error (MAE) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 7.10.3 MAE ≤ RMSE . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 7.10.4 Coefficient of determination (R2) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 7.11 Linear regression learning curve . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108 7.12 Decision tree regressor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 7.12.1 Hyperparameter: max_depth . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 112 7.13 Overfitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115 7.13.1 Parametric models: regularization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 116 7.13.2 Tree models: min_samples_leaf . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 7.14 Bias-variance tradeoff and decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 7.15 Monotonic constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 7.16 Quantile regression and prediction intervals . . . . . . . . . . . . . . . . . . . . . . . . . . 122 7.16.1 Pinball loss function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123 7.17 Conformal prediction intervals . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 7.17.1 Conformalized quantile regression (CQR) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 126 7.17.2 Locally-weighted conformal regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 7.17.3 Prediction interval metric: Winkler interval score . . . . . . . . . . . . . . . . . . . . . . . . . . 129 7.18 Expectile regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130 7.19 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131 8 Classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133 8.1 Nearest neighbors classifier (k-NN) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134 8.1.1 Nearest centroid classifier . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 8.2 Bernoulli regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 8.2.1 Linear regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 8.2.2 Probit regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136 8.2.3 Logistic regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137
Page
7
8.3 Log loss function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 138 8.4 Calculating κ1 and κ0 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140 8.4.1 Scikit-learn LogisticRegression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140 8.4.2 statsmodels Logit . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140 8.4.3 Gradient descent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 141 8.5 Probabilistic classification baseline model . . . . . . . . . . . . . . . . . . . . . . . . . . . 142 8.6 Dichotomization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 8.6.1 Linear classification: the perceptron . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 8.6.2 Logistic classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 143 8.7 Logistic regression and the XOR problem . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 8.8 Decision tree classifier . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 146 8.9 Classifier output: predict vs predict_proba . . . . . . . . . . . . . . . . . . . . . . . . . . . 149 8.10 Probabilistic metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150 8.10.1 Strictly proper scoring rules . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150 8.10.2 Calibration-refinement decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 150 8.10.3 AUC ROC score . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151 8.11 Classifier calibration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152 8.11.1 Reliability diagrams . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 152 8.11.2 Venn-ABERS calibration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 8.12 Classification metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155 8.12.1 Decision threshold and dichotomization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155 8.12.2 Classification baseline model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 8.12.3 Confusion matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 8.12.4 Accuracy score . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 158 8.12.5 Precision and recall . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 8.12.6 F1 score . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 8.12.7 Negative rates . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 8.12.8 Matthews correlation coefficient (MCC) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160 8.13 Multiclass classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160 8.13.1 One-vs-rest . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160 8.13.2 Multinomial logistic regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160 8.13.3 Multiclass metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162 8.14 Imbalanced classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162 8.14.1 The things you see people do to imbalanced data . . . . . . . . . . . . . . . . . . . . . . . . 164 8.14.2 A question of i.i.d. datasets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 165 8.15 Overfitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 166 8.16 No free lunch theorem . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 9 GLM and GAM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169 9.1 Generalized linear models (GLM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 169 9.2 Kolmogorov-Arnold superposition theorem . . . . . . . . . . . . . . . . . . . . . . . . . . 171 9.3 Generalized additive models (GAM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 172
Page
8
10 Ensemble estimators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174 10.1 Random Forest . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174 10.1.1 Bootstrapping: row subsampling with replacement . . . . . . . . . . . . . . . . . . . . . . . . 174 10.1.2 Feature subsampling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 175 10.1.3 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177 10.2 Extremely Randomized Trees . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 178 10.3 Weak learners and boosting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179 10.3.1 AdaBoost (Adaptive Boosting) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179 10.4 Gradient boosted decision trees (GBDT) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 180 10.4.1 The ‘big four’ GBDT . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181 10.4.2 Feature quantization: histogram gradient boosting . . . . . . . . . . . . . . . . . . . . . . . . 182 10.4.3 Tree growth strategy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 182 10.4.4 Dropout: DART . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183 10.4.5 Categorical features: native support . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 183 10.4.6 Missing values: native support . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 184 10.5 Extrapolation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 184 10.5.1 Linear tree . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185 10.6 Bias-variance-diversity decomposition and tradeoff . . . . . . . . . . . . . . . . . . . 187 10.7 Convex combination of model predictions (CCMP) . . . . . . . . . . . . . . . . . . . . 187 10.8 Stacking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 189 11 Hyperparameter optimization (HPO) . . . . . . . . . . . . . . . . . . . . . . . . . . 192 11.1 The Optimizer’s Curse: optimistic bias . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 192 11.2 Linear and logistic regression: None . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193 11.3 Decision trees: max_depth and min_samples_leaf . . . . . . . . . . . . . . . . . . . . . . 194 11.4 Random Forest: n_estimators (?) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 195 11.5 Gradient boosted decision trees: learning_rate . . . . . . . . . . . . . . . . . . . . . . 197 11.5.1 Overfitting: early stopping . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197 11.5.2 Some other GBDT hyperparameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 199 11.6 Automated search routines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 200 11.6.1 Specialized packages . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202 11.7 ...now change the random_state . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 202 12 Feature engineering and selection . . . . . . . . . . . . . . . . . . . . . . . . . . . 203 12.1 Feature engineering . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 203 12.1.1 Interaction and cross features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204 12.1.2 Bucketing of continuous features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204 12.1.3 Power transforms: Yeo-Johnson . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204 12.1.4 User defined transform . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205 12.1.5 Temporal features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205 12.1.6 External secondary features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205
Page
9
12.2 Feature selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205 12.2.1 Correlation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206 12.2.2 Permutation importance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 206 12.2.3 SHAP exact explainer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 208 12.2.4 Stepwise regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209 12.2.5 LASSO . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210 12.2.6 Maximum relevance minimum redundancy (mRMR) . . . . . . . . . . . . . . . . . . . . . . . 210 12.2.7 Boruta trick . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 210 12.2.8 Categorical features: chi-squared (χ2) test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211 12.2.9 Native feature importance plots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 211 12.3 Principal component analysis (PCA) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213 13 Tabular foundation models (TFM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 216 13.1 Artificial neural networks: simple examples . . . . . . . . . . . . . . . . . . . . . . . . . . 216 13.2 Why no traditional ANN? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 219 13.3 Foundation models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220 13.3.1 Pretraining corpus . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 221 13.3.2 TFM architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 222 13.3.3 In-context learning (ICL) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223 13.3.4 Classification and regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 224 Further reading . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 228
Page
10
This book is based on my five day course which I had the pleasure of teaching in the following Spanish cities: A Coruña, Algeciras, Alicante, Bilbao, Cáceres, Granada, Huesca, Jaén, Madrid, Málaga, Murcia, Sevilla, Valencia, Valladolid, and Zaragoza. I would like to take this opportunity to say a big thank you to all of my students ¡un gran placer! This ‘Green’ edition builds upon the material originally imparted in the aforementioned course. This book uses the python programming language largely in conjunction with the scikit-learn machine learning library, and pandas for data manipulation. All example notebooks use the Jupyter Notebook environment and many of the sample code snippets are also available as GitHub gist files. Special thanks to: Carlos Ortega, Santiago Mota Herce, Norberto Malpica, Charles H. Martin, Roland Stevenson, Adrian Olszewski, Ambros Marzetta, Valeriy Manokhin, and everyone on Kaggle. I would also very much like to thank Dr. Valeriy Manokhin and Miguel Pérez Michaus for their critical reading of prior versions of this book, as well as feedback from both Dr. David Holzmüller and Dr. Noah Hollmann particularly regarding Chapter 13. For Cristina
Page
11
1. Introduction Statisticians torture their data to fit the model. Data scientists torture their models to fit the data. 1.1 The β̂ and the ŷ A statistical model, primarily designed for explanatory or inference purposes, will use a simple predefined function (based on the assumption that the data can be modelled by a parametric distribution) and will then estimate the best parameters, β̂ , for said function. A non-parametric, i.e. making no assumptions about the distribution of the data, machine learning model forgoes explanatory power for predictive power, with the emphasis on producing the best possible predictions, ŷ, for previously unseen data (at the end of the day we are not interested in predicting the data that we already have!). Equation: Fundamental assumption: our data is composed of signal + noise which is symmetrically distributed either side of the signal (i.e. E[ε] = 0). In one dimension we can write this as y = f (x)︸︷︷︸ signal + ε︸︷︷︸ noise (1.1) In multiple dimensions the signal f (X) due to the data generating process will form a hypersurface. Our objective is to produce the best possible fit to this hypersurface. Suffice to say that in machine learning we do not know the original f (X) and our challenge, somewhat akin to an inverse problem, is to find an approximation to f (X) having been given a (sparse) smattering of (noisy) points on y. This short book is dedicated to introducing us to the essentials of a number of techniques for producing such a fit: this book is all about the ŷ.
Page
12
1.2 Estimators and approximators 2 1.2 Estimators and approximators Generalized linear models such as linear regression and logistic regression are statistical estimators; their purpose is to estimate values of the parameters of a model, for example β̂1 and β̂0, that has been fitted to samples taken from a population that has a well defined distribution. On the other hand in approximation theory the target that one fits to is usually sampled via an evenly spread dense grid of points on the domain. Indeed both decision trees and artificial neural networks are universal approximators. However, in machine learning neither are we so much interested in the parameters of a model, but rather the predictions of said model, nor are we provided with a nice grid of points having noiseless targets; we work with the data that we are given. In some sense machine learning is situated somewhere between estimation and approximation. Perhaps for that reason the algorithms that we employ are sometimes known as learners, rather than as estimators or as approximators. 1.3 Transductive vs inductive models A transductive model is one in which the data is the model; when one makes a prediction one consults the training data itself. Examples are the k-nearest neighbors regressor and classifier. On the other hand an inductive model has generalization in mind; it will distill the training data into a intermediate compact form and it is this compact representation that is used for prediction. An example would be linear regression; in the univariate case the form is a straight line, and any number of training data points become represented by just two values, namely β1 and β0. When making a prediction the training data is no longer relevant and one instead consults the model. The majority of this book is concerned with creating such inductive models. Figure 1.1: Different types of inference (adapted from the book “The Nature of Statistical Learning Theory” by Vladimir N. Vapnik). 1.4 Prediction vs forecast The etymology of the verb to predict is to say before, but before exactly what is ill defined. To avoid getting embroiled in semantics I shall simply provide an example. Say we use gradient descent as our optimization algorithm (or solver) to learn from our set of empirical
Page
13
1.5 Errors and residuals 3 data points (our dataset) in order to estimate the parameters β̂1 and β̂0 for the model ŷ(x) = β̂1x+ β̂0 for the one-dimensional independent variable, x, whose data had a range [a,b]. If we now use this hypothesis or fit (model + parameters) to calculate the value ŷ(x = c) where a < c < b this is a (point) prediction. If c < a or c > b this is an extrapolation, and when the independent variable is time and c > b this is called a (point) forecast. 1.5 Errors and residuals The error is the difference between the observed value and the ideal value. For example, a chocolate bar may have the ideal value printed on the packet, say 100g. However, no chocolate bar has ever been made that weights exactly 100g, and the difference between the actual weight and ideal 100g is the error. On the other hand the residual is the difference between the predicted value and the observed value. Note however that in this text when we use the word error correctly speaking we should really use the word residual, and we do this simply because of the prevalence of the word error in the machine learning literature, for example the mean absolute error (MAE) or the root mean squared error (RMSE). 1.6 Sources of uncertainty: aleatoric and epistemic Aleatoric, or stochastic, uncertainty is due to inherent random noise in a system/generating process. Epis- temic, and Knightian, uncertainty is due to a lack of data or knowledge; for example having an incomplete dataset. Figure 1.2: Aleatoric and epistemic uncertainty in a dataset. Aleatoric uncertainty is considered to be irreducible, whereas epistemic uncertainty is reducible; one could fill gaps present in the data. Knightian uncertainty could take the form of invisible features that influence the target, but are either not provided, or whose collection was perhaps never even contemplated, or our data could have missing values. Knowing how the uncertainty is partitioned could indicate whether further data acquisition could lead to better predictions. However, in practice a neat decomposition into the aforementioned components is not empirically practicable.
Page
14
1.7 Confidence and prediction intervals 4 1.7 Confidence and prediction intervals Confidence intervals1 quantify the uncertainty in the estimated parameters β̂ . Prediction intervals quantify the uncertainty in the estimated predictions ŷ. We shall see how to calculate regression confidence intervals in Section 7.3.4, and prediction intervals using conformal prediction in Section 7.17. 1.8 Explainability and interpretability • Explainability: the ability to understand both the ‘how’ and the ‘why’ of a prediction • Interpretability: the ability of understand the ‘why’ of a prediction Explainability is the forte of statistical models; the ‘how’ is in-built from the very start by using simple models for the data. On the other hand the focus of machine learning (ML) is almost exclusively on prediction performance. Only the simplest ML models are interpretable; decisions trees are in principle very easy to explain and are often touted as being a ‘white box’ algorithm, but even at depth 3 (i.e. 8 leaves) and having several features then in reality their predictions quickly become very hard to interpret. There are packages such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-Agnostic Explanations) to facilitate post hoc interpretability. However, it has been seen that often even the data scientists themselves do not know how to use these packages properly2: “...results indicate that data scientists over-trust and misuse interpretability tools. Furthermore, few of our participants were able to accurately describe the visualizations output by these tools”. which does not bode well. To make matters worse of late some of the uses that people make of the SHAP technique, such as posing counterfactual questions, have been brought into question3, 4. At the end of the day machine learning models are designed for performance, and generally interpretability takes a back seat. However, in many circumstances model interpretability is of paramount importance; indeed under the EU General Data Protection Regulation (GDPR) 2016/679 (Article 15 1h) “The data subject shall have the right to obtain. . . the following information: the existence of automated decision-making. . . meaningful information about the logic involved, as well as the significance and the envisaged consequences of such processing for the data subject.” With the ever increasing rôle that machine learning models play in modern society the creation of inter- pretability tools is an active area of development. 1.9 Correlation and causation Caveat emptor: standard machine learning can only identify correlative patterns, but as the mantra goes correlation does not imply causation. A famous example is the purported correlation between ice cream sales and death by drowning. Indeed we could create a perfectly valid predictive model ŷ = β̂1x where x is the number of ice creams sold, and ŷ is our prediction for the number of drownings. However, suffice to say ice cream sales do not cause drownings. Here there is an invisible feature (known as a 1Morey et al. “The fallacy of placing confidence in confidence intervals”, Psychonomic Bulletin & Review 23 pp. 103-123 (2016) 2Kaur et al. “Interpreting Interpretability: Understanding Data Scientists’ Use of Interpretability Tools for Machine Learning”, Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems pp. 1-14 (2020) 3Huang, Marques-Silva “The Inadequacy of Shapley Values for Explainability”, arXiv:2302.08160 (2023) 4Bilodeau, Jaques, Wei Koh, Kim “Impossibility Theorems for Feature Attribution”, arXiv:2212.11870 (2024)
Page
15
1.9 Correlation and causation 5 confounding factor) involved, namely temperature; when there is hot weather more people buy ice creams, and more people go swimming and unfortunately drown. We can see that in reality our interpretable model is highly misleading, and does not constitute an explanatory model. Recommended reading Papers • Leo Breiman “Statistical Modeling: The Two Cultures”, Statistical Science 16 pp. 199-231 (2001) • Galit Shmueli “To Explain or to Predict?”, Statistical Science 25 pp. 289-310 (2010) • Hrushikesh N. Mhaskar, Efstratios Tsoukanis, Ameya D. Jagtap “An Approximation Theory Perspec- tive on Machine Learning”, arXiv:2506.02168 (2025) • Gruber et al. “Sources of Uncertainty in Machine Learning - A Statisticians’ View”, arXiv:2305.16703 (2023) Books Explainability, interpretability, and trustworthiness • Denis Rothman “Hands-On Explainable AI (XAI) with Python”, Packt Publishing Limited (2020) • Serg Masís “Interpretable Machine Learning with Python”, Packt Publishing Limited (2023) • Christoph Molnar “Interpretable Machine Learning”, (online book) • Kush R. Varshney “Trustworthy Machine Learning”, (2022) • Mucsányi, Kirchhof, Nguyen, Rubinstein, Oh “Trustworthy Machine Learning”, arXiv:2310.08215 (2023) Causal analysis • Matheus Facure “Causal Inference in Python”, O’Reilly Media, Inc. (2023) • Aleksander Molak “Causal Inference and Discovery in Python”, Packt Publishing Limited (2023) Packages • SHAP • SHAP-IQ - Shapley Interaction Quantification • LIME • ConformaSight Global Explainer Package (paper) • EconML • CausalML
Page
16
2. Statistics 60% of the time, it works every time Brian Fantana in ‘Anchorman’ In this chapter we briefly cover some essential statistical concepts that are useful when it comes to under- standing our data and the workings of machine learning algorithms. 2.1 Centrality: mean, median, and mode The mean and the median are measures of the central tendency, or location, of a distribution. Equation: The Kolmogorov generalised mean is given by M(x) := g−1 ( 1 n n ∑ i=1 g(xi) ) (2.1) If g(x) = 1(x) we have the arithmetic mean of an additive sample of size n, given by x̄ = 1 n n ∑ i=1 xi (2.2) The mean of the hypothetical population (i.e. the expectation value, E(x)), from which the sample was supposedly obtained, is denoted by µ .
Page
17
2.1 Centrality: mean, median, and mode 7 If g(x) = log(x) and where xi > 0 we have the geometric mean, given by GM(x) = ( n ∏ i=1 xi ) 1 n (2.3) and is used for values that are combined multiplicatively. If g(x) = (1/x) and where xi > 0 we have the harmonic mean, given by HM(x) = n ∑ n i=1 1 xi (2.4) The median (x̃) of a set of values is a number such that half of the values are below the median value, and the other half are above the median value. The median is said to be a robust statistic in that it is less influenced by a few extreme or outlier values. For example here we calculate the mean and the median for an array of values using the python numpy library import numpy as np from scipy import stats np.mean([1, 2, 3, 4, 5, 6, 1001]) # arithmetic mean 146.0 stats.gmean([1, 2, 3, 4, 5, 6, 1001]) # geometric mean 6.867897363489863 stats.hmean([1, 2, 3, 4, 5, 6, 1001]) # harmonic mean 2.855978316248548 np.median([1, 2, 3, 4, 5, 6, 1001]) 4.0 The mode is the most frequently occurring value, and can be useful when working with a discrete distribution (such as count data), categorical features, or for quickly identifying a zero-inflated distribution.
Page
18
2.1 Centrality: mean, median, and mode 8 Figure 2.1: An example of zero-inflated count data. An example of zero-inflated data would be the number of items purchased when visiting a web site; most visits result in no purchase. Remark: For our intents and purposes it is interesting to consider the mean, median, and mode not so much as simply being summary statistics, but rather to think of them as being elemental baseline models of a set of numbers having one degree of freedom. 2.1.1 Law of large numbers The law of large numbers states that, as the number of samples (n) increases, the mean of the sample converges on the mean of the population that the samples were drawn from x̄→ µ as n→ ∞. (2.5) This has very important implications for our machine learning model; our model, if it was trained on a sufficiently large sample, should then be transferable (i.e. make just as good predictions) for a different sample taken from the same population. Transferability is an absolutely vital quality of a machine learning model as at the end of the day its sole function is to provide trustworthy predictions for data (i.e. Xnew) that it has never seen before.
Page
19
2.2 Dispersion: range, variance, MAD, and quartiles 9 Figure 2.2: The results of 1,000 simulations (green lines) for the mean of flipping a fair coin for a total of 400 flips. The horizontal dashed black line represents the mean of the 1,000 trajectories at that step, and the upper and lower black curves represents the ±1 standard deviation of the trajectories at that step. The (orange) histogram on the right is the distribution of the final mean value for each trajectory. 2.2 Dispersion: range, variance, MAD, and quartiles In Figure 2.3 we have two distributions that have exactly the same mean value. Figure 2.3: Two distributions with same central tendency. We can see that centrality only tells part of the story and evidently we also need to be able to describe how spread out a distribution is as well. The range (R) of a distribution is simply the difference between the
Page
20
2.2 Dispersion: range, variance, MAD, and quartiles 10 largest and smallest values, i.e. R(x) = max(x)−min(x) (2.6) Equation: The variance, Var or σ2, is given by σ 2(x) := 1 n n ∑ i=1 (xi− x)2 (2.7) and √ Var, or σ , is known as the standard deviation. Remark: When calculating the variance or standard deviation by default, unlike numpy, pandas applies the Bessel correction of n−1. Equation: The sample covariance, between the ordered pair x and y, is given by cov(x,y) = 1 (n−1) n ∑ i=1 (xi− x)(yi− y) (2.8) We have Var(x+ y) =Var(x)+Var(y)+2cov(x,y) (2.9) In 2-dimensions we can now also construct a variance-covariance matrix: K = [ Var(x) cov(x,y) cov(y,x) Var(y) ] (2.10) (note that cov(x,y) = cov(y,x) so this matrix is symmetric). import numpy as np K = np.cov(df, rowvar=False) The median absolute deviation is the robust statistic for dispersion. Equation: The median absolute deviation (MAD) is given by MAD(x) := median(|xi− x̃|) (2.11) 2.2.1 Quantiles, quartiles and the interquartile range (IQR) The median is a special case of a quantile and is often denoted as Q2. Two other quantiles of note are Q1, below which lies 25% of the values, and Q3, above which also lies 25% of the data. Suffice to say the other 50% of the data has values between Q1 and Q3 and is known as the interquartile range (IQR).