
Derive the Bias–Variance–Noise Decomposition for Squared Error
Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…
Read tutorialProbabilistic and statistical reasoning used to formulate, fit, compare, or interpret classical machine-learning methods.
Tagged articles
34 articles in this tag.

Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…
Read tutorial
You have a thousand bootstrap replicates sitting in a NumPy array. You have a histogram. What you do not have is the two numbers you actually need to…
Read tutorial
A random forest hands you an out-of-bag score, and someone nearby says, "Each tree leaves out about a third of the data." That number feels like a setting.…
Read tutorial
Your model scores 0.82 accuracy on the test set. That feels like a fact—something solid you can report and defend. Run the experiment again on a different…
Read tutorial
You have one sample, one number, and no honest way to say how much that number would move if you collected the data again. The bootstrap does not create…
Read tutorial
A model returns 0.4. The default cutoff turns that into "negative," and nobody asked whether that was the right call.
Read tutorial
Two models. Mean cross-validation scores of 0.842 and 0.847. The instinct is immediate: the second one wins.
Read tutorial
A satisfaction rating of 4 when the truth was 5 is not the same kind of mistake as predicting 1. Most models cannot tell the difference. This one can — but…
Read tutorial
You can fit a logistic regression in one line and read its coefficients in the next. What almost nobody shows you is where the loss function comes from.…
Read tutorial
A model degrades in production. Three engineers offer three diagnoses: "the inputs drifted," "the class balance changed," "the meaning changed." All three…
Read tutorial
You picked 2^18 buckets. You trained the model. The metrics came back slightly worse than your explicit-vocabulary baseline, and now you are staring at the…
Read tutorial
You can write the Gaussian mixture density in a single line. The trouble starts the moment you take the log.
Read tutorial
Picture a point sitting exactly between two well-separated blobs of data. K-means will force it into one group or the other, even when the data could…
Read tutorial
Linear regression draws a straight line through continuous numbers. Logistic regression draws an S-curve that outputs probabilities. They look like…
Read tutorial
Naive Bayes and logistic regression can classify the same dataset with nearly the same accuracy, yet they behave very differently on small samples, missing…
Read tutorial
You train a model, split your data randomly, and the test score looks great. Then the model meets real-world data and quietly falls apart. The usual…
Read tutorial
Fit KNN regression with k=1 and the prediction curve threads through every training point. Set k=50 and the same curve flattens toward a horizontal line.…
Read tutorial
Your model passed validation. Then it went to production and started predicting one class far more often than it should. The features look normal. The…
Read tutorial
You fit LinearDiscriminantAnalysis, plot the boundary, and get a straight line. Then someone asks why it is straight, and the honest answer is "because the…
Read tutorial
A logistic regression prediction is not a label. It is a probability—and the label is a decision you make on top of that probability.
Read tutorial
You do not need to finish a mathematics degree before you train your first model. You need enough math to understand what the model is doing—and you can…
Read tutorial
Two analysts open the same CSV. Both see the same 12% of rows with a blank in the income column. One drops those rows and reports a mean of 54,000. The…
Read tutorial
You run a grid search. The leaderboard prints a winner at 0.91. You ship it, and the first honest measurement lands near 0.86. Nothing broke. The model did…
Read tutorial
A multiclass model hands you K raw scores. The obvious move — divide each by their sum — collapses the first time a score goes negative. Here is the fix,…
Read tutorial
The intuition article told you Naive Bayes "combines prior and evidence." Then you tried to write $P(x_1, x_2, \dots, x_n \mid y)$ for real data, and the…
Read tutorial
Most classifiers earn their keep by capturing relationships between features. Naive Bayes does the opposite: it assumes those relationships don't exist.…
Read tutorial
A coefficient can be unbiased on average and still be wrong in your one dataset. Worse, it can be unbiased and still predict badly, and unbiased and still…
Read tutorial
A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.
Read tutorial
A count target is not a permission slip. It is a hypothesis about how the mean and variance move together — and you can test it.
Read tutorial
A regression model hands you one number. The person acting on that number needs to know how much weight it can bear. A prediction interval is how you show…
Read tutorial
Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.
Read tutorial
Everyone repeats that a forest reduces variance. Almost nobody writes down the formula that governs the reduction — and that formula contains a term that…
Read tutorial
Here is the uncomfortable truth about missing data: in a real dataset, you never see the values that went missing. You cannot check whether your imputation…
Read tutorial
A prediction interval built from a model's own residuals has no guarantee at all. Split conformal prediction replaces that hope with a finite-sample…
Read tutorial