
Derive the Bias–Variance–Noise Decomposition for Squared Error
Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…
Read tutorialSupervised prediction of numeric outcomes, including regression objectives, models, residual diagnostics, and uncertainty around numeric predictions.
Tagged articles
34 articles in this tag.

Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…
Read tutorial
You fit ridge, you fit lasso, lasso wins by a hair, and you quietly file away "lasso is better." That conclusion is a coin flip wearing a lab coat. One…
Read tutorial
You fit a linear model. The coefficients look sensible. Each feature seems to contribute its own share, and the numbers align with your intuition. Then you…
Read tutorial
Linear regression draws a straight line through continuous numbers. Logistic regression draws an S-curve that outputs probabilities. They look like…
Read tutorial
Fit KNN regression with k=1 and the prediction curve threads through every training point. Set k=50 and the same curve flattens toward a horizontal line.…
Read tutorial
Two models trained on the same data can look at a new point and give you sharply different answers. Neither one is broken. Each is answering a different…
Read tutorial
Fit a lasso on a dataset with a dozen features and something strange happens. A few coefficients come back as exactly 0.0. Not 0.0031. Not 0.0007. Zero.…
Read tutorial
You fit a line, plot the residuals, and they look like a clean cloud. No funnel, no curve, no obvious outlier. It is tempting to read that plot as a…
Read tutorial
You have probably heard the advice: "Just use a random forest. Trees handle everything." It sounds practical. It is also how many beginners end up with a…
Read tutorial
Most introductions to linear regression make it sound like drawing a line through some dots. That description is true and almost useless. Any line can be…
Read tutorial
Call LinearRegression().fit(X, y) and coefficients appear in milliseconds. The library never shows you the equation it solved, why that equation has a…
Read tutorial
A model that fits is not yet a model you can trust. Fitting takes one line. Trust takes a split, a reading of the numbers, and a look at where the errors…
Read tutorial
You train your first classifier. It hits 85% accuracy on the test set. You feel great—until you realize that a rule which ignores your data entirely would…
Read tutorial
A model that scores 90% accurate can still fail at the exact job it was built for. The headline number looks impressive, but it may be hiding a model that…
Read tutorial
You open a dataset, and the last few columns all look like answers. One column says the fruit is an apple. The next says it is red. A third says it is…
Read tutorial
A coefficient can be unbiased on average and still be wrong in your one dataset. Worse, it can be unbiased and still predict badly, and unbiased and still…
Read tutorial
You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the…
Read tutorial
A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.
Read tutorial
A count target is not a permission slip. It is a hypothesis about how the mean and variance move together — and you can test it.
Read tutorial
Fit degree 1 through 9, watch training error fall to almost zero, and pick degree 9. That is the trap. Training error measures how well a model remembers…
Read tutorial
You fit a straight line to data that clearly curves, and the line misses everything. Not by a little—systematically. It sits above the points in one…
Read tutorial
A regression model hands you one number. The person acting on that number needs to know how much weight it can bear. A prediction interval is how you show…
Read tutorial
Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.
Read tutorial
A respectable R² can sit on top of a model that is wrong in a structured way. The score summarizes average fit; it says nothing about where the fit fails.…
Read tutorial
A classification tree asks which split makes the labels purest. A regression tree has no labels to purify — only numbers that scatter. So what is it…
Read tutorial
You have data. You have a prediction goal. And you're stuck on a question that feels like it should be simple: should I use regression or classification?
Read tutorial
Regularization is not a magic switch that makes every model better. It is a deliberate trade: you give up a little accuracy on your training data to gain…
Read tutorial
A regression model reports a solid R-squared, so you call it done. Then you plot the predictions against the actual values and notice something…
Read tutorial
Two features that measure almost the same thing can make ordinary least squares produce coefficients that swing wildly while the predictions barely move.…
Read tutorial
One row can steer a line. Add a single extreme observation to an otherwise clean dataset, refit, and watch the slope tilt — even though every other point…
Read tutorial
A prediction interval built from a model's own residuals has no guarantee at all. Split conformal prediction replaces that hope with a finite-sample…
Read tutorial
One point sits far from the rest of the scatter, and the fitted line visibly tilts toward it. Move that point twice as far away. Squared error answers with…
Read tutorial
You have heard both terms. You can even recite rough definitions. Then you open your own dataset, stare at the columns, and freeze. Is this supervised or…
Read tutorial
Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset…
Read tutorial