
Derive AdaBoost’s Weight Updates From Exponential Loss
AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorialObjective functions that quantify prediction penalties and shape fitted parameters, model behavior, or the tradeoff among error types.
Tagged articles
18 articles in this tag.

AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorial
A model returns 0.4. The default cutoff turns that into "negative," and nobody asked whether that was the right call.
Read tutorial
You can fit a logistic regression in one line and read its coefficients in the next. What almost nobody shows you is where the loss function comes from.…
Read tutorial
Fit the next tree to the error. That instruction works beautifully for squared loss and then quietly falls apart the moment you switch to log loss. The…
Read tutorial
Fit K-means twice on the same data with the same k, and you can get two different clusterings. The second run often reports a lower cost than the first.…
Read tutorial
Most introductions to linear regression make it sound like drawing a line through some dots. That description is true and almost useless. Any line can be…
Read tutorial
Call LinearRegression().fit(X, y) and coefficients appear in milliseconds. The library never shows you the equation it solved, why that equation has a…
Read tutorial
A logistic regression prediction is not a label. It is a probability—and the label is a decision you make on top of that probability.
Read tutorial
Most people assume a machine learning model learns to be right. It doesn't. A model learns to minimize a number you hand it—and that number is the loss…
Read tutorial
A multiclass model hands you K raw scores. The obvious move — divide each by their sum — collapses the first time a score goes negative. Here is the fix,…
Read tutorial
You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the…
Read tutorial
A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.
Read tutorial
Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.
Read tutorial
A classification tree asks which split makes the labels purest. A regression tree has no labels to purify — only numbers that scatter. So what is it…
Read tutorial
You have data. You have a prediction goal. And you're stuck on a question that feels like it should be simple: should I use regression or classification?
Read tutorial
Two features that measure almost the same thing can make ordinary least squares produce coefficients that swing wildly while the predictions barely move.…
Read tutorial
One row can steer a line. Add a single extreme observation to an otherwise clean dataset, refit, and watch the slope tilt — even though every other point…
Read tutorial
One point sits far from the rest of the scatter, and the fitted line visibly tilts toward it. Move that point twice as far away. Squared error answers with…
Read tutorial