
Derive AdaBoost’s Weight Updates From Exponential Loss
AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorialObjectives and procedures for fitting model parameters, including losses, gradients, iterative updates, and convergence behavior.
Tagged articles
27 articles in this tag.

AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorial
K-means does not discover the real categories hiding in your data. It draws geometric boundaries around points that happen to sit close together. Those…
Read tutorial
A fully grown decision tree can score near-perfect on the data it was trained on and still lose to a single split on data it has never seen. The usual…
Read tutorial
You can write theta = theta - lr * grad from memory and still freeze when someone asks where grad came from. That gap is the whole problem. The update rule…
Read tutorial
You can fit a logistic regression in one line and read its coefficients in the next. What almost nobody shows you is where the loss function comes from.…
Read tutorial
You can write the Gaussian mixture density in a single line. The trouble starts the moment you take the log.
Read tutorial
Random forests grow many trees independently and average their votes. Gradient boosting grows trees one at a time, each one aimed at the mistakes the…
Read tutorial
Fit the next tree to the error. That instruction works beautifully for squared loss and then quietly falls apart the moment you switch to log loss. The…
Read tutorial
A model starts with random parameters and ends up making sharp predictions. Nobody tells it the answer. It just keeps nudging itself in the right direction…
Read tutorial
"Gradient descent works" and "gradient descent is guaranteed to converge" sound like the same sentence. They are not. One is a hope backed by experience.…
Read tutorial
You set a learning rate, hit run, and the loss curve does something strange. Maybe it barely moves. Maybe it bounces. Maybe it turns into nan. The number…
Read tutorial
Fit K-means twice on the same data with the same k, and you can get two different clusterings. The second run often reports a lower cost than the first.…
Read tutorial
Fit a lasso on a dataset with a dozen features and something strange happens. A few coefficients come back as exactly 0.0. Not 0.0031. Not 0.0007. Zero.…
Read tutorial
You fit a line, plot the residuals, and they look like a clean cloud. No funnel, no curve, no obvious outlier. It is tempting to read that plot as a…
Read tutorial
Most introductions to linear regression make it sound like drawing a line through some dots. That description is true and almost useless. Any line can be…
Read tutorial
Call LinearRegression().fit(X, y) and coefficients appear in milliseconds. The library never shows you the equation it solved, why that equation has a…
Read tutorial
You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the…
Read tutorial
You can call PCA(), read explained_variance_ratio_, and get a working result. Then someone asks why the principal directions are eigenvectors of the…
Read tutorial
A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.
Read tutorial
Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.
Read tutorial
A classification tree asks which split makes the labels purest. A regression tree has no labels to purify — only numbers that scatter. So what is it…
Read tutorial
Two features that measure almost the same thing can make ordinary least squares produce coefficients that swing wildly while the predictions barely move.…
Read tutorial
One point sits far from the rest of the scatter, and the fitted line visibly tilts toward it. Move that point twice as far away. Squared error answers with…
Read tutorial
Most explanations of support vector machines stop at the picture: two clouds of points, a line between them, a gap on either side. That picture is correct,…
Read tutorial
The primal SVM asks for a weight vector in feature space. The trained model never hands you one. It hands you a sum over training points instead — and that…
Read tutorial
XGBoost has a reputation problem. It gets treated like a secret weapon—a mysterious algorithm that wins competitions and powers production systems while…
Read tutorial
Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by…
Read tutorial