
Derive AdaBoost’s Weight Updates From Exponential Loss
AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorialArticles whose authoritative article variant is theory and whose primary outcome is formal derivation, assumptions, or mathematical reasoning.
Tagged articles
47 articles in this tag.

AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorial
Retrain the same model on a different random split and the predictions move. Nobody changed the code. Nobody changed the hyperparameters. The wobble is not…
Read tutorial
You have a thousand bootstrap replicates sitting in a NumPy array. You have a histogram. What you do not have is the two numbers you actually need to…
Read tutorial
A random forest hands you an out-of-bag score, and someone nearby says, "Each tree leaves out about a third of the data." That number feels like a setting.…
Read tutorial
Two forecasters can post the same Brier score and deserve opposite verdicts. One is honest but says almost nothing. The other is sharp but lies about its…
Read tutorial
A model returns 0.4. The default cutoff turns that into "negative," and nobody asked whether that was the right call.
Read tutorial
Two models. Mean cross-validation scores of 0.842 and 0.847. The instinct is immediate: the second one wins.
Read tutorial
A satisfaction rating of 4 when the truth was 5 is not the same kind of mistake as predicting 1. Most models cannot tell the difference. This one can — but…
Read tutorial
Two points sit the same distance from a dense blob. DBSCAN labels one a cluster member and the other noise. That looks arbitrary until you stop asking how…
Read tutorial
A fully grown decision tree can score near-perfect on the data it was trained on and still lose to a single split on data it has never seen. The usual…
Read tutorial
You fit a tree, print it, and there it is: the root splits on worst radius, not mean texture. The library made a choice. You can't say why.
Read tutorial
You can write theta = theta - lr * grad from memory and still freeze when someone asks where grad came from. That gap is the whole problem. The update rule…
Read tutorial
You can fit a logistic regression in one line and read its coefficients in the next. What almost nobody shows you is where the loss function comes from.…
Read tutorial
A model degrades in production. Three engineers offer three diagnoses: "the inputs drifted," "the class balance changed," "the meaning changed." All three…
Read tutorial
You picked 2^18 buckets. You trained the model. The metrics came back slightly worse than your explicit-vocabulary baseline, and now you are staring at the…
Read tutorial
Two houses sit side by side in your table. One is 1,800 square feet with 3 bedrooms. The other is 2,000 square feet with 5 bedrooms. To a k-NN model, the…
Read tutorial
You can write the Gaussian mixture density in a single line. The trouble starts the moment you take the log.
Read tutorial
Fit the next tree to the error. That instruction works beautifully for squared loss and then quietly falls apart the moment you switch to log loss. The…
Read tutorial
"Gradient descent works" and "gradient descent is guaranteed to converge" sound like the same sentence. They are not. One is a hope backed by experience.…
Read tutorial
Change one parameter, get a different tree. Same points, same distance metric, same greedy loop — yet the dendrogram reorganizes itself. That parameter is…
Read tutorial
You fit an Isolation Forest, call decision_function, and get a column of numbers. You sort them, pick a cutoff, and ship the alerts. But if someone asks…
Read tutorial
Fit K-means twice on the same data with the same k, and you can get two different clusterings. The second run often reports a lower cost than the first.…
Read tutorial
Fit KNN regression with k=1 and the prediction curve threads through every training point. Set k=50 and the same curve flattens toward a horizontal line.…
Read tutorial
Fit a lasso on a dataset with a dozen features and something strange happens. A few coefficients come back as exactly 0.0. Not 0.0031. Not 0.0007. Zero.…
Read tutorial
You fit a line, plot the residuals, and they look like a clean cloud. No funnel, no curve, no obvious outlier. It is tempting to read that plot as a…
Read tutorial
You fit LinearDiscriminantAnalysis, plot the boundary, and get a straight line. Then someone asks why it is straight, and the honest answer is "because the…
Read tutorial
Call LinearRegression().fit(X, y) and coefficients appear in milliseconds. The library never shows you the equation it solved, why that equation has a…
Read tutorial
A point can sit closer to its neighbors than anything else in the dataset and still be the anomaly. That is the contradiction Local Outlier Factor was…
Read tutorial
Two analysts open the same CSV. Both see the same 12% of rows with a blank in the income column. One drops those rows and reports a mean of 54,000. The…
Read tutorial
You run a grid search. The leaderboard prints a winner at 0.91. You ship it, and the first honest measurement lands near 0.86. Nothing broke. The model did…
Read tutorial
A multiclass model hands you K raw scores. The obvious move — divide each by their sum — collapses the first time a score goes negative. Here is the fix,…
Read tutorial
The intuition article told you Naive Bayes "combines prior and evidence." Then you tried to write $P(x_1, x_2, \dots, x_n \mid y)$ for real data, and the…
Read tutorial
A coefficient can be unbiased on average and still be wrong in your one dataset. Worse, it can be unbiased and still predict badly, and unbiased and still…
Read tutorial
You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the…
Read tutorial
You can call PCA(), read explained_variance_ratio_, and get a working result. Then someone asks why the principal directions are eigenvectors of the…
Read tutorial
A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.
Read tutorial
A model can keep the same weights, the same features, and the same threshold, and still report precision 0.9 on your test set and precision 0.3 in…
Read tutorial
Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.
Read tutorial
Everyone repeats that a forest reduces variance. Almost nobody writes down the formula that governs the reduction — and that formula contains a term that…
Read tutorial
A classification tree asks which split makes the labels purest. A regression tree has no labels to purify — only numbers that scatter. So what is it…
Read tutorial
Two features that measure almost the same thing can make ordinary least squares produce coefficients that swing wildly while the predictions barely move.…
Read tutorial
A prediction interval built from a model's own residuals has no guarantee at all. Split conformal prediction replaces that hope with a finite-sample…
Read tutorial
One point sits far from the rest of the scatter, and the fitted line visibly tilts toward it. Move that point twice as far away. Squared error answers with…
Read tutorial
Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing…
Read tutorial
Most explanations of support vector machines stop at the picture: two clouds of points, a line between them, a gap on either side. That picture is correct,…
Read tutorial
The primal SVM asks for a weight vector in feature space. The trained model never hands you one. It hands you a sum over training points instead — and that…
Read tutorial
Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by…
Read tutorial