
Derive AdaBoost’s Weight Updates From Exponential Loss
AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorialCombining multiple predictive learners through bagging, boosting, voting, stacking, or related aggregation strategies.
Tagged articles
16 articles in this tag.

AdaBoost's algorithm listing looks like three separate rules stitched together: fit a weak learner, compute a coefficient, multiply the sample weights by…
Read tutorial
If you already understand decision trees, you know the dilemma: a deep tree memorizes the training data and fails on new data, while a shallow tree is too…
Read tutorial
A random forest hands you an out-of-bag score, and someone nearby says, "Each tree leaves out about a third of the data." That number feels like a setting.…
Read tutorial
A stacking classifier can score beautifully on your test set for a reason that has nothing to do with skill: the meta-model was trained on predictions the…
Read tutorial
One score is a snapshot. A distribution is a measurement. If you fit Random Forest and Extra Trees once, watch Extra Trees edge ahead by 0.01 accuracy, and…
Read tutorial
Random forests grow many trees independently and average their votes. Gradient boosting grows trees one at a time, each one aimed at the mistakes the…
Read tutorial
Fit the next tree to the error. That instruction works beautifully for squared loss and then quietly falls apart the moment you switch to log loss. The…
Read tutorial
A single accuracy number tells you almost nothing. A controlled experiment tells you whether that number means anything at all.
Read tutorial
Everyone repeats that a forest reduces variance. Almost nobody writes down the formula that governs the reduction — and that formula contains a term that…
Read tutorial
oob_score=True looks like a free validation set. It is not free, and it is not quite a validation set. It is a running estimate computed on the same rows…
Read tutorial
A single decision tree can feel like a confident liar: it nails your training data, then wobbles on new data, and reshapes its entire structure when one…
Read tutorial
"Extra Trees" does not mean more trees. It means extra randomness — and that single distinction changes how the ensemble learns, how fast it trains, and…
Read tutorial
Your stacking classifier reports 0.97 accuracy during training. On held-out data, it drops to 0.81 — barely better than the single best base model. Nothing…
Read tutorial
More models do not automatically mean a better model. An ensemble only wins when its members make different mistakes—and the way you combine them…
Read tutorial
XGBoost has a reputation problem. It gets treated like a secret weapon—a mysterious algorithm that wins competitions and powers production systems while…
Read tutorial
Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by…
Read tutorial