XGBoost’s Regularized Objective: Derive Leaf Weights and Split Gain
Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by…

Key topics
Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by every impurity measure you learned from decision trees. XGBoost picks the second one — or refuses to split at all. If that has ever felt like the library ignoring your intuition, the gap is not in the library. It is in the objective.
Once you see the objective, the leaf weights and the split decision stop being constants handed down by the library. They become algebra you can do on paper.
This article assumes you already know that gradient boosting adds trees sequentially and that each tree corrects the previous ensemble. If that is still fuzzy, the earlier XGBoost overview is the right place to start. Here we go one level down: the actual quantity XGBoost minimizes at each round, and what falls out of it.
What the Objective Is Actually Optimizing
At boosting round , the model's prediction for sample is the running sum of all previous trees plus the new tree :
The full objective for the round is data loss plus a complexity penalty on every tree built so far:
The penalty is inside the objective, not bolted on afterward. That distinction is the whole article.
Because the previous trees are already fixed, only is free. We approximate the loss with a second-order Taylor expansion around the current prediction :
Here and are the first and second derivatives of the loss with respect to the current prediction, evaluated at :
Once the previous round is fixed, and are numbers, not functions. That is what makes the rest of the derivation tractable.
Drop the constant term — it does not depend on , so it cannot change which tree we choose. What remains is the working objective:
Note: This is where XGBoost diverges from generic gradient boosting. Classical gradient boosting fits the new tree to the negative gradient . Here we keep both and , so the tree sees the local curvature of the loss, not just its slope. The approximation is only as good as the local quadratic fit — with a large learning rate or a badly initialized model, that fit is a weaker guide, which is one reason shrinkage still matters.
Knowledge check
Check your understanding
Answer this question before you continue.
The Regularization Term and Why It Has Two Parts
For a tree with leaves and leaf weights , XGBoost defines:
Two parameters, two jobs:
- charges a flat price per leaf. Every leaf you create costs , whether or not it earns its keep. It is a per-split toll.
- is an L2 penalty on leaf magnitude. It does not care how many leaves exist, only how loud each one is allowed to be.
I think of it as structure versus volume. controls how many leaves the tree is allowed to grow. controls how much any single leaf is allowed to say. The derivation below will show exactly where each one lands.
XGBoost also supports an L1 penalty on leaf weights (reg_alpha), which can drive some weights to exactly zero and induce sparsity. But L1 is non-differentiable at zero, so it needs soft-thresholding rather than the clean closed form we are about to derive. We stick to the L2 case here, and that is a deliberate choice, not an oversight.
Knowledge check
Check your understanding
Answer this question before you continue.
Grouping by Leaf: From Per-Sample to Per-Leaf
A tree predicts a constant inside each leaf. So , where maps sample to its leaf index. Define as the set of training indices landing in leaf , and define the per-leaf gradient and Hessian sums:
Substituting into the working objective turns the sum over samples into a sum over leaves:
This is the step most readers skip, and then wonder why the closed form appears out of nowhere. The reorganization is the point: once the tree structure is fixed, the leaves are independent. There is no term. Each leaf is its own one-variable quadratic, and that independence is what makes a closed-form solution possible at all.
Picture a small tree with three leaves. Each leaf collects its own and from the samples that fall into it. The whole objective is just three separate quadratics plus a flat charge.
Deriving the Optimal Leaf Weight
Each leaf contributes a convex quadratic in :
Because for convex losses and , the coefficient on is positive, so the minimum is where the derivative is zero. No iterative search needed.
Solving:
Read the formula. The numerator is the summed gradient direction. The denominator is summed curvature plus the penalty. The minus sign is there because points in the direction that increases loss, so the optimal step moves against it.
Why in the denominator matters: it inflates the effective curvature, so the same gradient signal produces a smaller step. A leaf with a strong signal but a small gets damped hardest. That is the flat-region protection — leaves where the loss barely curves get pulled toward zero instead of taking a confident step on thin evidence.
Two sanity checks you can run mentally:
- As , the formula collapses to the Newton step .
- As , .
Substituting back into the objective gives the structural score of the tree:
Common mistake: The score is negative when the tree helps — a lower objective is better, and the data-loss term contributes a negative quantity. Readers often expect a positive "score" and then misread the sign of the gain. Keep the sign convention straight: negative is good.
Knowledge check
Check your understanding
Answer this question before you continue.
From Structural Score to Split Gain
A split replaces one leaf with a left and a right child. The change in objective is the child score minus the parent score. Writing it out:
The term appears once per split because the split adds exactly one leaf. That is why is a hard floor rather than a soft penalty: a split is kept only when the bracketed improvement strictly exceeds .
The bracketed term is not guaranteed to be non-negative. It is non-negative when the split genuinely reduces the regularized objective, but a bad partition can push the child-score sum below the parent score — the split makes things worse before is even applied. The worked example below shows exactly this case. So does two jobs: it vetoes weak-but-positive splits, and it adds a further penalty on top of splits that were already harmful.
Two contrasts worth holding onto:
| Criterion | What it measures | Uses current model? | Per-leaf price? |
|---|---|---|---|
| Decision-tree impurity gain | Label purity of the partition | No | No |
| First-order gradient boosting | Reduction in gradient-only loss | Yes | No |
| XGBoost split gain | Reduction in second-order regularized loss | Yes | Yes () |
The shape of the formula is similar to impurity gain — parent minus children — but the quantity being measured is entirely different. Impurity criteria score partitions by label purity and have no notion of the current model's residuals, no Hessian weighting, and no per-leaf price. Without , the gain reduces to a gradient-only criterion and loses the curvature weighting that makes small- leaves cautious.
Knowledge check
Check your understanding
Answer this question before you continue.
A Worked Example Where Gamma Changes the Answer
Let's make the arithmetic visible. Use squared-error loss, where and for every sample. That keeps the Hessian sums honest: is just the count of samples in leaf .
Suppose a parent leaf has and . Two candidate splits:
| Split | ||||
|---|---|---|---|---|
| A | 5 | 2 | 1 | 2 |
| B | 3 | 2 | 3 | 2 |
Take . Compute the parent score first:
Split A:
Split B:
With , split A wins with gain and split B is rejected outright — its bracketed term is negative, meaning the split would make the objective worse. This is the case where the L2 penalty alone kills a split.
Now raise to . Split A's gain becomes , still positive, still kept. Raise to and split A drops below the floor: the split is rejected even though it reduces training loss. That is acting as a hard floor.
Now change instead. Set and recompute with :
Split A's gain is now . Split B's is . Both are rejected. Under small , split A looked strong; under large , the same split loses to the penalty. changes which splits clear zero, not just how large the gains are.
| Split | Gain, | Gain, | Gain, |
|---|---|---|---|
| A | 0.733 | −0.067 | −0.202 |
| B | −0.600 | −1.400 | −0.536 |
The arithmetic is fully visible on purpose. Reproduce it on paper before you trust it.
What This Derivation Does and Does Not Explain
This is the exact objective for the L2 case with a fixed tree structure. Real implementations add approximate split finding, histogram binning, sparsity-aware handling of missing values, and column subsampling. None of those change the algebra above, but all of them change runtime and behavior.
A positive gain is a statement about the training objective at this round, not a guarantee of better held-out performance. The regularized objective is a proxy for generalization, not a proof of it.
The second-order approximation is local. With a large learning rate or a badly initialized model, the quadratic fit is a weaker guide, and the step the formula recommends may overshoot.
and also interact with max_depth and min_child_weight, which are structural constraints outside . Tuning one while ignoring the others produces confusing results.
When this derivation is the right tool: you are writing a custom objective, debugging why a split was rejected, or reasoning about why a leaf value came out smaller than the raw gradient suggested. When it is not: routine model fitting, where the library's defaults and a validation split will teach you more per minute.
The Decision Rule to Carry Forward
When a split is rejected, check which parameter did it. If vetoed the split, the bracketed improvement was real but too small to pay the toll — lower if you want more structure. If flattened the leaf weights, the split may have been accepted but its leaves are whispering instead of speaking — lower if you want sharper leaf values. Those are different diagnoses with different fixes.
Here is the next move. Take the numbers from the worked example, change by a factor of ten, and predict the new leaf weights before you compute them. If your prediction matches, the objective is yours. If it does not, the gap between prediction and computation is the lesson — and it is worth more than another read-through.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


