Skip to content
intermediate

XGBoost’s Regularized Objective: Derive Leaf Weights and Split Gain

Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by…

Published 2026-10-02Updated 2026-10-0411 min read
A serene view of fluffy white clouds against a vibrant blue sky, perfect for backgrounds.
A serene view of fluffy white clouds against a vibrant blue sky, perfect for backgrounds. Photo by luian C on Pexels.

Two candidate splits sit in front of you. One carves the data into two clean, balanced groups. The other produces a lopsided partition that looks worse by every impurity measure you learned from decision trees. XGBoost picks the second one — or refuses to split at all. If that has ever felt like the library ignoring your intuition, the gap is not in the library. It is in the objective.

Once you see the objective, the leaf weights and the split decision stop being constants handed down by the library. They become algebra you can do on paper.

This article assumes you already know that gradient boosting adds trees sequentially and that each tree corrects the previous ensemble. If that is still fuzzy, the earlier XGBoost overview is the right place to start. Here we go one level down: the actual quantity XGBoost minimizes at each round, and what falls out of it.

What the Objective Is Actually Optimizing

At boosting round tt, the model's prediction for sample ii is the running sum of all previous trees plus the new tree ftf_t:

y^i(t)=y^i(t−1)+ft(xi)\hat{y}_i^{(t)} = \hat{y}_i^{(t-1)} + f_t(x_i)

The full objective for the round is data loss plus a complexity penalty on every tree built so far:

Obj(t)=∑i=1nl(yi,y^i(t))+∑k=1tΩ(fk)\text{Obj}^{(t)} = \sum_{i=1}^{n} l\left(y_i, \hat{y}_i^{(t)}\right) + \sum_{k=1}^{t} \Omega(f_k)

The penalty is inside the objective, not bolted on afterward. That distinction is the whole article.

Because the previous trees are already fixed, only ftf_t is free. We approximate the loss with a second-order Taylor expansion around the current prediction y^i(t−1)\hat{y}_i^{(t-1)}:

Obj(t)≈∑i=1n[l(yi,y^i(t−1))+gift(xi)+12hift(xi)2]+Ω(ft)+const\text{Obj}^{(t)} \approx \sum_{i=1}^{n} \left[ l(y_i, \hat{y}_i^{(t-1)}) + g_i f_t(x_i) + \tfrac{1}{2} h_i f_t(x_i)^2 \right] + \Omega(f_t) + \text{const}

Here gig_i and hih_i are the first and second derivatives of the loss with respect to the current prediction, evaluated at y^i(t−1)\hat{y}_i^{(t-1)}:

gi=∂y^(t−1) l(yi,y^(t−1)),hi=∂y^(t−1)2 l(yi,y^(t−1))g_i = \partial_{\hat{y}^{(t-1)}} \, l(y_i, \hat{y}^{(t-1)}), \qquad h_i = \partial^2_{\hat{y}^{(t-1)}} \, l(y_i, \hat{y}^{(t-1)})

Once the previous round is fixed, gig_i and hih_i are numbers, not functions. That is what makes the rest of the derivation tractable.

Drop the constant term l(yi,y^i(t−1))l(y_i, \hat{y}_i^{(t-1)}) — it does not depend on ftf_t, so it cannot change which tree we choose. What remains is the working objective:

Obj~(t)=∑i=1n[gift(xi)+12hift(xi)2]+Ω(ft)\tilde{\text{Obj}}^{(t)} = \sum_{i=1}^{n} \left[ g_i f_t(x_i) + \tfrac{1}{2} h_i f_t(x_i)^2 \right] + \Omega(f_t)

Note: This is where XGBoost diverges from generic gradient boosting. Classical gradient boosting fits the new tree to the negative gradient −gi-g_i. Here we keep both gig_i and hih_i, so the tree sees the local curvature of the loss, not just its slope. The approximation is only as good as the local quadratic fit — with a large learning rate or a badly initialized model, that fit is a weaker guide, which is one reason shrinkage still matters.

Knowledge check

Check your understanding

Answer this question before you continue.

With the previous boosting round fixed, what additional local loss information does the stated XGBoost approximation use beyond the gradient-only view?
Comparison Reasoning

Focus: Distinguish the local loss information used by XGBoost’s second-order approximation from gradient-only boosting.

The Regularization Term and Why It Has Two Parts

For a tree with TT leaves and leaf weights w1,…,wTw_1, \dots, w_T, XGBoost defines:

Ω(f)=γT+12λ∑j=1Twj2\Omega(f) = \gamma T + \tfrac{1}{2} \lambda \sum_{j=1}^{T} w_j^2

Two parameters, two jobs:

  • γ\gamma charges a flat price per leaf. Every leaf you create costs γ\gamma, whether or not it earns its keep. It is a per-split toll.
  • λ\lambda is an L2 penalty on leaf magnitude. It does not care how many leaves exist, only how loud each one is allowed to be.

I think of it as structure versus volume. γ\gamma controls how many leaves the tree is allowed to grow. λ\lambda controls how much any single leaf is allowed to say. The derivation below will show exactly where each one lands.

XGBoost also supports an L1 penalty on leaf weights (reg_alpha), which can drive some weights to exactly zero and induce sparsity. But L1 is non-differentiable at zero, so it needs soft-thresholding rather than the clean closed form we are about to derive. We stick to the L2 case here, and that is a deliberate choice, not an oversight.

Knowledge check

Check your understanding

Answer this question before you continue.

A learner wants to increase the flat cost of adding leaves without directly changing the L2 penalty on the size of each leaf weight. Which parameter matches that goal?
Scenario Interpretation

Focus: Identify whether gamma or lambda directly penalizes tree structure versus leaf-weight magnitude.

Grouping by Leaf: From Per-Sample to Per-Leaf

A tree predicts a constant inside each leaf. So ft(xi)=wq(xi)f_t(x_i) = w_{q(x_i)}, where q(xi)q(x_i) maps sample ii to its leaf index. Define IjI_j as the set of training indices landing in leaf jj, and define the per-leaf gradient and Hessian sums:

Gj=∑i∈Ijgi,Hj=∑i∈IjhiG_j = \sum_{i \in I_j} g_i, \qquad H_j = \sum_{i \in I_j} h_i

Substituting into the working objective turns the sum over nn samples into a sum over TT leaves:

Obj~(t)=∑j=1T[Gjwj+12(Hj+λ)wj2]+γT\tilde{\text{Obj}}^{(t)} = \sum_{j=1}^{T} \left[ G_j w_j + \tfrac{1}{2}(H_j + \lambda) w_j^2 \right] + \gamma T

This is the step most readers skip, and then wonder why the closed form appears out of nowhere. The reorganization is the point: once the tree structure qq is fixed, the leaves are independent. There is no w1w2w_1 w_2 term. Each leaf is its own one-variable quadratic, and that independence is what makes a closed-form solution possible at all.

Picture a small tree with three leaves. Each leaf collects its own GjG_j and HjH_j from the samples that fall into it. The whole objective is just three separate quadratics plus a flat γT\gamma T charge.

Deriving the Optimal Leaf Weight

Each leaf contributes a convex quadratic in wjw_j:

ℓj(wj)=Gjwj+12(Hj+λ)wj2\ell_j(w_j) = G_j w_j + \tfrac{1}{2}(H_j + \lambda) w_j^2

Because hi≥0h_i \geq 0 for convex losses and λ>0\lambda > 0, the coefficient on wj2w_j^2 is positive, so the minimum is where the derivative is zero. No iterative search needed.

∂ℓj∂wj=Gj+(Hj+λ)wj=0\frac{\partial \ell_j}{\partial w_j} = G_j + (H_j + \lambda) w_j = 0

Solving:

wj∗=−GjHj+λw_j^{*} = -\frac{G_j}{H_j + \lambda}

Read the formula. The numerator is the summed gradient direction. The denominator is summed curvature plus the penalty. The minus sign is there because GjG_j points in the direction that increases loss, so the optimal step moves against it.

Why λ\lambda in the denominator matters: it inflates the effective curvature, so the same gradient signal produces a smaller step. A leaf with a strong signal but a small HjH_j gets damped hardest. That is the flat-region protection — leaves where the loss barely curves get pulled toward zero instead of taking a confident step on thin evidence.

Two sanity checks you can run mentally:

  • As λ→0\lambda \to 0, the formula collapses to the Newton step −Gj/Hj-G_j / H_j.
  • As λ→∞\lambda \to \infty, wj∗→0w_j^{*} \to 0.

Substituting wj∗w_j^{*} back into the objective gives the structural score of the tree:

Obj~∗=−12∑j=1TGj2Hj+λ+γT\tilde{\text{Obj}}^{*} = -\frac{1}{2} \sum_{j=1}^{T} \frac{G_j^2}{H_j + \lambda} + \gamma T

Common mistake: The score is negative when the tree helps — a lower objective is better, and the data-loss term contributes a negative quantity. Readers often expect a positive "score" and then misread the sign of the gain. Keep the sign convention straight: negative is good.

Knowledge check

Check your understanding

Answer this question before you continue.

For a leaf with summed gradient $G=-6$, summed Hessian $H=3$, and $lambda=1$, what is its optimal weight under the article’s L2 formula?
Output Prediction

Focus: Apply the closed-form optimal leaf-weight equation to given gradient, Hessian, and regularization sums.

From Structural Score to Split Gain

A split replaces one leaf with a left and a right child. The change in objective is the child score minus the parent score. Writing it out:

Gain=12[GL2HL+λ+GR2HR+λ−(GL+GR)2HL+HR+λ]−γ\text{Gain} = \tfrac{1}{2} \left[ \frac{G_L^2}{H_L + \lambda} + \frac{G_R^2}{H_R + \lambda} - \frac{(G_L + G_R)^2}{H_L + H_R + \lambda} \right] - \gamma

The γ\gamma term appears once per split because the split adds exactly one leaf. That is why γ\gamma is a hard floor rather than a soft penalty: a split is kept only when the bracketed improvement strictly exceeds γ\gamma.

The bracketed term is not guaranteed to be non-negative. It is non-negative when the split genuinely reduces the regularized objective, but a bad partition can push the child-score sum below the parent score — the split makes things worse before γ\gamma is even applied. The worked example below shows exactly this case. So γ\gamma does two jobs: it vetoes weak-but-positive splits, and it adds a further penalty on top of splits that were already harmful.

Two contrasts worth holding onto:

CriterionWhat it measuresUses current model?Per-leaf price?
Decision-tree impurity gainLabel purity of the partitionNoNo
First-order gradient boostingReduction in gradient-only lossYesNo
XGBoost split gainReduction in second-order regularized lossYesYes (γ\gamma)

The shape of the formula is similar to impurity gain — parent minus children — but the quantity being measured is entirely different. Impurity criteria score partitions by label purity and have no notion of the current model's residuals, no Hessian weighting, and no per-leaf price. Without hih_i, the gain reduces to a gradient-only criterion and loses the curvature weighting that makes small-HH leaves cautious.

Knowledge check

Check your understanding

Answer this question before you continue.

According to the article’s split-gain rule, when is a split kept after accounting for gamma?
Single Choice

Focus: Use the split-gain expression to state the strict condition for retaining a candidate split.

A Worked Example Where Gamma Changes the Answer

A comparison of two candidate splits from the same parent leaf. Split A has gain 0.733 at gamma zero and −0.067 at gamma 0.8, changing from keep to reject. Split B has gains −0.600 and −1.400, so it is rejected in both cases.
Compare the candidates before and after the per-split penalty: gamma vetoes A, while B already worsens the objective.

Let's make the arithmetic visible. Use squared-error loss, where gi=y^i−yig_i = \hat{y}_i - y_i and hi=1h_i = 1 for every sample. That keeps the Hessian sums honest: HjH_j is just the count of samples in leaf jj.

Suppose a parent leaf has GP=6G_P = 6 and HP=4H_P = 4. Two candidate splits:

SplitGLG_LHLH_LGRG_RHRH_R
A5212
B3232

Take λ=1\lambda = 1. Compute the parent score first:

Parent=GP2HP+λ=365=7.2\text{Parent} = \frac{G_P^2}{H_P + \lambda} = \frac{36}{5} = 7.2

Split A:

253+13=8.667⇒GainA=12(8.667−7.2)−γ=0.733−γ\frac{25}{3} + \frac{1}{3} = 8.667 \quad \Rightarrow \quad \text{Gain}_A = \tfrac{1}{2}(8.667 - 7.2) - \gamma = 0.733 - \gamma

Split B:

93+93=6.0⇒GainB=12(6.0−7.2)−γ=−0.6−γ\frac{9}{3} + \frac{9}{3} = 6.0 \quad \Rightarrow \quad \text{Gain}_B = \tfrac{1}{2}(6.0 - 7.2) - \gamma = -0.6 - \gamma

With γ=0\gamma = 0, split A wins with gain 0.7330.733 and split B is rejected outright — its bracketed term is negative, meaning the split would make the objective worse. This is the case where the L2 penalty alone kills a split.

Now raise γ\gamma to 0.50.5. Split A's gain becomes 0.2330.233, still positive, still kept. Raise γ\gamma to 0.80.8 and split A drops below the floor: the split is rejected even though it reduces training loss. That is γ\gamma acting as a hard floor.

Now change λ\lambda instead. Set λ=10\lambda = 10 and recompute with γ=0\gamma = 0:

Parent=3614=2.571,Split A=2512+112=2.167,Split B=912+912=1.5\text{Parent} = \frac{36}{14} = 2.571, \quad \text{Split A} = \frac{25}{12} + \frac{1}{12} = 2.167, \quad \text{Split B} = \frac{9}{12} + \frac{9}{12} = 1.5

Split A's gain is now 12(2.167−2.571)=−0.202\tfrac{1}{2}(2.167 - 2.571) = -0.202. Split B's is 12(1.5−2.571)=−0.536\tfrac{1}{2}(1.5 - 2.571) = -0.536. Both are rejected. Under small λ\lambda, split A looked strong; under large λ\lambda, the same split loses to the penalty. λ\lambda changes which splits clear zero, not just how large the gains are.

SplitGain, γ=0,λ=1\gamma=0, \lambda=1Gain, γ=0.8,λ=1\gamma=0.8, \lambda=1Gain, γ=0,λ=10\gamma=0, \lambda=10
A0.733−0.067−0.202
B−0.600−1.400−0.536

The arithmetic is fully visible on purpose. Reproduce it on paper before you trust it.

What This Derivation Does and Does Not Explain

This is the exact objective for the L2 case with a fixed tree structure. Real implementations add approximate split finding, histogram binning, sparsity-aware handling of missing values, and column subsampling. None of those change the algebra above, but all of them change runtime and behavior.

A positive gain is a statement about the training objective at this round, not a guarantee of better held-out performance. The regularized objective is a proxy for generalization, not a proof of it.

The second-order approximation is local. With a large learning rate or a badly initialized model, the quadratic fit is a weaker guide, and the step the formula recommends may overshoot.

γ\gamma and λ\lambda also interact with max_depth and min_child_weight, which are structural constraints outside Ω\Omega. Tuning one while ignoring the others produces confusing results.

When this derivation is the right tool: you are writing a custom objective, debugging why a split was rejected, or reasoning about why a leaf value came out smaller than the raw gradient suggested. When it is not: routine model fitting, where the library's defaults and a validation split will teach you more per minute.

The Decision Rule to Carry Forward

When a split is rejected, check which parameter did it. If γ\gamma vetoed the split, the bracketed improvement was real but too small to pay the toll — lower γ\gamma if you want more structure. If λ\lambda flattened the leaf weights, the split may have been accepted but its leaves are whispering instead of speaking — lower λ\lambda if you want sharper leaf values. Those are different diagnoses with different fixes.

Here is the next move. Take the numbers from the worked example, change λ\lambda by a factor of ten, and predict the new leaf weights before you compute them. If your prediction matches, the objective is yours. If it does not, the gap between prediction and computation is the lesson — and it is worth more than another read-through.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

In the worked example, with gamma set to zero, what happens to both candidate splits when lambda increases from 1 to 10?
Question 1 of 2Scenario Interpretation

Focus: Interpret how increasing lambda can change split acceptance, not merely shrink leaf weights.

A candidate split has positive gain under the regularized objective. Which conclusion is justified by the article?
Question 2 of 2Misconception Check

Focus: Explain what a positive XGBoost split gain establishes and what it does not establish about generalization.

References

  1. Higgs Boson Discovery with Boosted Treesproceedings.mlr.press
  2. Exploring XGBoost: A Deep Dive — ROCm Blogsrocm.blogs.amd.com
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.