Gradient Boosting as Additive Function Fitting: Derive the Pseudo-Residuals
Fit the next tree to the error. That instruction works beautifully for squared loss and then quietly falls apart the moment you switch to log loss. The…

Key topics
Fit the next tree to the error. That instruction works beautifully for squared loss and then quietly falls apart the moment you switch to log loss. The target you compute is no longer the error. It is something else wearing the same name.
The intuition article in this series gave you the loop: build a weak learner, correct the ensemble, repeat. This piece replaces the special case with the general rule. By the end, you will be able to write down the negative gradient of any differentiable loss and derive the boosting target yourself, instead of trusting the library to hand it to you.
The Additive Model and the Empirical Loss
Fix the object we are optimizing before touching any derivative.
An additive model builds a prediction as a sum of stagewise learners:
Here is an initial constant, is the learner added at stage , and is a shrinkage multiplier we will return to. The initial model is problem-specific. For least-squares regression, the conventional choice is the mean of the training targets; for binary classification with log loss, it is the log-odds of the base rate.
The empirical loss over training examples is
Read that carefully. is a function of the entire prediction vector , not of a finite parameter vector. There is no to differentiate against. The variable is the function itself.
Two constraints define the game:
- Forward stagewise. At stage we add one and never revisit earlier stages. No joint re-optimization.
- Restricted function class. Each is drawn from a class , usually fixed-depth regression trees. That restriction is what makes the step tractable. It is also what makes the result greedy rather than globally optimal.
State the assumptions up front, because they bound everything that follows. The per-example loss must be differentiable in its second argument. The training set is finite. The class is closed under the operations we apply later — specifically, we need to be able to fit a learner to arbitrary real-valued targets, which regression trees can do.
Knowledge check
Check your understanding
Answer this question before you continue.
Why the Ideal Next Learner Is Intractable
The exact stagewise objective is clean to write and impossible to solve directly:
To solve this exactly, you would enumerate every function in and evaluate the total loss after adding each one. For a tree class, that means searching over every possible partition of the feature space at every depth. The search space is combinatorial and the loss surface is not convex in the tree structure.
So we cheat, but we cheat in a principled direction. Instead of solving the stagewise problem, we take a step in function space that mirrors gradient descent in parameter space. In parameter space, we cannot jump to the minimum, so we compute the gradient and move a small distance along it. In function space, we cannot enumerate , so we compute the steepest-descent direction and fit the best available learner to it.
Note: This is a greedy heuristic. It does not solve the stagewise problem exactly. It solves a first-order approximation of it, and the approximation is only as good as the learner class and the step size allow.
Deriving the Negative Gradient Target
Now the derivation. It is short, and every symbol has a job.
Differentiate the empirical loss with respect to the current prediction for a single training example. Define
This is the per-example gradient: how much the loss changes when you nudge the model's prediction for example by a small amount, holding every other prediction fixed.
The steepest-descent direction in function space is the negative gradient, evaluated at the current model's predictions. So the ideal next learner would satisfy
We cannot fit an arbitrary function, so we fit the best function in to the values . Define the pseudo-residual
and train as a regression model with targets .
That is the whole trick. The "residual" label is a historical accident: for squared loss, happens to equal the ordinary residual. For every other loss, it is a gradient-derived quantity with different units, a different range, and a different meaning.
Convention matters here. You will see the update written as and also as . Both are correct; they differ in whether the learner is fit to (add) or to (subtract). Pick one and stay consistent, because mixing conventions flips the sign of your entire ensemble. The squared-loss case aligns pseudo-residuals with ordinary residuals only when the sign convention is held fixed.
One more boundary. If is not differentiable in the prediction — 0-1 loss is the canonical example — the derivation does not apply. You cannot take the gradient of a step function. This is why classification uses smooth surrogates like log loss or the logistic loss: they approximate the decision boundary you care about while remaining differentiable. Hinge loss is a partial exception: it is differentiable almost everywhere, but its flat regions and kink mean the gradient carries no information at some points, so the derivation applies only where the gradient is defined and nonzero.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: Squared Loss and Log Loss
Two cases make the derivation concrete. One where pseudo-residuals are residuals, one where they visibly are not.
Squared loss
Differentiate with respect to :
The negative gradient is
That is the ordinary residual. The derivation recovers the intuition exactly, under the convention that the learner is added to the ensemble.
Knowledge check
Check your understanding
Answer this question before you continue.
Log loss for binary classification
Let be the log-odds score and the predicted probability. The per-example loss is
Differentiate with respect to . Using :
The negative gradient is
This is a residual in probability space — the gap between the label and the predicted probability — but it is not a residual in score space. The tree is fit to values in while the ensemble accumulates scores in . The units do not match, and that mismatch has a practical consequence: the leaf values produced by fitting to are not on the same scale as the raw label, so implementations refine leaf values after the tree structure is fixed.
| Loss | Gradient | Pseudo-residual | Range of |
|---|---|---|---|
| Squared | Same as label | ||
| Log loss |
Same derivation, two different objects. The table is the whole point of this section.
Step Size and Line Search
Fitting to the pseudo-residuals gives you a direction. It does not give you a distance.
After the tree structure is fixed, solve a one-dimensional line search for the multiplier that minimizes the loss along that direction:
For squared loss, this has a closed form: is the least-squares coefficient of the residual on the fitted values, which reduces to a simple ratio of sums. For log loss and most other losses, there is no closed form and you solve it numerically or approximate it.
In practice, libraries do not run a full line search. They use a fixed shrinkage parameter — the learning rate — and often compute per-leaf optimal values instead of a single global . Shrinkage and line search are related but not identical: line search finds the best step along the direction, shrinkage deliberately takes a smaller step than optimal to leave room for later learners to correct.
The tradeoff is direct. Small steps need more learners to reach the same loss, which costs training time and memory. Large steps risk overshooting the minimum and destabilizing the ensemble, especially when the loss surface is steep near the current predictions. The learning rate and the number of estimators are coupled; tuning one without the other is guesswork.
Knowledge check
Check your understanding
Answer this question before you continue.
From Derivation to the Boosting Procedure
Here is the algorithm, with each step mapped to the derivation that justifies it. The two step-size branches are kept separate, because they produce different updates.
- Initialize to a constant that minimizes the loss — the mean for squared loss, the log-odds for log loss. This is the starting point of the additive model.
- Compute pseudo-residuals at the current predictions. This is the negative gradient step.
- Fit a learner to the targets using the restricted class . This is the projection of the steepest-descent direction onto the available function class.
- Choose the step size. Either solve the line search for , or fix a shrinkage parameter . These are alternatives, not a single combined step.
- Update the ensemble. With line search: . With fixed shrinkage: . Use one branch consistently; do not mix and in the same update.
- Repeat until a stopping condition is met.
Implementations diverge from this textbook version in three places worth knowing. XGBoost and similar libraries use a second-order Taylor approximation of the loss, replacing the gradient-only step with a Newton step that uses both gradient and Hessian. They compute per-leaf optimal values rather than a single global step size. And they add explicit regularization terms on leaf weights and tree complexity. None of these change the core derivation; they refine the step.
For stopping, validation-based early stopping is the honest choice. Training loss decreases monotonically as you add learners, so it cannot tell you when to stop. A held-out set can.
When Pseudo-Residuals Are Not Residuals
This is the section that prevents the most expensive debugging mistakes.
Pseudo-residuals equal ordinary residuals only for squared loss under a consistent sign convention. For every other loss, they are gradient-derived targets with different units and ranges. Three consequences follow.
Debugging. A shrinking pseudo-residual magnitude does not mean the model is fitting the label better in the original units. For log loss, is bounded in regardless of how wrong the score is. A model can have small pseudo-residuals and still produce badly calibrated probabilities.
Loss choice. Switching loss changes the target distribution. Tree depth and learning rate tuned for squared loss may not transfer to log loss, because the targets have different variance, different scale, and different sensitivity to outliers. Re-tune when you switch.
Assumption boundary. The derivation requires differentiability and a function class rich enough to approximate the gradient direction. When is too weak — trees that are too shallow, or a linear class on a nonlinear problem — boosting stalls regardless of iteration count. Adding more learners to a class that cannot represent the gradient direction just accumulates the same correction over and over.
Common mistake: Treating pseudo-residuals as "the error" and expecting them to shrink toward zero in the label's units. They shrink toward zero in the gradient's units, which is a different thing.
What to Do Next
The derivation is a decision rule, not a ritual. When you can write down for your loss, you can derive the boosting target yourself and stop treating the library as a black box. When you cannot — because the loss is non-differentiable, or the function class cannot represent the gradient direction — the derivation stops being the right tool, and you need a different approach.
The practical next step is to run a boosting experiment and watch how learning rate and tree depth interact with the derived update. Fit a model with a small learning rate and many estimators, then a large learning rate and few. Compare the validation curves. The shape of those curves is the line search and the shrinkage parameter doing their work in the open.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


