Skip to content
intermediate

Squared Error vs Huber Loss: How Large Residuals Change the Objective

One point sits far from the rest of the scatter, and the fitted line visibly tilts toward it. Move that point twice as far away. Squared error answers with…

Published 2026-10-02Updated 2026-10-0410 min read
Close-up of a futuristic toy robot with blue eyes, showcasing modern technology indoors.
Close-up of a futuristic toy robot with blue eyes, showcasing modern technology indoors. Photo by Pavel Danilyuk on Pexels.

One point sits far from the rest of the scatter, and the fitted line visibly tilts toward it. Move that point twice as far away. Squared error answers with four times the penalty and twice the pull. Huber answers with a penalty that grows slowly and a pull that does not grow at all. That gap is the whole article, and it is a design decision, not a default.

The Residual Is the Unit of Blame

Every regression loss starts with the same small object: the residual. For a prediction y^i\hat{y}_i and a target yiy_i, define

ri=yi−y^ir_i = y_i - \hat{y}_i

I use the target-minus-prediction convention throughout. The sign tells you which side of the prediction the target fell on; the magnitude tells you how far off you were. Flip the convention and the loss values stay identical, because both losses here depend on ∣ri∣|r_i| or ri2r_i^2.

Two levels of aggregation matter, and mixing them up causes real confusion:

  • Per-sample loss ℓ(ri)\ell(r_i): the penalty for one prediction.
  • Objective J(θ)=∑iℓ(ri)J(\theta) = \sum_i \ell(r_i): the total penalty the fitting procedure minimizes over all training points.

The model family stays fixed. We are still fitting a linear predictor y^=Xθ\hat{y} = X\theta. The only thing changing is the function that converts a residual into a number the optimizer wants to shrink. That function encodes a judgment: what kind of error do we consider acceptable? Squared error and Huber give different answers.

Knowledge check

Check your understanding

Answer this question before you continue.

If a learner switches from target-minus-prediction to prediction-minus-target, what happens to the squared-error and Huber loss values for each example?
Single Choice

Focus: Explain how the residual sign convention relates to the size of squared and Huber losses.

Squared Error: A Quadratic Tax on Being Wrong

The per-point squared error is

ℓsq(r)=12r2\ell_{\text{sq}}(r) = \frac{1}{2}r^2

The 12\frac{1}{2} is cosmetic. It cancels when you differentiate, which is why many implementations write it that way. The summed objective is J(θ)=12∑iri2J(\theta) = \frac{1}{2}\sum_i r_i^2, and minimizing it is ordinary least squares.

Differentiate with respect to the residual:

dℓsqdr=r\frac{d\ell_{\text{sq}}}{dr} = r

The penalty grows quadratically, but the gradient with respect to the residual grows linearly. That distinction is where the trouble lives.

Double a residual and its contribution to the objective quadruples, while its residual-gradient doubles. A point ten times further away than a typical point contributes one hundred times as much to the objective and exerts ten times the residual-gradient. The fit does not merely notice the extreme point; it reorganizes itself around it.

This is the mechanism behind the mean-seeking behavior of least squares. Minimizing squared error drives the fitted value toward the conditional mean of the targets. That is exactly right when large errors are genuinely worse than small ones and the data is clean. It is exactly wrong when a handful of extreme points are measurement noise, because the mean is not robust to them.

Common mistake: Treating the outlier problem as a data-cleaning problem that a robust loss will solve. The loss changes how much influence a point has. It does not decide whether the point deserves that influence.

Knowledge check

Check your understanding

Answer this question before you continue.

For squared error defined as one-half the residual squared, a residual changes from 3 to 6. How do that example's loss and residual-gradient change?
Comparison Reasoning

Focus: Predict how doubling a residual changes squared-error loss and its residual-gradient.

Huber Loss: Quadratic Near Zero, Linear in the Tail

Two aligned plots compare squared error and Huber loss against residual r. In the loss plot, the curves match within ±δ, then Huber grows linearly while squared error curves upward. In the gradient plot, squared error follows a rising diagonal, while Huber follows it within ±δ and stays flat at ±δ outside that range.
Huber preserves squared-error behavior near zero but caps the residual-gradient in the tails.

Huber loss keeps the quadratic penalty for small residuals and switches to a linear penalty once a residual crosses a threshold δ\delta:

ℓδ(r)={12r2∣r∣≤δδ(∣r∣−12δ)∣r∣>δ\ell_\delta(r) = \begin{cases} \frac{1}{2}r^2 & |r| \le \delta \\ \delta\left(|r| - \frac{1}{2}\delta\right) & |r| > \delta \end{cases}

The two pieces meet at ∣r∣=δ|r| = \delta. Check the value: at r=δr = \delta, the quadratic piece gives 12δ2\frac{1}{2}\delta^2, and the linear piece gives δ(δ−12δ)=12δ2\delta(\delta - \frac{1}{2}\delta) = \frac{1}{2}\delta^2. Same number. Check the slope: the quadratic piece has slope δ\delta at that point, and the linear piece has slope δ\delta for all r>δr > \delta. Same slope. The loss is continuous and smooth at the changepoint, which is why gradient-based fitting behaves well.

Now differentiate:

dℓδdr={r∣r∣≤δδ⋅sign(r)∣r∣>δ\frac{d\ell_\delta}{dr} = \begin{cases} r & |r| \le \delta \\ \delta \cdot \text{sign}(r) & |r| > \delta \end{cases}

Inside the quadratic region, the residual-gradient is the residual itself. Outside it, the residual-gradient is capped at magnitude δ\delta. No matter how far a point drifts, its residual-gradient never exceeds δ\delta.

That cap is the entire robustness story. Interpret δ\delta as the boundary between "ordinary noise" and "large deviation." It is a modeling decision about what counts as an outlier, not a universal constant.

Two limits are worth memorizing. As δ→∞\delta \to \infty, every residual falls in the quadratic region and Huber becomes squared error. As δ→0\delta \to 0, the quadratic region shrinks to nothing and Huber approaches absolute error, whose residual-gradient is always ±1\pm 1. Huber lives between those two poles, and δ\delta is the dial.

One property keeps Huber practical: it is convex. The piecewise definition looks like a kink, but the slopes match at the boundary, so the function curves smoothly and gradient descent still converges reliably.

From Residual-Gradient to Parameter Pull

Here is the bridge that keeps this from becoming a story about a single number. The residual-gradient dℓdr\frac{d\ell}{dr} is not the same thing as the gradient with respect to the model parameters. For a linear predictor y^i=xi⊤θ\hat{y}_i = x_i^\top \theta, the chain rule gives

∂ℓ(ri)∂θ=dℓdri⋅∂ri∂θ=dℓdri⋅(−xi)\frac{\partial \ell(r_i)}{\partial \theta} = \frac{d\ell}{dr_i} \cdot \frac{\partial r_i}{\partial \theta} = \frac{d\ell}{dr_i} \cdot (-x_i)

The parameter gradient for one example is the residual-gradient multiplied by the feature vector. Huber caps the first factor at δ\delta; it does not cap the second. A point with large feature values still contributes a larger parameter update than a point with small feature values, even when both sit in the linear region.

So the honest statement is narrower than "Huber caps a point's pull." Huber caps the loss-side factor. The feature values still scale the parameter contribution. If you want to reason about which points steer the fit, you need both pieces: how far the residual is, and how large the feature vector is.

Knowledge check

Check your understanding

Answer this question before you continue.

Two examples are both in Huber's linear tail and have the same residual sign. One has a feature vector with a larger magnitude. Which statement about their parameter-gradient contributions is supported?
Misconception Check

Focus: Distinguish Huber's cap on the residual-gradient from the feature-scaled parameter gradient.

Worked Comparison: One Extreme Residual, Two Objectives

Numbers beat adjectives. Take five points with residuals of −1-1, 0.50.5, −0.5-0.5, 11, and one extreme point at r=10r = 10. Fix δ=1\delta = 1.

Residual rrSquared error 12r2\frac{1}{2}r^2Huber, δ=1\delta=1Squared residual-gradient rrHuber residual-gradient
−1-10.50.5−1-1−1-1
0.50.50.1250.1250.50.50.50.5
−0.5-0.50.1250.125−0.5-0.5−0.5-0.5
110.50.51111
101050.09.5101011

The four ordinary points behave identically under both losses. The extreme point is where the objectives diverge. Squared error charges it 50; Huber charges it 9.5. More importantly, squared error assigns it a residual-gradient of 10, while Huber assigns it a residual-gradient of 1 — the same as a residual of 1.

Now double the extreme residual to r=20r = 20:

Residual rrSquared errorHuber, δ=1\delta=1Squared residual-gradientHuber residual-gradient
2020200.019.5202011

Squared error's penalty quadrupled and its residual-gradient doubled. Huber's penalty roughly doubled and its residual-gradient did not move at all. The point can run to infinity and its loss-side factor stays pinned at δ\delta. Its parameter contribution still scales with its features, but the residual no longer amplifies it.

You can reproduce and modify this in a few lines:

import numpy as np

def huber(delta, r):
    r = np.asarray(r, dtype=float)
    return np.where(np.abs(r) <= delta, 0.5 * r**2,
                    delta * (np.abs(r) - 0.5 * delta))

r = np.array([-1.0, 0.5, -0.5, 1.0, 10.0])
print("squared:", 0.5 * r**2)
print("huber:  ", huber(1.0, r))

Change delta to 0.5 or 5.0 and watch the extreme point's contribution move. That sensitivity is the next section.

Knowledge check

Check your understanding

Answer this question before you continue.

Using Huber loss with δ = 1, what are the loss and residual-gradient for r = 10?
Output Prediction

Focus: Calculate Huber loss and residual-gradient for a residual beyond the threshold.

What Delta Actually Controls

δ\delta is not a magic number. It is a residual-scale threshold that controls where the loss transitions from quadratic to linear. Moving it changes the fit's personality.

  • Small δ\delta: most residuals land in the linear region. The fit behaves closer to median-seeking, which can ignore large errors that are genuinely informative.
  • Large δ\delta: most residuals land in the quadratic region. The fit behaves almost like least squares, and the robustness you wanted largely disappears.

δ\delta is also scale-dependent. A residual of 10 is extreme if your targets range from 0 to 20 and unremarkable if they range from 0 to 10,000. Set δ\delta relative to the spread of the target, not as a fixed constant copied from a tutorial.

The practical consequence of a bad choice is symmetric. Too small, and you underfit real signal because the model treats legitimate variation as noise. Too large, and you inherit the exact outlier sensitivity you switched losses to avoid.

In practice, δ\delta gets tuned alongside other hyperparameters on a validation set. The criterion is not "which δ\delta makes the residuals look cleanest." The criterion is the evaluation metric or downstream cost you actually care about — validation RMSE, mean absolute error, or a business cost tied to specific error sizes. Pick the δ\delta that wins on that metric, not the one that matches your intuition about which points are outliers.

Tip: Before tuning δ\delta, plot the residuals from a squared-error fit. Large residuals clustering in a region you suspect is noisy is a reason to try Huber. Large residuals that look like real rare events are a reason to investigate the data, not to down-weight the points. Residual inspection is diagnosis, not proof.

Robust Is Not the Same as Correct

Here is the misconception worth killing: a robust loss does not fix unusual observations. It changes how much influence they have. Whether that influence should be reduced is a question about the data, not about the loss function.

If the extreme value is genuine and important — a real spike in demand, a real failure event — down-weighting it is a modeling error dressed up as robustness. You have told the model to ignore the thing you most needed it to learn.

Huber is also not immune to outliers. It still has a residual-gradient in the linear region; that gradient is just capped. Less sensitive is a different claim from insensitive, and the difference matters when many points sit in the tail.

It helps to see Huber as one option among several:

ApproachBehavior on large residualsWhen it fits
Squared errorQuadratic penalty, strong pullClean data; large errors genuinely worse
Absolute errorLinear penalty, constant pullMedian-seeking; heavy-tailed noise
HuberQuadratic then linear at δ\deltaYou believe a subset of large residuals is noise
Target transformChanges the error scale itselfSkewed targets; multiplicative error structure

My rule is simple. Use squared error when large errors are genuinely worse and the data is clean. Consider Huber when you have a concrete reason to believe a subset of large residuals is noise. Investigate the points before choosing either.

One evaluation trap: a model trained with Huber can still be scored with RMSE. The training loss shapes the fit; the evaluation metric answers a separate question about the errors you care about. Do not assume they should match.

Where to Go Next

Before switching losses, inspect the residuals. Ask whether the large ones are noise or signal. That question determines the loss; the loss does not answer it for you.

Then run the experiment. Fit the same data twice — once with squared error, once with Huber — and plot the residuals for both. The places where the two fits disagree are the real output. Those disagreements tell you which points were steering the squared-error fit, and whether you wanted them to.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A team copies δ = 1 from a tutorial even though its target values are on a scale of thousands. Which approach best follows the article's guidance?
Question 1 of 2Scenario Interpretation

Focus: Choose a scale-aware way to tune Huber's threshold using the target scale and the evaluation objective.

A large residual is a verified, important rare demand spike rather than measurement noise. What is the most appropriate conclusion when considering Huber loss?
Question 2 of 2Scenario Interpretation

Focus: Explain why robust loss should not automatically down-weight an unusual but genuine and important observation.

References

  1. scipy.special.huber — SciPy v1.18.0 Manualdocs.scipy.org
  2. 1.5. Stochastic Gradient Descent — scikit-learn 1.9.1 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.