Skip to content
intermediate

Quantile Regression: Derive the Pinball-Loss Objective

Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.

Published 2026-10-02Updated 2026-10-049 min read
Professional woman standing confidently in a data center, surrounded by glowing servers.
Professional woman standing confidently in a data center, surrounded by glowing servers. Photo by Christina Morillo on Pexels.

Your model keeps under-predicting the values you actually care about, and averaging the errors away does not fix it.

That is the symptom. The cause is usually a quiet assumption buried in the loss function: squared error decides that the thing worth predicting is the conditional mean. If your real cost is asymmetric — a missed high demand hurts more than an overstocked shelf — then you are optimizing the wrong functional of the distribution. This article derives the objective that fixes it, from the asymmetry up, so you can regenerate the formula instead of memorizing it.

Why the Mean Is the Wrong Target

Picture a supply planner. Under-predicting demand means empty shelves and lost sales. Over-predicting means warehouse cost and markdowns. Those two errors do not cost the same, yet a model trained on squared error treats a +10 miss and a −10 miss as identical. It will happily sit in the middle of the conditional distribution, which is exactly where the expensive errors live.

The mean is a summary, not a universal target. When the conditional distribution is skewed or heavy-tailed, the mean can sit in a region where almost no observation actually lands. You get a number that minimizes average squared distance and answers a question nobody asked.

Here is the reframe that matters: a loss function is not a scoreboard bolted on after training. It is the definition of what the model is trying to estimate. Squared error buys you the conditional mean. Absolute error buys you the conditional median. An asymmetric absolute error buys you a conditional quantile. Change the loss, and you change the target — not just the score.

If you have already worked through least-squares fitting and residuals, you have the machinery. We are changing the objective, not the fitting.

Notation and the Quantile Function

Before the derivation, pin down every symbol.

  • xx: the input vector.
  • yy: the target.
  • y^\hat{y}: the prediction.
  • τ∈(0,1)\tau \in (0,1): the quantile level.
  • F(y∣x)F(y \mid x): the conditional CDF of yy given xx.

The τ\tau-quantile of the conditional distribution is the smallest value qq such that F(q∣x)≥τF(q \mid x) \ge \tau. When the CDF is continuous and strictly increasing, this is simply the value where F(q∣x)=τF(q \mid x) = \tau. When the CDF has jumps — a discrete or degenerate distribution — the quantile is a set, and we will come back to that.

The distinction that trips people up: an unconditional quantile is one number for the whole sample. A conditional quantile is a function of xx. That function is the object a quantile regression model estimates. Two different inputs can have two different 90th percentiles, and the model is supposed to capture that.

One assumption to state plainly: the derivation below is a population statement about the true conditional distribution. It says what the loss is minimized by in the limit of infinite data. It is not a guarantee about a finite fitted model.

Knowledge check

Check your understanding

Answer this question before you continue.

Which statement best describes what a conditional 90th-quantile model estimates?
Comparison Reasoning

Focus: Distinguish a conditional quantile function from a single unconditional quantile.

Building the Pinball Loss From the Asymmetry

Start from absolute error, ∣y−y^∣|y - \hat{y}|, and ask what happens if you weight the two sides differently. Weight under-prediction by τ\tau and over-prediction by 1−τ1 - \tau:

Lτ(y,y^)={τ(y−y^),y≥y^(1−τ)(y^−y),y<y^L_\tau(y, \hat{y}) = \begin{cases} \tau (y - \hat{y}), & y \ge \hat{y} \\ (1 - \tau)(\hat{y} - y), & y < \hat{y} \end{cases}

The compact form is the same object:

Lτ(y,y^)=(y−y^)(τ−I(y<y^))L_\tau(y, \hat{y}) = (y - \hat{y})\bigl(\tau - \mathbb{I}(y < \hat{y})\bigr)

Check the boundary case. At τ=0.5\tau = 0.5, both branches carry weight 0.50.5, so the loss is half the absolute error and the minimizer is the median. The median is not a separate method; it is the special case.

Geometrically, this is a kinked, asymmetric V. The kink sits at y^=y\hat{y} = y, and the two arms have different slopes. That asymmetry is the whole mechanism: the optimizer slides toward whichever side is penalized more steeply until the two forces balance.

You will see this called the pinball loss, the quantile loss, and the check function. Same object, different names across papers and libraries.

Knowledge check

Check your understanding

Answer this question before you continue.

At quantile level τ = 0.9, how are the two kinds of error weighted in the article's pinball loss?
Misconception Check

Focus: Identify how the quantile level sets the relative pinball-loss weights on under- and over-predictions.

Deriving the Minimizer: Why the Optimum Is the Quantile

A conditional CDF curve crosses a horizontal line at quantile level tau. To the left of the crossing, the derivative F(y-hat given x) minus tau is negative and an arrow points right; to the right, it is positive and an arrow points left. The crossing marks the quantile.
The expected-loss gradient points toward the prediction where the conditional CDF reaches τ.

Now the actual derivation. Fix xx and consider the expected loss over yy:

E[Lτ(y,y^)∣x]=τ∫y^∞(y−y^) f(y∣x) dy+(1−τ)∫−∞y^(y^−y) f(y∣x) dy\mathbb{E}\bigl[L_\tau(y, \hat{y}) \mid x\bigr] = \tau \int_{\hat{y}}^{\infty} (y - \hat{y})\, f(y \mid x)\, dy + (1 - \tau) \int_{-\infty}^{\hat{y}} (\hat{y} - y)\, f(y \mid x)\, dy

Differentiate with respect to y^\hat{y}. The first integral contributes −τ(1−F(y^∣x))-\tau(1 - F(\hat{y} \mid x)); the second contributes (1−τ)F(y^∣x)(1 - \tau) F(\hat{y} \mid x). Collect:

∂∂y^E[Lτ∣x]=(1−τ)F(y^∣x)−τ(1−F(y^∣x))=F(y^∣x)−τ\frac{\partial}{\partial \hat{y}} \mathbb{E}\bigl[L_\tau \mid x\bigr] = (1 - \tau) F(\hat{y} \mid x) - \tau\bigl(1 - F(\hat{y} \mid x)\bigr) = F(\hat{y} \mid x) - \tau

Set it to zero:

F(y^∣x)=τF(\hat{y} \mid x) = \tau

The population minimizer is the value where the conditional CDF crosses τ\tau — the τ\tau-quantile.

Read the derivative's sign on either side and the behavior becomes obvious. Below the quantile, F(y^∣x)<τF(\hat{y} \mid x) < \tau, so the gradient is negative and pushes the prediction up. Above it, the gradient is positive and pushes down. The equilibrium is exactly where the CDF crosses τ\tau.

In plain causal language: when τ>0.5\tau > 0.5, under-prediction is steeper, so the optimizer drifts upward until roughly τ\tau of the probability mass sits below the prediction. That is the asymmetry doing its job.

What breaks when assumptions fail? If the conditional distribution is discontinuous or degenerate, FF jumps over τ\tau and the minimizer is a set of values rather than a unique point. Finite-sample fits inherit that ambiguity — you may see the loss surface flat over a range.

Knowledge check

Check your understanding

Answer this question before you continue.

For a continuous, strictly increasing conditional CDF, what condition characterizes the population prediction that minimizes expected pinball loss at level τ?
Single Choice

Focus: Connect the expected pinball-loss derivative to the conditional quantile in the continuous, strictly increasing case.

A Worked Example You Can Check by Hand

Take a tiny fixed sample: y=[1,2,3,4,10]y = [1, 2, 3, 4, 10]. Try τ=0.9\tau = 0.9 and evaluate the pinball loss at a few candidate predictions. For each candidate, split the points into under-predictions (y≥y^y \ge \hat{y}, weighted by τ\tau) and over-predictions (y<y^y < \hat{y}, weighted by 1−τ1 - \tau).

y^\hat{y}Under-predictionsOver-predictionsLoss at τ=0.9\tau = 0.9
30.9(1−3)+0.9(2−3)+0.9(3−3)+0.9(4−3)+0.9(10−3)0.9(1-3) + 0.9(2-3) + 0.9(3-3) + 0.9(4-3) + 0.9(10-3)—0.9(0+1+0+1+7)=8.10.9(0+1+0+1+7) = 8.1
40.9(4−4)+0.9(10−4)0.9(4-4) + 0.9(10-4)0.1(4−1)+0.1(4−2)+0.1(4−3)0.1(4-1) + 0.1(4-2) + 0.1(4-3)0.9(0+6)+0.1(3+2+1)=6.00.9(0+6) + 0.1(3+2+1) = 6.0
100.9(10−10)0.9(10-10)0.1(10−1)+0.1(10−2)+0.1(10−3)+0.1(10−4)0.1(10-1) + 0.1(10-2) + 0.1(10-3) + 0.1(10-4)0.9(0)+0.1(9+8+7+6)=3.00.9(0) + 0.1(9+8+7+6) = 3.0

The minimum lands on the sample 90th quantile, which for five points is the largest value, 10. The asymmetry is visible in the table: at y^=3\hat{y} = 3 every point is an under-prediction and the steep τ\tau weight dominates; at y^=10\hat{y} = 10 every point is an over-prediction and the shallow 1−τ1 - \tau weight keeps the total small. Now repeat at τ=0.5\tau = 0.5 and the minimum lands on the median, 3 — the MAE special case.

You can verify the sweep in a few lines:

import numpy as np

y = np.array([1, 2, 3, 4, 10])
candidates = np.arange(0, 12, 0.5)

def pinball(y, yhat, tau):
    return np.mean(np.where(y >= yhat, tau * (y - yhat),
                            (1 - tau) * (yhat - y)))

for tau in (0.5, 0.9):
    losses = [pinball(y, c, tau) for c in candidates]
    print(tau, candidates[int(np.argmin(losses))])

The output confirms the algebra: the loss curve is asymmetric, and its lowest point is not the mean.

Reading the Quantile Level as a Lever

τ\tau is not a confidence level. It is the fraction of the conditional distribution expected to fall at or below the prediction. At τ=0.9\tau = 0.9, you are saying: about 90% of the mass sits below this value.

Fit several τ\tau values and you trace out the shape of the conditional distribution. Spread, skew, and heteroscedasticity become visible instead of being averaged away. If the conditional spread widens with xx, a single mean model hides it; a fan of quantile models exposes it.

Choose τ\tau from the cost ratio between over- and under-prediction, not from a default. If a miss high costs three times a miss low, the ratio points you toward a τ\tau above 0.5.

Knowledge check

Check your understanding

Answer this question before you continue.

A model predicts the conditional 90th quantile for a particular input x. What does that level mean in the article's interpretation?
Scenario Interpretation

Focus: Interpret a quantile level as a conditional probability-mass target rather than an interval confidence level.

What Quantile Prediction Does Not Guarantee

This is where the expensive mistake lives. Stack a 0.1 model and a 0.9 model and you get an interval-shaped output. Nothing in the pinball objective enforces 80% coverage.

Coverage is a property of the joint behavior of both quantile estimates plus the data distribution. Each quantile can be individually well-fit and the pair can still miscover. Worse, independently fitted quantiles can invert — the 0.1 prediction lands above the 0.9 prediction, producing a negative-width interval. That is quantile crossing, and it is a symptom of fitting each level in isolation.

The population result is asymptotic. A flexible model can fit the training quantiles without generalizing the coverage.

Common mistake: reading a quantile pair as a calibrated interval. Measure empirical coverage on held-out data, and treat calibration as a separate step rather than a byproduct of the loss.

When to Reach for Quantile Regression

Use it when the cost of error is asymmetric, when the conditional spread matters, or when the mean is a poor summary of a skewed target. Skip it when a symmetric loss matches the real cost, when you need a single point estimate, or when you need guaranteed coverage and have not budgeted for calibration.

There is a robustness angle worth noting: because the loss grows linearly rather than quadratically, extreme targets exert less pull on the fit than they would under squared error. That is a feature when outliers are noise and a liability when they are signal.

This article derived the objective. The estimator, solver, and tuning choices are a separate topic.

The One-Line Result, and What to Do With It

The pinball loss is minimized by the conditional quantile. That is the whole derivation compressed into a sentence.

Turn it into a decision rule: choose τ\tau from the cost asymmetry, fit several quantiles to see the conditional spread, and never read a quantile pair as a calibrated interval without measuring coverage.

Here is your next move. Pick one skewed target you already have. Fit a median model and a high-quantile model, then compute empirical coverage on held-out data against the nominal level. The gap you see is the difference between predicting a quantile and owning an interval — and it is the gap that separates a model that looks right from one that survives contact with a real decision.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A team stacks predictions from separately fitted 0.1 and 0.9 quantile models and calls the result an 80% interval. Which conclusion is supported by the article?
Question 1 of 2Misconception Check

Focus: Explain why fitting lower and upper conditional quantiles does not by itself guarantee nominal interval coverage.

For a demand-planning decision, a miss that under-predicts demand costs three times as much as a miss that over-predicts it. Which quantile-level choice follows the article's cost-asymmetry guidance?
Question 2 of 2Scenario Interpretation

Focus: Choose the direction of the quantile level implied by asymmetric under- and over-prediction costs.

References

  1. 3.4. Metrics and scoring: quantifying the quality of predictions — scikit-learn 1.9.1 documentationscikit-learn.org
  2. The Earth Mover's Pinball Loss: Quantiles for Histogram ...proceedings.mlr.press
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.