Skip to content
intermediate

Why the Mean, Median, or Class Prevalence Is the Best Constant Baseline

You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the…

Published 2026-10-02Updated 2026-10-0410 min read
Close-up of a car speedometer display reading 76 km/h with a digital dashboard interface.
Close-up of a car speedometer display reading 76 km/h with a digital dashboard interface. Photo by Damir K on Pexels.

You build a baseline, get a number, and have no idea whether that number is the best a constant can do or just a habit copied from a tutorial. Here is the uncomfortable part: the mean, the median, and the class prevalence are not conventions. Each one is the closed-form solution to a specific loss function. Change the loss, and the "obvious" constant changes with it.

This article derives all three results from scratch, then turns each derivation into a baseline decision rule you can actually use.

The Constant Predictor as an Optimization Problem

A constant predictor is a model with no feature dependence. It ignores the input and outputs one value for every example. For regression, that value is a real number. For binary classification, it is a probability.

Let the training sample be y1,y2,…,yny_1, y_2, \dots, y_n. We want to choose a single scalar θ\theta that minimizes the empirical risk:

L(θ)=1n∑i=1nℓ(θ,yi)L(\theta) = \frac{1}{n} \sum_{i=1}^{n} \ell(\theta, y_i)

Here ℓ\ell is the loss function — the penalty for predicting θ\theta when the truth is yiy_i. We will minimize three of them: squared error, absolute error, and binary log loss.

The structural fact that makes this tractable: because the predictor is constant, the optimization collapses to a one-dimensional problem in a single scalar. That is why closed-form solutions exist. A model with features has one parameter per feature; a constant model has one parameter, period.

Three assumptions are baked in, and they matter:

  • The sample is the object being optimized over. We are finding the best constant for this data, not for the population it came from.
  • The loss is the one we declare. A different loss gives a different answer.
  • The result is optimal for that sample and that loss — nothing more.

Note: A baseline must be evaluated under the same loss and the same split as the model it competes with. If you score your baseline with a different metric than your model, the comparison is meaningless.

If you need a refresher on constructing baselines and comparing them under a shared evaluation design, that groundwork is assumed here. This article is about why the constant takes the form it does.

Knowledge check

Check your understanding

Answer this question before you continue.

What does minimizing empirical risk for a constant predictor establish?
Misconception Check

Focus: Distinguish the sample-specific constant optimum from a claim about population performance.

Why the Mean Minimizes Squared Error

Start with squared error. The loss for one prediction is (θ−yi)2(\theta - y_i)^2, so the empirical risk is:

L(θ)=1n∑i=1n(θ−yi)2L(\theta) = \frac{1}{n} \sum_{i=1}^{n} (\theta - y_i)^2

This is a smooth, convex function of θ\theta. Differentiate with respect to θ\theta:

dLdθ=1n∑i=1n2(θ−yi)=2n∑i=1n(θ−yi)\frac{dL}{d\theta} = \frac{1}{n} \sum_{i=1}^{n} 2(\theta - y_i) = \frac{2}{n} \sum_{i=1}^{n} (\theta - y_i)

Set the derivative to zero:

2n∑i=1n(θ−yi)=0\frac{2}{n} \sum_{i=1}^{n} (\theta - y_i) = 0

Divide both sides by 2n\frac{2}{n} and split the sum:

∑i=1nθ−∑i=1nyi=0\sum_{i=1}^{n} \theta - \sum_{i=1}^{n} y_i = 0

Since θ\theta is constant, ∑i=1nθ=nθ\sum_{i=1}^{n} \theta = n\theta:

nθ=∑i=1nyi⟹θ=1n∑i=1nyi=yˉn\theta = \sum_{i=1}^{n} y_i \quad \Longrightarrow \quad \theta = \frac{1}{n} \sum_{i=1}^{n} y_i = \bar{y}

The optimal constant is the sample mean. The second derivative is 2n>0\frac{2}{n} > 0, so this stationary point is the unique minimum. Convexity means there are no local-minimum traps — the mean is the global answer.

Interpretation: the mean is the point of balance. The minimized squared error equals the sample variance, which is why a constant model under MSE absorbs exactly the spread of the data and nothing else.

The failure mode follows directly from the mechanism. Squared error weights large deviations quadratically. A single extreme value drags the mean toward it and inflates the baseline error. This is not a separate rule to memorize — it is what the squared term does.

Connect to practice: when your evaluation metric is RMSE or MSE, the mean is the correct constant baseline. Reporting it tells you how much error a featureless model already absorbs.

Knowledge check

Check your understanding

Answer this question before you continue.

For targets 2, 4, and 10, which constant minimizes the sum of squared errors?
Single Choice

Focus: Apply the squared-error derivation to identify the optimal constant from observed targets.

Why the Median Minimizes Absolute Error

Now switch to absolute error. The loss is ∣θ−yi∣|\theta - y_i|, and the empirical risk is:

L(θ)=1n∑i=1n∣θ−yi∣L(\theta) = \frac{1}{n} \sum_{i=1}^{n} |\theta - y_i|

This function is piecewise linear with kinks at each data value yiy_i. It is not differentiable everywhere, so the derivative argument from the previous section breaks down. We need a different tool.

Use the slope-counting argument. For a candidate value θ\theta that sits between two data points, the slope of LL depends on how many points lie to the left versus the right. Each point to the left contributes a slope of −1-1; each point to the right contributes +1+1. If kk points are below θ\theta and n−kn - k are above, the slope is:

dLdθ=1n(−k+(n−k))=n−2kn\frac{dL}{d\theta} = \frac{1}{n}\bigl(-k + (n - k)\bigr) = \frac{n - 2k}{n}

The minimum sits where this slope crosses zero — that is, where k=n/2k = n/2. The counts balance when half the data lies on each side. That point is a median.

But the slope argument only covers regions between data points. At a data value yjy_j, the function has a kink, and the derivative does not exist. To handle that case, use the one-sided slopes. Just to the left of yjy_j, the slope is n−2kleftn\frac{n - 2k_{\text{left}}}{n}, where kleftk_{\text{left}} counts points strictly below yjy_j. Just to the right, it is n−2krightn\frac{n - 2k_{\text{right}}}{n}, where krightk_{\text{right}} counts points at or below yjy_j. A minimum occurs at yjy_j when the slope is non-positive on the left and non-negative on the right — that is, when neither side holds more than half the observations:

kleft≤n2andkright≥n2k_{\text{left}} \le \frac{n}{2} \quad \text{and} \quad k_{\text{right}} \ge \frac{n}{2}

This condition defines the full set of medians. When nn is odd, exactly one data value satisfies it. When nn is even, an entire interval between the two middle values satisfies it, and every point in that interval is a minimizer. The minimizer is non-unique, and that is a real property of the objective, not a numerical artifact. Repeated values only widen the set of points that satisfy the condition.

The result generalizes. The α\alpha-quantile minimizes the tilted absolute (pinball) loss, with the median as the α=0.5\alpha = 0.5 special case. When over- and under-prediction cost different amounts, the optimal constant shifts toward a quantile other than the median.

Failure mode: absolute error is robust to outliers, but its objective is flat or kinked. That is why it is harder to optimize with gradient methods and why the minimizer can be non-unique.

Connect to practice: when your metric is MAE, the median is the correct constant baseline. The gap between the mean-baseline error and the median-baseline error is a quick diagnostic for how much outliers are distorting your target.

Knowledge check

Check your understanding

Answer this question before you continue.

For sorted targets 1, 4, 8, and 13, which statement describes the constants that minimize total absolute error?
Scenario Interpretation

Focus: Recognize the full set of absolute-error minimizers for an even-sized sample.

Why Class Prevalence Minimizes Binary Log Loss

Classification is where learners most often get confused, because the predicted probability and the predicted label are two different decisions.

Define a constant probabilistic classifier: it outputs one probability pp for every input. Let n+n_+ be the number of positive examples and n−n_- the number of negatives, with n=n++n−n = n_+ + n_-. Binary log loss is the average negative log-likelihood of the observed labels:

L(p)=−1n[n+log⁡p+n−log⁡(1−p)]L(p) = -\frac{1}{n} \left[ n_+ \log p + n_- \log(1 - p) \right]

Differentiate with respect to pp:

dLdp=−1n[n+p−n−1−p]\frac{dL}{dp} = -\frac{1}{n} \left[ \frac{n_+}{p} - \frac{n_-}{1 - p} \right]

Set to zero:

n+p=n−1−p\frac{n_+}{p} = \frac{n_-}{1 - p}

Cross-multiply and solve:

n+(1−p)=n−p⟹n+=p(n++n−)=pn⟹p=n+nn_+(1 - p) = n_- p \quad \Longrightarrow \quad n_+ = p(n_+ + n_-) = pn \quad \Longrightarrow \quad p = \frac{n_+}{n}

The optimal constant probability is the observed positive rate — the class prevalence. It is the same value regardless of which class you call positive.

Now separate two decisions that beginners conflate. The optimal constant probability is prevalence. The optimal constant hard label depends on the decision threshold and the cost of each error type. These are not the same question, and answering one does not answer the other.

There is a calibration connection worth noting, but it is narrower than it first appears. When prevalence is computed on the same sample you are optimizing over, the constant predictor is calibrated on that sample — its predicted probability matches the observed frequency by construction. That is an in-sample property. It does not guarantee calibration on a separate validation set or on future data, where the true positive rate may differ. Treat it as a reference point for checking whether a model's probabilities are trustworthy, not as a promise that survives the split.

Failure mode: with imbalanced classes, prevalence is a small number. Log loss rewards a confident-looking but nearly useless predictor. Always pair the log-loss baseline with a threshold-dependent metric such as accuracy, precision, or recall.

Connect to practice: when your metric is log loss or cross-entropy, the prevalence baseline is the number to compare against. The gap to a real model measures how much signal the features actually carry.

Knowledge check

Check your understanding

Answer this question before you continue.

A training sample has 18 positive labels out of 60 examples. Which constant probability minimizes its binary log loss?
Single Choice

Focus: Calculate the constant probability that minimizes binary log loss from the observed positive rate.

One Objective, Three Answers: A Comparison

Three side-by-side comparisons show a mean marker balancing sample values, a median marker splitting ordered values into two groups, and a positive-class fraction bar representing prevalence.
The loss function determines which summary of the training data is the optimal constant.
Loss functionOptimal constantMetric that makes it the right baseline
Squared errorSample meanMSE, RMSE
Absolute errorSample medianMAE
Binary log lossClass prevalenceLog loss, cross-entropy
Tilted absolute (pinball)α\alpha-quantileQuantile loss

The unifying principle: each optimal constant is the summary statistic that its loss function is designed to be minimized by. The pairing is not arbitrary — it falls out of the derivation.

For asymmetric loss, the optimal constant shifts toward a quantile other than the median. The median is a special case of a more general rule, not the whole story.

One boundary to flag: these results are for a single constant. They say nothing about whether a model with features can do better. That is a separate empirical question answered by validation, not by this derivation.

Turning the Derivation Into a Baseline Decision

Here is the workflow I would use.

Pick the constant baseline that matches your evaluation metric, not the one that is easiest to compute. The derivation tells you which constant is optimal for which loss. Ignoring that pairing means you are not measuring what you think you are measuring.

Three mistakes I see repeatedly:

Computing the mean baseline and then evaluating with MAE. This makes the baseline look worse than it is and inflates the apparent value of your model. The mean is not the MAE minimizer. The median is.

Reporting only the constant baseline's score without the metric's scale. A log-loss of 0.4 means nothing without knowing the prevalence. A MAE of 12 means nothing without knowing the target's range.

Treating the constant baseline as a formality rather than a diagnostic. A large gap between baseline and model on a small dataset is often noise, not signal. The baseline is a measurement instrument, not a checkbox.

A practical check: compute all three constants on the same training split and compare their errors under the matching metrics. This shows you how much the loss choice changes the story.

Common mistake: Fitting the constant on the full dataset and then scoring it on the same data. The constant should be fit on training data and scored on validation data, consistent with the split discipline you use for any other model.

The constant baseline is a solved optimization problem. The mean, median, and prevalence are each the provably optimal constant for one specific loss. Knowing which one applies turns the baseline from a ritual into a measurement instrument.

Your next step: take a dataset you already have. Compute the mean, the median, and the class prevalence on the training split. Score each under its matching metric. Those three numbers are the reference your candidate model should be compared against under the same evaluation design. If your model barely clears them, that is not a failure — it is evidence that the features are carrying little signal, and it is worth more than a higher score you cannot explain.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A team evaluates regression predictions with mean absolute error and wants the best constant baseline for that metric. Which constant should it fit?
Question 1 of 2Comparison Reasoning

Focus: Select the optimal constant baseline for a specified evaluation loss.

A model is fit on a training split and compared on a validation split using MSE. Which baseline procedure follows the article's guidance?
Question 2 of 2Scenario Interpretation

Focus: Apply the article's split and metric consistency guidance when evaluating a constant baseline.

References

  1. Constant predictorsee104.stanford.edu
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.