Skip to content
intermediate

Decompose the Brier Score Into Reliability, Resolution, and Uncertainty

Two forecasters can post the same Brier score and deserve opposite verdicts. One is honest but says almost nothing. The other is sharp but lies about its…

Published 2026-10-02Updated 2026-10-0411 min read
Serene ocean waves under a darkening sky, capturing the mysterious twilight sea.
Serene ocean waves under a darkening sky, capturing the mysterious twilight sea. Photo by Eyüpcan Timur on Pexels.

Two forecasters can post the same Brier score and deserve opposite verdicts. One is honest but says almost nothing. The other is sharp but lies about its probabilities. A single number cannot tell you which one you are holding.

That is why the decomposition exists. The Brier score is the mean squared difference between your predicted probability and the 0/1 outcome, bounded in [0, 1], lower is better. It is a proper scoring rule: you minimize it by reporting your honest probability. But it fuses two different failure modes into one number — probabilities that do not match observed frequencies, and probabilities that never move away from the base rate. The decomposition pulls them apart.

This article works one stated grouped-probability decomposition and interprets its terms. It does not re-derive calibration curves or compare calibration methods. If you already know what a reliability diagram shows and why ranking and calibration are separate questions, you have the prerequisite. The job here is to compute the three terms by hand and read each one as its own diagnostic.

Why One Brier Score Cannot Answer Two Questions

Consider two forecasters on the same set of events.

Forecaster A always predicts the base rate. If 30% of outcomes are positive, A says 0.30 every time. A is perfectly calibrated in the aggregate — its stated probability matches the observed frequency. But A carries zero information. It never separates a high-risk case from a low-risk one.

Forecaster B is aggressive. It says 0.95 when it is confident and 0.05 when it is not, and it ranks cases well. But when it says 0.95, the event happens only 70% of the time. B is informative and badly miscalibrated.

Both can land near the same Brier score. A scores respectably by refusing to commit. B scores badly by committing and being wrong about the magnitude. The raw score treats these as similar. The decomposition does not.

The Brier score mixes two questions: are your probabilities honest? and do your probabilities separate cases? A third quantity — the difficulty of the problem itself — sets the reference both of them work against.

Knowledge check

Check your understanding

Answer this question before you continue.

Two forecasters obtain the same Brier score. One always predicts the base rate, while the other makes sharp predictions that are sometimes badly miscalibrated. What can you conclude from the shared score alone?
Scenario Interpretation

Focus: Explain why an identical Brier score does not establish that two forecasters have the same calibration or ability to separate cases.

Notation and the Grouped Setup

Define the pieces before using them.

You have NN forecasts. Each forecast ii has a predicted probability pi∈[0,1]p_i \in [0, 1] and a realized binary outcome yi∈{0,1}y_i \in \{0, 1\}. The raw score is:

BS=1N∑i=1N(pi−yi)2\text{BS} = \frac{1}{N} \sum_{i=1}^{N} (p_i - y_i)^2

Now group. Bucket the forecasts into KK bins by predicted probability. For bin kk:

  • nkn_k — number of forecasts in the bin
  • fkf_k — the bin's probability value (the probability shared by forecasts in that bin)
  • oko_k — the observed frequency, i.e. (number of 1s in bin kk) / nkn_k

Define the overall base rate oˉ\bar{o} as the mean of all outcomes across all NN forecasts.

Note: The three-term identity is exact only when every forecast inside a bin shares the same probability value. With real-valued probabilities, binning introduces within-bin variation, and the three terms no longer sum exactly to the raw score. Extra within-bin terms appear. This is a known limitation, not a rounding error — I will return to it.

The bin count is a choice, and it changes the answer. More on that below.

Knowledge check

Check your understanding

Answer this question before you continue.

A probability bin contains 8 forecasts, of which 3 have outcome 1. What is the bin's observed frequency, o_k?
Single Choice

Focus: Compute a bin's observed frequency from its number of outcomes and forecast count.

The Three Terms, One at a Time

Three labeled terms combine into the Brier score: uncertainty adds to the baseline, resolution is subtracted, and reliability is added. A worked-example line shows 0.24 minus 0.0317 plus 0.0147 equals 0.223.
Read uncertainty as the base-rate score, then subtract resolution and add the reliability penalty.

Each term gets its own mechanism.

Reliability is the nkn_k-weighted average of (fk−ok)2(f_k - o_k)^2. It measures how far your stated probabilities drift from observed frequencies. If you say 0.80 and the event happens 80% of the time in that bin, the term contributes nothing. Lower is better. A perfectly calibrated forecaster has reliability zero.

Resolution is the nkn_k-weighted average of (ok−oˉ)2(o_k - \bar{o})^2. It measures how far the outcome rate inside each bin moves away from the base rate. If every bin has the same outcome rate as the overall base rate, resolution is zero — the model separates nothing. Higher is better. A forecaster that always predicts the base rate has resolution zero by construction.

Uncertainty is oˉ(1−oˉ)\bar{o}(1 - \bar{o}). It depends only on the outcome prevalence, not on the model at all. It is maximal at a 50/50 base rate and zero when the outcome never or always occurs. No model can change it.

The identity:

BS=Reliability−Resolution+Uncertainty\text{BS} = \text{Reliability} - \text{Resolution} + \text{Uncertainty}

Read the signs carefully, because this is where readers get it wrong. Miscalibration adds error. Informative separation subtracts error. Uncertainty is the score of the uninformative base-rate forecast — the reference point that resolution improves on and reliability penalizes.

That reference framing matters. Uncertainty is not a floor on the score you can achieve. A model that is perfectly calibrated and perfectly separating scores zero, even when the outcome is genuinely uncertain. What uncertainty tells you is how much a base-rate-only forecast would score on this problem. Resolution is the credit you earn by beating that baseline; reliability is the penalty you pay for probabilities that do not match reality.

Knowledge check

Check your understanding

Answer this question before you continue.

Suppose each bin's forecast probability is changed to equal that bin's observed frequency, while the bin ordering and outcomes stay fixed. According to the article's recalibrated example, what changes?
Comparison Reasoning

Focus: Predict how perfect bin-level recalibration affects reliability, resolution, uncertainty, and the Brier score when bin ordering is unchanged.

A Worked Example on Ten Forecasts

Ten forecasts, grouped into three bins.

Binfkf_knkn_k1soko_k
10.20410.25
20.50310.333
30.80320.667

Total outcomes: 1+1+2=41 + 1 + 2 = 4 out of 10, so oˉ=0.4\bar{o} = 0.4.

Raw Brier score. Using the individual forecasts (bin probability applied to each member):

  • Bin 1: four forecasts at 0.20, one outcome 1 and three outcomes 0. Errors: (0.2−1)2=0.64(0.2-1)^2 = 0.64 once, (0.2−0)2=0.04(0.2-0)^2 = 0.04 three times. Sum =0.64+0.12=0.76= 0.64 + 0.12 = 0.76.
  • Bin 2: three forecasts at 0.50, one outcome 1 and two outcomes 0. Errors: (0.5−1)2=0.25(0.5-1)^2 = 0.25 once, (0.5−0)2=0.25(0.5-0)^2 = 0.25 twice. Sum =0.75= 0.75.
  • Bin 3: three forecasts at 0.80, two outcomes 1 and one outcome 0. Errors: (0.8−1)2=0.04(0.8-1)^2 = 0.04 twice, (0.8−0)2=0.64(0.8-0)^2 = 0.64 once. Sum =0.72= 0.72.

Total =0.76+0.75+0.72=2.23= 0.76 + 0.75 + 0.72 = 2.23. Divide by N=10N = 10: BS = 0.223.

Reliability.

REL=110[4(0.20−0.25)2+3(0.50−0.333)2+3(0.80−0.667)2]\text{REL} = \frac{1}{10}\left[4(0.20 - 0.25)^2 + 3(0.50 - 0.333)^2 + 3(0.80 - 0.667)^2\right] =110[4(0.0025)+3(0.0278)+3(0.0178)]= \frac{1}{10}\left[4(0.0025) + 3(0.0278) + 3(0.0178)\right] =110[0.01+0.0833+0.0533]=0.0147= \frac{1}{10}\left[0.01 + 0.0833 + 0.0533\right] = 0.0147

Resolution.

RES=110[4(0.25−0.4)2+3(0.333−0.4)2+3(0.667−0.4)2]\text{RES} = \frac{1}{10}\left[4(0.25 - 0.4)^2 + 3(0.333 - 0.4)^2 + 3(0.667 - 0.4)^2\right] =110[4(0.0225)+3(0.0045)+3(0.0713)]= \frac{1}{10}\left[4(0.0225) + 3(0.0045) + 3(0.0713)\right] =110[0.09+0.0135+0.2139]=0.0317= \frac{1}{10}\left[0.09 + 0.0135 + 0.2139\right] = 0.0317

Uncertainty.

UNC=0.4×0.6=0.24\text{UNC} = 0.4 \times 0.6 = 0.24

Verify the identity.

REL−RES+UNC=0.0147−0.0317+0.24=0.223\text{REL} - \text{RES} + \text{UNC} = 0.0147 - 0.0317 + 0.24 = 0.223

That matches the raw score. The arithmetic closes.

Interpretation. The score is 0.223. A base-rate-only forecast on this problem would score 0.24, so this model beats the baseline by 0.0317 through resolution and gives back 0.0147 to miscalibration. The net gain over the baseline is small but real. Reliability is low, so the probabilities are roughly honest. Resolution is modest, so the bins separate, but not dramatically. The model is doing something, just not much.

A recalibrated variant. Suppose you recalibrate so each bin's stated probability matches its observed frequency: bin 1 becomes 0.25, bin 2 becomes 0.333, bin 3 becomes 0.667. The bin ordering is unchanged, so resolution stays at 0.0317. Reliability drops to zero. The new score is 0−0.0317+0.24=0.2080 - 0.0317 + 0.24 = 0.208. The improvement came entirely from the term you targeted. That is the check you want: move one term, confirm the score moved for the reason you intended.

Knowledge check

Check your understanding

Answer this question before you continue.

In the ten-forecast example, 4 outcomes are positive. What is the uncertainty term?
Single Choice

Focus: Calculate uncertainty from the overall outcome prevalence in the worked example.

What Each Term Is Allowed to Claim

The decomposition is a diagnostic, not a verdict. Each term licenses a specific claim.

Reliability isolates systematic over- or under-confidence. It is the term recalibration is designed to reduce. If reliability is large relative to resolution, your probabilities are drifting from reality in a way a calibration step can address.

Resolution reflects whether the model separates high-risk from low-risk cases. If resolution is near zero, the model is not distinguishing anything — it is a base-rate machine. No amount of recalibration fixes that. You need better features or more model capacity.

Uncertainty is a property of the problem, not the model. This has a sharp practical consequence: comparing raw Brier scores across datasets with different base rates compares difficulty as much as skill. A score of 0.15 on a 50/50 problem and a score of 0.15 on a 5% base-rate problem are not the same achievement.

Common mistake: Assuming a low Brier score means good calibration. It does not. A model can score low by being well calibrated, or by being uninformative on an easy problem. High resolution does not guarantee usable probabilities either — a sharp model can be sharp and wrong.

The components are not independently optimizable. Driving reliability to zero does not automatically improve the score, because recalibration can flatten the probabilities and reduce resolution. Pushing resolution up can raise reliability if the model becomes overconfident. The terms trade against each other. Report the decomposition alongside a calibration view, not instead of it.

Where the Decomposition Breaks Down

The three-term reading is trustworthy only inside its assumptions.

Bin-count sensitivity. Reliability and resolution both change as you change the number of bins. Uncertainty does not. With too few bins, both terms are understated — a single bin forces resolution to zero by definition. With too many bins, each bin holds too few forecasts and the observed frequencies become noisy. The bin count is a modeling choice, and it moves the answer.

Within-bin variation. When forecasts inside a bin differ, the identity is no longer exact. Additional within-bin variance and covariance terms appear, and the clean three-term sum breaks. If you stratify on distinct probability values instead of binning, the within-bin terms vanish and the identity holds exactly.

Small samples per bin. Noisy observed frequencies inflate apparent reliability and resolution. A bin with three forecasts tells you almost nothing about the true conditional frequency.

It describes the past, not the future. The decomposition describes the forecasts you scored. It says nothing about distribution shift or whether the probabilities will stay calibrated on new data.

When not to lean on it. Very small evaluation sets, heavily imbalanced outcomes where uncertainty is near zero, and settings where you only need a ranking rather than a probability. If you never act on the probability magnitude, the decomposition is measuring something you do not use.

A Practical Reading Routine

Turn the decomposition into a repeatable sequence.

  1. Collect aligned pairs. Gather each predicted probability and its realized outcome from a held-out set, not from training data.
  2. Compute the raw Brier score. This is your headline number.
  3. Bin the probabilities. Choose a bin count you can defend, then compute reliability, resolution, and uncertainty.
  4. Read the terms in order. Is the score dominated by miscalibration, by weak separation, or is it close to the base-rate baseline?
  5. Choose the next action from the dominant term. Reliability large relative to resolution: recalibrate. Resolution near zero: revisit features or model capacity. Score near the base-rate baseline: the model is adding little, so the bottleneck is the model, not the problem's difficulty.
  6. Re-run after any change. Confirm the improvement came from the term you intended to move.

My rule is simple: read reliability first, resolution second, and treat uncertainty as the baseline you are trying to beat. Reliability tells you whether you can trust the number. Resolution tells you whether the number is worth trusting. Uncertainty tells you what "worth trusting" costs on this problem.

The same grouped reasoning extends to multi-class problems through the ranked probability score, where the decomposition generalizes and the same three roles reappear. The natural next step is to apply this routine to your own held-out predictions and pair the decomposition with a calibration plot — the plot shows you the shape of the miscalibration, and the reliability term tells you how much it costs.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Two models each have a raw Brier score of 0.15, but one was evaluated on a dataset with a 50% base rate and the other on a dataset with a 5% base rate. What is the most careful interpretation?
Question 1 of 2Comparison Reasoning

Focus: Explain why an equal raw Brier score on datasets with different base rates does not imply equal predictive achievement.

A held-out evaluation shows reliability is large relative to resolution. Following the article's reading routine, which next action best targets the dominant issue?
Question 2 of 2Scenario Interpretation

Focus: Choose a next modeling action based on whether reliability or resolution is the main diagnostic concern.

References

  1. Simplifying and generalising Murphy's Brier score ...ore.exeter.ac.uk
  2. NOTES AND CORRESPONDENCE Two Extra Components in ...www.cptec.inpe.br
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.