Skip to content
intermediate

Derive Softmax Probabilities and the Multiclass Cross-Entropy Objective

A multiclass model hands you K raw scores. The obvious move — divide each by their sum — collapses the first time a score goes negative. Here is the fix,…

Published 2026-10-02Updated 2026-10-049 min read
A human brain model placed on a blue plate, viewed from above against a pastel background.
A human brain model placed on a blue plate, viewed from above against a pastel background. Photo by Amel Uzunovic on Pexels.

A multiclass model hands you K raw scores. The obvious move — divide each by their sum — collapses the first time a score goes negative. Here is the fix, and where the loss function actually comes from.

From Binary Logits to K Class Scores

You already know the binary case. A linear model produces one score, the sigmoid squashes it into a probability, and the complement gives you the other class. One score, one probability, one decision.

Multiclass problems break that symmetry. With K classes, there is no "the other class" to fall back on. So we give every class its own parameters. For class jj, with weight vector wjw_j and bias bjb_j:

zj=wj⊤x+bjz_j = w_j^\top x + b_j

Stack the results into a vector z=(z1,…,zK)z = (z_1, \dots, z_K). These are logits — raw, unnormalized scores. They can be negative, they can be huge, and they do not sum to anything meaningful. A logit of 4.24.2 does not mean "42% likely." It means nothing on its own.

Our goal is a map from z∈RKz \in \mathbb{R}^K to a probability vector pp where every pj>0p_j > 0 and ∑jpj=1\sum_j p_j = 1. Two properties matter: the map should be monotonic (a larger score must produce a larger probability) and differentiable (so gradient-based optimization can find the weights).

Knowledge check

Check your understanding

Answer this question before you continue.

A model assigns class 1 a logit of 4.2. What can you conclude from that value alone?
Misconception Check

Focus: Interpret a multiclass model's class scores as raw logits rather than probabilities.

Why Naive Normalization Fails

The tempting candidate is pj=zj/∑kzkp_j = z_j / \sum_k z_k. Try it on z=(2.0,1.0,−1.0)z = (2.0, 1.0, -1.0). The denominator is 2.02.0, and you get p=(1.0,0.5,−0.5)p = (1.0, 0.5, -0.5). That is not a probability distribution. It has a negative entry, and it sums to 1.01.0 only by accident.

Two requirements are violated at once: non-negativity, and a denominator that is guaranteed positive. If all scores are negative, the denominator flips sign and every "probability" inverts.

We need a transformation that is always positive and strictly increasing. The exponential is the smallest such change: exp⁡(zj)>0\exp(z_j) > 0 for every real zjz_j, and it preserves order. Apply it, then normalize.

Knowledge check

Check your understanding

Answer this question before you continue.

For scores (2.0, 1.0, -1.0), what is the decisive problem with dividing each score by their sum?
Scenario Interpretation

Focus: Recognize why dividing raw scores by their sum does not generally produce probabilities.

Deriving the Softmax Function

Two steps, each fixing one broken requirement.

Step 1 — guarantee positivity. Replace each score with ezje^{z_j}. Every term is now strictly positive.

Step 2 — guarantee the sum is 1. Divide each term by the total:

pj=exp⁡(zj)∑k=1Kexp⁡(zk)p_j = \frac{\exp(z_j)}{\sum_{k=1}^{K} \exp(z_k)}

That is the softmax function. The name comes from the fact that it is a smooth approximation of argmax: push one score far above the rest and its probability approaches 1 while the others approach 0.

Check the two properties. Each pjp_j is a positive number divided by a sum that is larger than it, so 0<pj<10 < p_j < 1. And the numerator of the whole vector sums to exactly the denominator, so ∑jpj=1\sum_j p_j = 1. Both requirements hold.

Note: With K=2K = 2, softmax reduces to the sigmoid. Let z=(z1,z2)z = (z_1, z_2) and factor out ez1e^{z_1}: you get p1=1/(1+e−(z1−z2))p_1 = 1/(1 + e^{-(z_1 - z_2)}), which is the sigmoid applied to the score difference. That is the check that softmax is a genuine generalization, not a separate invention.

One practical detail. Exponentials overflow fast — e1000e^{1000} is not a number your machine will store. The standard fix is to subtract the maximum score before exponentiating:

pj=exp⁡(zj−max⁡kzk)∑kexp⁡(zk−max⁡kzk)p_j = \frac{\exp(z_j - \max_k z_k)}{\sum_k \exp(z_k - \max_k z_k)}

This changes nothing mathematically, because the same factor e−max⁡kzke^{-\max_k z_k} cancels from numerator and denominator. It just keeps the largest exponent at e0=1e^0 = 1.

That cancellation is not a coincidence. It reveals that softmax is shift-invariant: add any constant cc to every score and the output is identical. Only score differences carry information. This redundancy is a real modeling fact, not a curiosity — it is why one class's weights can be treated as a reference, and why the parameters are not uniquely identified without regularization.

Knowledge check

Check your understanding

Answer this question before you continue.

If the same constant is added to every class score before applying softmax, what happens to the probabilities?
Comparison Reasoning

Focus: Explain why adding a shared constant to every logit leaves softmax probabilities unchanged.

From Scores to a Likelihood

So far softmax is just a mapping. To get a loss function with a principled origin, we need a probabilistic model.

Assume that given input xx, the label yy is drawn from a categorical distribution with probabilities p=softmax(z)p = \text{softmax}(z). Encode the true label as a one-hot vector yy, where yj=1y_j = 1 for the correct class and 00 elsewhere.

The probability of observing the correct label is then:

P(y∣x)=∏j=1Kpj yjP(y \mid x) = \prod_{j=1}^{K} p_j^{\,y_j}

Every factor where yj=0y_j = 0 contributes 11. The single factor where yj=1y_j = 1 selects the predicted probability of the true class. The one-hot encoding turns a product over all classes into a lookup.

For a dataset of NN examples, the likelihood is the product of per-example likelihoods:

P(Y∣X)=∏i=1NP(y(i)∣x(i))P(Y \mid X) = \prod_{i=1}^{N} P(y^{(i)} \mid x^{(i)})

This factorization assumes the examples are conditionally independent given their features. That assumption is about the dataset, not about the categorical model itself. When examples are genuinely dependent — successive tokens in a sentence, consecutive frames in a video — you cannot factor the joint likelihood this way, and you need a model that represents the dependence. The per-example categorical loss still shows up inside those models; it just gets combined with a sequence structure rather than multiplied naively across independent rows.

We want to maximize this quantity — the parameters that make the observed labels most probable are the ones we keep. But products of many small probabilities underflow to zero, and products are awkward to differentiate. So we take the log next.

Knowledge check

Check your understanding

Answer this question before you continue.

For one example with true class 2 encoded by one-hot vector y = (0, 1, 0), what does ∏ⱼ pⱼʸʲ evaluate to?
Single Choice

Focus: Use one-hot encoding to identify the true-class probability in the categorical likelihood.

Deriving the Cross-Entropy Objective

The log turns the product into a sum:

log⁡P(Y∣X)=∑i=1Nlog⁡P(y(i)∣x(i))=∑i=1N∑j=1Kyj(i)log⁡pj(i)\log P(Y \mid X) = \sum_{i=1}^{N} \log P(y^{(i)} \mid x^{(i)}) = \sum_{i=1}^{N} \sum_{j=1}^{K} y_j^{(i)} \log p_j^{(i)}

The inner sum over classes collapses again, because only one yjy_j is nonzero. For each example, you are left with a single term: the log-probability of the true class.

Maximizing a sum is the same as minimizing its negative. Flip the sign:

L=−∑i=1Nlog⁡py(i)(i)L = -\sum_{i=1}^{N} \log p_{y^{(i)}}^{(i)}

This is the negative log-likelihood, and it is also the cross-entropy between the one-hot target distribution and the predicted distribution. Averaging over examples gives the per-example loss that scikit-learn minimizes for multinomial logistic regression.

Read the loss value directly. It is the negative log-probability assigned to the correct class. Assign probability 1.01.0 to the right answer and the loss is 00. Assign probability 0.010.01 and the loss is about 4.64.6. As that probability approaches zero, the loss grows without bound. The objective punishes confident wrong answers hardest, which is exactly what you want.

A Worked Example by Hand

Three class scores, 2.0, 1.0, and 0.1, flow through exponentiation and normalization to probabilities 0.659, 0.242, and 0.099. Class 1 is marked as the true class, yielding a cross-entropy loss of about 0.417.
Follow the score-to-probability transformation, then see how the true class probability determines the loss.

Take three classes and the score vector z=(2.0,1.0,0.1)z = (2.0, 1.0, 0.1).

ClassScore zjz_jexp⁡(zj)\exp(z_j)Probability pjp_j
12.07.3890.659
21.02.7180.242
30.11.1050.099

The denominator is 7.389+2.718+1.105=11.2127.389 + 2.718 + 1.105 = 11.212. Dividing each exponential by it gives p=(0.659,0.242,0.099)p = (0.659, 0.242, 0.099), which sums to 1.01.0.

Suppose the true class is class 1. The loss is −log⁡(0.659)≈0.417-\log(0.659) \approx 0.417.

Now make the model confidently wrong. Use z=(−1.0,3.0,0.5)z = (-1.0, 3.0, 0.5), where class 2 dominates but class 1 is the truth.

ClassScore zjz_jexp⁡(zj)\exp(z_j)Probability pjp_j
1-1.00.3680.026
23.020.0860.943
30.51.6490.077

The loss is −log⁡(0.026)≈3.65-\log(0.026) \approx 3.65 — nearly nine times larger. The model assigned almost no probability to the correct answer, and the objective says so loudly.

Finally, verify shift invariance. Subtract the max score (2.02.0) from the first example: z′=(0.0,−1.0,−1.9)z' = (0.0, -1.0, -1.9). The exponentials become (1.0,0.368,0.150)(1.0, 0.368, 0.150), the denominator becomes 1.5181.518, and the probabilities are (0.659,0.242,0.099)(0.659, 0.242, 0.099) — identical. The absolute magnitude of the scores never mattered. Only their gaps did.

You can reproduce all of this in a few lines:

import numpy as np

def softmax(z):
    z = z - np.max(z)          # stability, not correctness
    e = np.exp(z)
    return e / e.sum()

def cross_entropy(z, true_class):
    return -np.log(softmax(z)[true_class])

Assumptions, Boundaries, and Common Misreadings

The derivation rests on one assumption worth naming: exactly one correct class per example. That is what makes the one-hot encoding valid and the sum collapse to a single term. It is also why this objective does not fit multilabel problems, where several labels can be true at once. For those, you use independent sigmoid outputs and a sum of binary cross-entropies instead.

A few more boundaries:

  • Softmax outputs are model-implied probabilities, not calibrated ones. A pjp_j of 0.90.9 does not mean the model is right 90% of the time. Calibration is a separate concern, and it often needs its own correction step.
  • Probabilities are relative. Because only score differences matter, a single logit has no standalone meaning. Two models with completely different score scales can produce identical probabilities.
  • The parameterization is redundant. The shift-invariance means the weights are not uniquely identified unless you fix a reference class or apply regularization.
  • One-vs-rest is a different model. There, you train K independent binary classifiers and the outputs are not forced to sum to 1. That can be an advantage when classes overlap, and a liability when you need a coherent distribution.

Common mistake: Treating softmax as a classifier. It is a normalization function. The decision comes from argmax, and the loss comes from the cross-entropy you just derived. Keep those three jobs separate in your head.

This same softmax-plus-cross-entropy pairing reappears in neural network output layers, so the derivation is not a dead end. It is the last layer of nearly every multiclass network you will build.

Where to Go Next

The derivation is short enough to hold in your head: softmax is the exponential-then-normalize answer to a constrained mapping problem, and cross-entropy is just the negative log-likelihood of a categorical model. Two requirements, two steps, one loss.

Do not take that on faith. Implement softmax and cross_entropy in NumPy, feed them the score vectors from the worked example, and confirm the numbers match by hand. Then fit a multinomial logistic regression model on a small dataset, pull out the coefficients, and compute the scores yourself. If your manual softmax matches the model's predict_proba, you understand the mechanism. If it does not, you have found the exact assumption you were carrying wrong — and that is the fastest way to learn it.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Why does minimizing the article's cross-entropy objective favor parameters that make observed labels more probable?
Question 1 of 2Comparison Reasoning

Focus: Connect minimizing categorical cross-entropy with maximizing the likelihood of observed labels.

A task allows several labels to be true for one example. Which conclusion follows from the article's stated assumptions?
Question 2 of 2Scenario Interpretation

Focus: Identify when the one-hot categorical cross-entropy setup is inappropriate because multiple labels may be true.

References

  1. 4.1. Softmax Regression — Dive into Deep Learning 1.0.3 documentationd2l.ai
  2. Unsupervised Feature Learning and Deep Learning Tutorialdeeplearning.stanford.edu
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.