Skip to content
intermediate

Naive Bayes: Derive the Classification Score From Bayes' Rule

The intuition article told you Naive Bayes "combines prior and evidence." Then you tried to write $P(x_1, x_2, \dots, x_n \mid y)$ for real data, and the…

Published 2026-10-02Updated 2026-10-0411 min read
A woman diving underwater in a clear blue swimming pool, captured from above.
A woman diving underwater in a clear blue swimming pool, captured from above. Photo by Demetra Ioannidou on Pexels.

The intuition article told you Naive Bayes "combines prior and evidence." Then you tried to write P(x1,x2,…,xn∣y)P(x_1, x_2, \dots, x_n \mid y) for real data, and the table exploded. This article is the escape hatch: start from Bayes' rule, watch the joint distribution become unmanageable, then let one assumption collapse it into a product of small one-dimensional pieces — and prove the result is a linear score in disguise.

Set Up the Notation Before the Algebra

Before any derivation, pin down the symbols. A derivation you can't read is just a formula wall.

  • yy is the class variable. It takes one of KK values, written C1,…,CKC_1, \dots, C_K (or just yy when the context is clear).
  • x=(x1,x2,…,xn)\mathbf{x} = (x_1, x_2, \dots, x_n) is the feature vector for one instance. Each xix_i is a single feature — a word count, a pixel, a measurement.
  • Our goal: for a new instance x\mathbf{x}, compute P(y∣x)P(y \mid \mathbf{x}) for every candidate class, then pick the class with the largest value.

Two quantities must be estimated from training data:

  1. The prior P(y)P(y) — how common each class is before you see any features.
  2. The class-conditional P(x∣y)P(\mathbf{x} \mid y) — how likely a feature vector is, given the class.

The prior is easy. The class-conditional is where the trouble starts. If you have nn binary features, the full joint P(x1,…,xn∣y)P(x_1, \dots, x_n \mid y) needs a separate parameter for every one of the 2n2^n feature combinations. Twenty features means over a million parameters per class. Real datasets don't have that many rows, so most combinations never appear, and the estimate is mostly zeros.

If you've read the intuition piece on how Naive Bayes combines prior and evidence, you already know the "naive" label refers to an independence assumption. Here we do the actual algebra of that assumption and see exactly where it enters.

Start From Bayes' Rule and Watch the Joint Explode

Bayes' rule gives the posterior directly:

P(y∣x)=P(y) P(x∣y)P(x)P(y \mid \mathbf{x}) = \frac{P(y)\, P(\mathbf{x} \mid y)}{P(\mathbf{x})}

Each term has a job. P(y)P(y) is the prior. P(x∣y)P(\mathbf{x} \mid y) is the likelihood — the class-conditional. P(x)P(\mathbf{x}) is the evidence: the probability of seeing this feature vector at all, regardless of class.

The evidence term expands as a sum over classes:

P(x)=∑y′P(y′) P(x∣y′)P(\mathbf{x}) = \sum_{y'} P(y')\, P(\mathbf{x} \mid y')

Notice what that means. P(x)P(\mathbf{x}) does not depend on yy. It is the same number for every candidate class. That is the first hint it will not affect the decision — and it's the reason the denominator eventually drops out entirely.

Now the bottleneck. For nn binary features, the naive joint P(x∣y)P(\mathbf{x} \mid y) requires on the order of 2(2n−1)2(2^n - 1) parameters to characterize. Under the independence assumption, you need only 2n2n. The gap is not a constant factor; it is exponential versus linear. At n=20n = 20, that's roughly two million parameters versus forty.

With finite data, the exponential version fails in a specific way: most feature combinations never appear in the training set, so their estimated probability is zero. A single zero in the joint wipes out the entire product. The classifier becomes useless — not because the math is wrong, but because the estimate has no data to stand on.

The assumption that rescues us is stated formally as:

P(xi∣y, x1,…,xi−1,xi+1,…,xn)=P(xi∣y)for all iP(x_i \mid y,\, x_1, \dots, x_{i-1}, x_{i+1}, \dots, x_n) = P(x_i \mid y) \quad \text{for all } i

In words: once you know the class, knowing every other feature tells you nothing extra about feature xix_i. That is the conditional independence assumption. It is almost always false in real data. We'll come back to why the classifier still works anyway.

Knowledge check

Check your understanding

Answer this question before you continue.

When using Bayes' rule to choose the most probable class for a fixed feature vector, why can the evidence term P(x) be omitted from the class comparison?
Single Choice

Focus: Explain why the evidence term can be omitted when comparing candidate classes.

Factor the Likelihood With Conditional Independence

A three-stage derivation: Bayes’ rule uses a joint feature likelihood; conditional independence factors it into a product of per-feature likelihoods; taking logs and comparing classes gives the prior log-probability plus a sum of feature log-probabilities.
Conditional independence makes the likelihood factorize; logs turn those factors into additive class scores.

Here is where the assumption does its work. Start with two features and apply the chain rule — a general property of probability that holds regardless of any assumption:

P(x1,x2∣y)=P(x1∣x2,y) P(x2∣y)P(x_1, x_2 \mid y) = P(x_1 \mid x_2, y)\, P(x_2 \mid y)

Now substitute the independence assumption. The term P(x1∣x2,y)P(x_1 \mid x_2, y) simplifies to P(x1∣y)P(x_1 \mid y):

P(x1,x2∣y)=P(x1∣y) P(x2∣y)P(x_1, x_2 \mid y) = P(x_1 \mid y)\, P(x_2 \mid y)

Generalize to nn features:

P(x1,…,xn∣y)=∏i=1nP(xi∣y)P(x_1, \dots, x_n \mid y) = \prod_{i=1}^{n} P(x_i \mid y)

Substitute this back into Bayes' rule to get the fundamental Naive Bayes equation:

P(y∣x)=P(y)∏i=1nP(xi∣y)P(x)P(y \mid \mathbf{x}) = \frac{P(y) \prod_{i=1}^{n} P(x_i \mid y)}{P(\mathbf{x})}

For classification, we only need the most probable class. Since P(x)P(\mathbf{x}) is constant across classes, it drops out of the comparison:

y^=arg⁡max⁡y  P(y)∏i=1nP(xi∣y)\hat{y} = \arg\max_y \; P(y) \prod_{i=1}^{n} P(x_i \mid y)

This is the classification rule. Each P(xi∣y)P(x_i \mid y) is a one-dimensional distribution estimated independently from counts. That decoupling is why Naive Bayes tolerates small datasets: you are not estimating a joint over nn dimensions, you are estimating nn separate one-dimensional distributions.

The distribution you choose for P(xi∣y)P(x_i \mid y) defines the model variant. Gaussian for continuous features, multinomial for counts, Bernoulli for binary presence/absence, categorical for discrete values. The derivation above is the same for all of them; only the per-feature distribution changes.

Knowledge check

Check your understanding

Answer this question before you continue.

For two features, which step specifically uses the Naive Bayes conditional-independence assumption?
Misconception Check

Focus: Distinguish the chain rule from the conditional-independence assumption used to factor the likelihood.

Why We Take Logs: From Products to Sums

The product form has a practical problem. Multiply enough small probabilities together and the result underflows to zero in floating point. With a few hundred features, every class scores zero and the argmax is meaningless.

The fix is to take the logarithm. Since log⁡\log is monotonically increasing, the argmax is unchanged:

y^=arg⁡max⁡y[log⁡P(y)+∑i=1nlog⁡P(xi∣y)]\hat{y} = \arg\max_y \left[ \log P(y) + \sum_{i=1}^{n} \log P(x_i \mid y) \right]

The product became a sum. That single transformation gives you two things.

Numerical stability. Summing log-probabilities keeps the score in a usable range instead of collapsing to zero. Log-probabilities are measured in nats — units of information — and they add cleanly.

Structural clarity. The score is now a weighted sum of per-feature contributions plus a prior term. That is exactly the form of a linear classifier in log space. Each feature votes, the votes add up, and the class with the highest total wins. This is why Naive Bayes is fast, interpretable, and easy to inspect: you can read off which feature contributed most to a decision.

Common mistake: Dropping the prior term because "it's just a constant." The prior is constant across features, not across classes. It shifts the decision boundary and matters enormously when classes are imbalanced. A spam filter trained on 99% ham will predict ham for almost everything if you ignore the prior.

Warning: The log score is not a calibrated probability. It ranks classes correctly more often than it estimates their true likelihood. Do not treat the raw score as "the probability this email is spam."

Knowledge check

Check your understanding

Answer this question before you continue.

What happens to the Naive Bayes class comparison when the product score is written in log form?
Comparison Reasoning

Focus: Explain why taking logarithms produces an additive score without changing the winning class.

Work a Small Example by Hand

Let's classify whether a message is spam based on two binary features: contains "free" (x1x_1) and contains "meeting" (x2x_2). Training data:

Messagex1x_1 (free)x2x_2 (meeting)Class
110spam
211spam
300ham
401ham
501ham

Priors. Spam: 2/5=0.42/5 = 0.4. Ham: 3/5=0.63/5 = 0.6.

Class-conditionals with Laplace smoothing. Without smoothing, P(x1=0∣spam)=0P(x_1 = 0 \mid \text{spam}) = 0, and any message without "free" would score zero for spam regardless of other evidence. Add 1 to every count and 2 to every denominator (two possible values per feature):

P(x1=1∣spam)=2+12+2=0.75,P(x1=0∣spam)=0+12+2=0.25P(x_1 = 1 \mid \text{spam}) = \frac{2+1}{2+2} = 0.75, \quad P(x_1 = 0 \mid \text{spam}) = \frac{0+1}{2+2} = 0.25 P(x2=1∣spam)=1+12+2=0.5,P(x2=0∣spam)=0.5P(x_2 = 1 \mid \text{spam}) = \frac{1+1}{2+2} = 0.5, \quad P(x_2 = 0 \mid \text{spam}) = 0.5 P(x1=1∣ham)=0+13+2=0.2,P(x1=0∣ham)=3+13+2=0.8P(x_1 = 1 \mid \text{ham}) = \frac{0+1}{3+2} = 0.2, \quad P(x_1 = 0 \mid \text{ham}) = \frac{3+1}{3+2} = 0.8 P(x2=1∣ham)=2+13+2=0.6,P(x2=0∣ham)=1+13+2=0.4P(x_2 = 1 \mid \text{ham}) = \frac{2+1}{3+2} = 0.6, \quad P(x_2 = 0 \mid \text{ham}) = \frac{1+1}{3+2} = 0.4

Score a new message with x1=1x_1 = 1 ("free") and x2=1x_2 = 1 ("meeting").

Raw probability form:

spam=0.4×0.75×0.5=0.15\text{spam} = 0.4 \times 0.75 \times 0.5 = 0.15 ham=0.6×0.2×0.6=0.072\text{ham} = 0.6 \times 0.2 \times 0.6 = 0.072

Log form:

log⁡(spam)=log⁡0.4+log⁡0.75+log⁡0.5≈−0.916+(−0.288)+(−0.693)=−1.897\log(\text{spam}) = \log 0.4 + \log 0.75 + \log 0.5 \approx -0.916 + (-0.288) + (-0.693) = -1.897 log⁡(ham)=log⁡0.6+log⁡0.2+log⁡0.6≈−0.511+(−1.609)+(−0.511)=−2.631\log(\text{ham}) = \log 0.6 + \log 0.2 + \log 0.6 \approx -0.511 + (-1.609) + (-0.511) = -2.631

Both forms agree: spam wins. The log scores are negative because probabilities are below 1, but the comparison is what matters.

Read the numbers. The feature "free" contributed −0.288-0.288 for spam versus −1.609-1.609 for ham — a swing of over 1.3 nats. The prior gave ham a head start of about 0.4 nats, but "free" overwhelmed it. That is the mechanism: each feature adds evidence, and the class with the largest total wins.

Knowledge check

Check your understanding

Answer this question before you continue.

For the worked message with x₁ = 1 and x₂ = 1, which class wins under the article's raw Naive Bayes scores?
Output Prediction

Focus: Use the article's smoothed class scores to identify the predicted class for the worked message.

spam: 0.4 × 0.75 × 0.5 = 0.15
ham: 0.6 × 0.2 × 0.6 = 0.072

See What the Independence Assumption Actually Changes

The example above proves the arithmetic. It does not yet show what the assumption costs you. To see that, compare the factorized likelihood against the true joint on a feature pair that is deliberately dependent.

Suppose two features are "contains free" and "contains prize." In spam, these words travel together: a message with "free" usually also has "prize." In ham, they almost never co-occur. The training counts might give:

P(x1=1∣spam)=0.6,P(x2=1∣spam)=0.6P(x_1 = 1 \mid \text{spam}) = 0.6, \quad P(x_2 = 1 \mid \text{spam}) = 0.6

The factorized likelihood multiplies them:

P(x1=1∣spam) P(x2=1∣spam)=0.36P(x_1 = 1 \mid \text{spam})\, P(x_2 = 1 \mid \text{spam}) = 0.36

But the actual class-conditional joint, read directly from the data, is:

P(x1=1,x2=1∣spam)=0.55P(x_1 = 1, x_2 = 1 \mid \text{spam}) = 0.55

The factorized model underestimates the joint by a wide margin. Why? Because it treats the two features as independent votes. In reality, "prize" adds almost no new evidence once you already know "free" is present — the two features are telling the same story. Naive Bayes counts that story twice, so its confidence is miscalibrated.

Now the crucial part. Suppose the ham joint for the same pair is P(x1=1,x2=1∣ham)=0.02P(x_1 = 1, x_2 = 1 \mid \text{ham}) = 0.02. The factorized ham score is 0.1×0.1=0.010.1 \times 0.1 = 0.01. Even though both factorized values are wrong, the ratio still favors spam by a wide margin. The ranking survives; the confidence does not.

That is the boundary the assumption draws. Conditional independence is a lie about the data, but it is often a lie that preserves the ordering of classes. When it does, Naive Bayes classifies well. When the dependence structure flips the ordering — for example, when two features are individually weak but jointly decisive for the wrong class — the classifier fails. The assumption is not "harmless." It is a bet that the ranking is robust to the distortion.

What the Assumption Does and Does Not Justify

The independence assumption is almost always false. Height and weight are not independent given sex. Word frequencies are not independent given topic. So why does Naive Bayes work?

Because classification only needs the ranking of classes to be correct, not the probabilities. Even when the posterior estimates are badly wrong, the argmax can still land on the right class. The errors from double-counting correlated features often cancel or reinforce the correct ranking.

Where the assumption bites:

  • Correlated features get double-counted. If two features carry the same signal, the score overstates confidence. The predicted class may be right, but the probability is inflated.
  • Posterior probabilities are not trustworthy. Use Naive Bayes for ranking and fast baselines, not for probability estimates you plan to act on directly.
  • Smoothing is not optional. Without it, a single unseen feature value zeros out the entire product for that class. Laplace smoothing is the standard fix.

My rule: reach for Naive Bayes when you need a fast, interpretable baseline on text or high-dimensional sparse data, and you care about ranking more than calibration. Reach for a discriminative model like logistic regression when you need calibrated probabilities or when features are strongly correlated and you can afford the training cost.

Prove It to Yourself

The derivation is short enough to hold in your head, but holding it is not the same as owning it. Two exercises will close the gap.

First, re-derive the two-feature case from memory. Write Bayes' rule, apply the chain rule, substitute independence, take the log. If you can do it without looking, the mechanism is yours.

Second, implement the log-score rule in a few lines of NumPy on the toy data above, then compare your predictions against scikit-learn's CategoricalNB on the same table. When your hand-computed scores match the library's output, you have proven the theory to yourself — and you will never again treat the formula as something you merely recognize.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

In the article's dependent-feature example, the factorized spam likelihood is 0.36 while the observed joint is 0.55. What conclusion about classification is justified?
Question 1 of 2Scenario Interpretation

Focus: Interpret how correlated features can distort Naive Bayes likelihoods and confidence without necessarily changing class ranking.

A feature value was never observed with one class in the training data. According to the article, what is the role of Laplace smoothing in this situation?
Question 2 of 2Scenario Interpretation

Focus: Explain how Laplace smoothing prevents an unseen feature value from zeroing a class score.

References

  1. 1.9. Naive Bayes — scikit-learn 1.9.0 documentationscikit-learn.org
  2. NAIVE BAYES AND LOGISTIC REGRESSION Machine ...www.cs.cmu.edu
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.