Skip to content
intermediate

Deriving the Logistic Regression Classification Objective

You can fit a logistic regression in one line and read its coefficients in the next. What almost nobody shows you is where the loss function comes from.…

Published 2026-10-02Updated 2026-10-048 min read
A vivid clownfish glides through the blue depths of an aquarium in Rio de Janeiro, showcasing its striking colors.
A vivid clownfish glides through the blue depths of an aquarium in Rio de Janeiro, showcasing its striking colors. Photo by Jonathan Borba on Pexels.

You can fit a logistic regression in one line and read its coefficients in the next. What almost nobody shows you is where the loss function comes from. LogisticRegression().fit() minimizes something called log loss, and most tutorials simply declare that it does. The formula arrives like a rule from above.

It is not a rule. It is a consequence. Once you build the probability model yourself, the loss function falls out of it, and the coefficients stop being opaque numbers.

This article assumes you already know the two-step picture: a linear score z=w⋅x+bz = w \cdot x + b, then a sigmoid that squashes it into a probability. That is the model. It is not yet the objective. We are going to derive the objective, then read the fitted parameters through it.

Why the Objective Needs Deriving

Here is the gap. You know the model produces a probability. You know scikit-learn minimizes log loss. What you cannot yet explain is why log loss and not squared error. When a formula is asserted rather than derived, it feels arbitrary, and arbitrary formulas are hard to debug when a model misbehaves.

The plan is short. Build a probability model. Write the likelihood of the labels you actually observed. Take the log. Show that maximizing that quantity is exactly minimizing cross-entropy. By the end, the loss is something you could have invented.

Notation and the Probability Model

Before any derivation, fix the symbols.

We have nn samples. Each sample ii has a feature vector xix_i and a label yi∈{0,1}y_i \in \{0, 1\}. The parameters are a weight vector ww and an intercept bb. The linear score is:

zi=w⋅xi+bz_i = w \cdot x_i + b

The model's estimated probability that yi=1y_i = 1 is the sigmoid of that score:

pi=σ(zi)=11+e−zip_i = \sigma(z_i) = \frac{1}{1 + e^{-z_i}}

Read pip_i carefully. It is a conditional probability, not a score and not a decision. It is the model's belief that this particular sample belongs to class 1, given its features.

Now the assumption that makes everything work. Each label is a Bernoulli draw with probability pip_i, and the observations are conditionally independent given the features. Independence is the load-bearing wall: it is what lets us multiply per-sample probabilities into a joint probability later.

We can write the probability of a single observed label in one compact expression:

P(yi∣xi)=piyi⋅(1−pi)1−yiP(y_i \mid x_i) = p_i^{y_i} \cdot (1 - p_i)^{1 - y_i}

Check both cases. If yi=1y_i = 1, the second factor becomes (1−pi)0=1(1-p_i)^0 = 1, leaving pip_i. If yi=0y_i = 0, the first factor becomes pi0=1p_i^0 = 1, leaving 1−pi1 - p_i. One expression, both labels. That compactness is not decoration; it is what makes the next step clean.

Note: Independence is an assumption, not a fact about your data. It fails for grouped or repeated measures, where samples within a group are correlated. The derivation still produces a usable estimator in that case, but the likelihood is no longer the true joint probability, so any uncertainty estimates built on it deserve suspicion.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article's probability-model assumptions, what does conditional independence let us do?
Single Choice

Focus: Explain how conditional independence supports constructing the joint likelihood from per-sample probabilities.

From Likelihood to Log-Likelihood

Four connected stages show a Bernoulli probability for one observed label, multiplication across samples to form the likelihood, a logarithm converting the product into a sum, and negation with averaging to produce cross-entropy loss.
Follow the probability model through the likelihood transformations to see where the cross-entropy objective comes from.

Multiply the per-sample probabilities across all nn samples. That product is the likelihood of the observed labels:

L(w,b)=∏i=1npiyi(1−pi)1−yiL(w, b) = \prod_{i=1}^{n} p_i^{y_i} (1 - p_i)^{1 - y_i}

This is the probability of the data you saw, as a function of the parameters. Maximum likelihood estimation says: pick the parameters that make the observed data most probable.

There is a practical problem. Every factor sits between 0 and 1. Multiply a few hundred of them and the product underflows to zero in floating point. Products are also painful to differentiate.

Take the log. The log is monotonic, so it does not move the maximum — the parameters that maximize LL are the same ones that maximize log⁡L\log L. This is the step most tutorials skip, and it is the whole reason the log is not a convenience trick. It is a structural simplification.

The product becomes a sum, and the exponents become multiplicative selectors:

log⁡L(w,b)=∑i=1n[yilog⁡(pi)+(1−yi)log⁡(1−pi)]\log L(w, b) = \sum_{i=1}^{n} \left[ y_i \log(p_i) + (1 - y_i) \log(1 - p_i) \right]

For each sample, exactly one term survives. If the true label is 1, you get log⁡(pi)\log(p_i). If it is 0, you get log⁡(1−pi)\log(1 - p_i). The log-likelihood is just a sum of per-sample log-probabilities.

Knowledge check

Check your understanding

Answer this question before you continue.

What happens when we take the logarithm of the likelihood product?
Comparison Reasoning

Focus: Describe how taking the logarithm changes the likelihood expression without changing its maximizing parameters.

Maximizing Likelihood Is Minimizing Cross-Entropy

Optimizers are built to minimize. So negate the log-likelihood and divide by nn to get an average:

Loss(w,b)=−1n∑i=1n[yilog⁡(pi)+(1−yi)log⁡(1−pi)]\text{Loss}(w, b) = -\frac{1}{n} \sum_{i=1}^{n} \left[ y_i \log(p_i) + (1 - y_i) \log(1 - p_i) \right]

That is the cross-entropy loss, also called log loss. Maximizing likelihood and minimizing negative log-likelihood are the same optimization problem, stated in the direction the optimizer expects. Nothing was added. The sign flipped.

Substitute pi=σ(zi)p_i = \sigma(z_i) and you have the loss in terms of the linear score. The sigmoid's derivative cancels cleanly against the log's derivative in the gradient, which is a hint that this pairing is not accidental.

Why not squared error? Suppose the true label is 1 and the model predicts p=0.01p = 0.01. Squared error gives roughly 0.980.98, a bounded penalty. Cross-entropy gives −log⁡(0.01)≈4.6-\log(0.01) \approx 4.6, and it grows without bound as the prediction approaches 0. Squared error on a sigmoid output also produces a non-convex objective with vanishing gradients for confidently wrong predictions. Cross-entropy stays convex in the parameters and punishes confident mistakes sharply.

Warning: Cross-entropy is the right objective when the model outputs a probability for a binary label. It is not automatically right for ranking or ordinal outcomes, where the target structure is not Bernoulli. Unequal decision costs are a separate issue: they usually change the decision threshold or the training setup, not the likelihood derivation itself.

Knowledge check

Check your understanding

Answer this question before you continue.

For a sample with true label 1, the model predicts p = 0.01. Approximately what per-sample cross-entropy loss does this produce?
Output Prediction

Focus: Estimate the cross-entropy contribution of a confidently incorrect positive-label prediction.

Worked Example: One Feature, Two Coefficients

Let's make this concrete with four points and one feature.

xix_iyiy_i
-10
00
11
21

Start with w=1w = 1, b=0b = 0. Compute zi=wxi+bz_i = w x_i + b, then pi=σ(zi)p_i = \sigma(z_i):

xix_iyiy_iziz_ipip_iper-sample loss
-10-10.2690.313
0000.5000.693
1110.7310.313
2120.8810.127

The average loss is (0.313+0.693+0.313+0.127)/4≈0.361(0.313 + 0.693 + 0.313 + 0.127) / 4 \approx 0.361.

Now nudge ww up to 1.51.5. The scores spread out, the probabilities sharpen, and the average loss drops to roughly 0.2760.276. The loss moved in the direction the derivation predicts: increasing the weight on a feature that separates the classes lowers the negative log-likelihood.

Look at the asymmetry. The sample at x=2x = 2 with a true label of 1 contributes only 0.1270.127 because the model is already confident and correct. If the model had predicted p=0.01p = 0.01 for that same label, the term would be 4.64.6 — a single sample dominating the sum. That is the behavior the derivation bought us: confident errors are expensive, confident correct answers are nearly free.

You can verify the hand calculation against scikit-learn on the same data. Fit LogisticRegression with regularization disabled and compare the reported loss. The numbers should agree.

Reading the Fitted Parameters

Now the payoff. A fitted coefficient is the change in the log-odds of the positive class per one-unit increase in that feature, holding the others fixed. Log-odds is log⁡(p/(1−p))\log(p / (1-p)), and the model is linear in it.

Exponentiate a coefficient and you get an odds ratio — a multiplicative change in odds. An odds ratio of 2 means a one-unit increase doubles the odds. That framing is often easier to hand to a stakeholder than a raw log-odds number.

There is a fast approximation worth knowing. Divide a coefficient by 4 to estimate the largest change in predicted probability a one-unit increase can produce. The maximum slope of the sigmoid is 0.250.25, so this is where the shortcut comes from. It is a rough read, not a precise one.

Common mistake: A coefficient is not a change in probability, and its magnitude is not comparable across features measured on different scales. Standardize before comparing.

One failure mode deserves a flag. When the classes are perfectly separable, the likelihood keeps increasing as the weights grow without bound, so the coefficients diverge. Regularization is the practical fix, and it is why scikit-learn applies it by default.

Knowledge check

Check your understanding

Answer this question before you continue.

Holding other features fixed, what does a coefficient c mean for a one-unit increase in its feature?
Misconception Check

Focus: Interpret a logistic-regression coefficient as a change in log-odds and its exponent as an odds ratio.

What the Derivation Buys You

The loss function is not a library convention. It is the negative log-likelihood of a Bernoulli model, and that is why it behaves the way it does.

The decision rule: when you see a probability output paired with cross-entropy loss, you are looking at a maximum-likelihood fit. When the loss and the output distribution disagree, the model is mismatched to the problem.

This same derivation generalizes. Swap the Bernoulli for a categorical distribution and you get softmax and multiclass cross-entropy. The pattern recurs anywhere a model outputs probabilities.

Your next move: take the gradient of the cross-entropy loss with respect to ww and bb and confirm it reduces to (pi−yi)⋅xi(p_i - y_i) \cdot x_i. That clean form is what makes gradient descent on logistic regression almost trivial to implement, and deriving it yourself is the natural continuation of everything above.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

In the worked example with w = 1 and b = 0, the sample x = 2 has label 1 and predicted probability 0.881. Why is its per-sample loss relatively small?
Question 1 of 2Scenario Interpretation

Focus: Relate a correct, confident prediction in the worked example to its per-sample loss.

A binary classifier outputs probabilities and is fitted by minimizing cross-entropy. According to the derivation, what is the key interpretation?
Question 2 of 2Comparison Reasoning

Focus: Connect binary probability modeling with cross-entropy as a maximum-likelihood objective.

References

  1. A Tutorial on Regression Analysis: From Linear Models to Deep Learning — Lecture Notes on Artificial Intelligence at Beihang Universityarxiv.org
  2. A Fast Hybrid Algorithm for Large-Scale ℓ1-Regularized ...jmlr.org
  3. Machine Learning Glossarydevelopers.google.com
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.