Skip to content
intermediate

Derive the Poisson Regression Objective With a Log Link

A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.

Published 2026-10-02Updated 2026-10-0410 min read
Aerial view of a traditional leather tannery in Fes, Morocco, showcasing colorful dye pits.
Aerial view of a traditional leather tannery in Fes, Morocco, showcasing colorful dye pits. Photo by Ramon Karolan on Pexels.

A straight line on count data will eventually promise you a negative number of events. The Poisson likelihood with a log link is what stops it.

Why Counts Need Their Own Objective

Suppose you are modeling the number of support tickets a customer files in a month, the number of failures a machine logs per shift, or the number of purchases a visitor makes in a session. Your target lives on {0,1,2,3,… }\{0, 1, 2, 3, \dots\}: non-negative integers, usually with a pile of zeros and a long right tail.

Ordinary least squares does not respect that support. Fit a line to counts and, for some feature values, the predicted mean drops below zero. A negative expected count is not a small numerical annoyance; it is a statement that cannot be true. OLS also assumes the noise has constant variance, but counts rarely behave that way — when the mean count is large, the spread tends to be large too.

The obvious patch is to log-transform the raw counts and run OLS on log⁡y\log y. It fails the moment a single observation is zero, because log⁡0\log 0 is undefined. Adding one before the log keeps the code running, but now you are fitting a different model than the one you described, and the coefficients no longer mean what you think they mean.

The real fix is not to transform the data. It is to model the mean through a link function and let the distribution handle the integer support. That is the generalized linear model move, and the Poisson case is the cleanest place to see it work.

Note: If the linear predictor, link function, and response distribution split still feels abstract, treat this article as the concrete instance. Everything here is one specific choice of distribution and one specific choice of link.

Notation and the Two Modeling Choices

Before any derivation, pin down the symbols.

We observe pairs (yi,xi)(y_i, \mathbf{x}_i) for i=1,…,ni = 1, \dots, n. Here yiy_i is a realized count and xi\mathbf{x}_i is a feature vector for observation ii. The parameters are a coefficient vector β\boldsymbol{\beta}, which includes an intercept. The linear predictor is

ηi=xi⊤β.\eta_i = \mathbf{x}_i^\top \boldsymbol{\beta}.

This quantity is unbounded: it can be any real number. That is the whole tension. Counts are non-negative integers; linear predictors are unrestricted reals. Two independent decisions bridge the gap.

Choice 1 — the response distribution. We assume

yi∣xi∼Poisson(λi),y_i \mid \mathbf{x}_i \sim \text{Poisson}(\lambda_i),

with probability mass function

P(y;λ)=e−λλyy!,y∈{0,1,2,… }.P(y; \lambda) = \frac{e^{-\lambda} \lambda^{y}}{y!}, \qquad y \in \{0, 1, 2, \dots\}.

The Poisson family already lives on the correct support. Its single parameter λ\lambda is both the mean and the variance, a fact we will return to when we discuss assumptions.

Choice 2 — the link. We connect the mean to the linear predictor with

log⁡(λi)=ηi,equivalentlyλi=exp⁡(ηi).\log(\lambda_i) = \eta_i, \qquad \text{equivalently} \qquad \lambda_i = \exp(\eta_i).

Why the log? Two reasons. First, exp⁡(⋅)\exp(\cdot) maps any real number to a strictly positive one, so the predicted mean can never go negative no matter what β\boldsymbol{\beta} or xi\mathbf{x}_i do. Second, the log is the canonical link for the Poisson family — the one that falls out of writing the Poisson PMF in exponential-family form. Canonical links tend to produce well-behaved, convex objectives, which is exactly what we will find.

We also assume the observations are independent given the features. The likelihood is a product, and that product only factors cleanly under independence.

Knowledge check

Check your understanding

Answer this question before you continue.

A fitted linear predictor is negative for one observation. What does the Poisson log-link model imply about that observation’s expected count?
Misconception Check

Focus: Explain how the log link maps an unrestricted linear predictor to a valid Poisson mean.

From the Poisson PMF to the Likelihood

Start with a single observation. Its contribution to the likelihood is just the PMF evaluated at the observed count:

P(yi;λi)=e−λiλiyiyi!.P(y_i; \lambda_i) = \frac{e^{-\lambda_i} \lambda_i^{y_i}}{y_i!}.

Now substitute the link. Since λi=exp⁡(ηi)=exp⁡(xi⊤β)\lambda_i = \exp(\eta_i) = \exp(\mathbf{x}_i^\top \boldsymbol{\beta}), the PMF becomes a function of β\boldsymbol{\beta}:

P(yi;β)=exp⁡ ⁣(−exp⁡(xi⊤β)) exp⁡(xi⊤β)yiyi!.P(y_i; \boldsymbol{\beta}) = \frac{\exp\!\big(-\exp(\mathbf{x}_i^\top \boldsymbol{\beta})\big) \, \exp(\mathbf{x}_i^\top \boldsymbol{\beta})^{y_i}}{y_i!}.

Independence lets us multiply these contributions across all nn observations. The likelihood is

L(β)=∏i=1nexp⁡ ⁣(−exp⁡(xi⊤β)) exp⁡(xi⊤β)yiyi!.L(\boldsymbol{\beta}) = \prod_{i=1}^{n} \frac{\exp\!\big(-\exp(\mathbf{x}_i^\top \boldsymbol{\beta})\big) \, \exp(\mathbf{x}_i^\top \boldsymbol{\beta})^{y_i}}{y_i!}.

This is the count regression likelihood in its raw form. Notice that the yi!y_i! terms do not contain β\boldsymbol{\beta} at all. They are constants with respect to the parameters, so they will drop out of any optimization. That is why most treatments write the objective without them.

Knowledge check

Check your understanding

Answer this question before you continue.

Given independent observations with Poisson means λᵢ = exp(xᵢᵀβ), which statement correctly describes the likelihood and the role of yᵢ!?
Single Choice

Focus: Construct the joint likelihood from independent Poisson observations and identify parameter-free factors.

Taking the Log: The Objective You Actually Optimize

A three-stage sequence: the linear predictor eta becomes the positive mean lambda equals exp eta; a single Poisson observation contributes y eta minus exp eta to the log-likelihood after dropping the parameter-free log factorial; independent observations combine into the negative log-likelihood sum of exp eta minus y eta.
The log link turns each count into a likelihood contribution; independence adds those contributions, and changing sign gives the objective to minimize.

Products of many small probabilities underflow fast, and differentiating a product is painful. Take the natural log. The product becomes a sum:

log⁡L(β)=∑i=1n[−exp⁡(ηi)+yiηi−log⁡(yi!)],\log L(\boldsymbol{\beta}) = \sum_{i=1}^{n} \Big[ -\exp(\eta_i) + y_i \eta_i - \log(y_i!) \Big],

where I have written ηi\eta_i for xi⊤β\mathbf{x}_i^\top \boldsymbol{\beta} to keep the line readable. Drop the β\boldsymbol{\beta}-free constant ∑ilog⁡(yi!)\sum_i \log(y_i!) and you have the log-likelihood up to an additive constant:

ℓ(β)=∑i=1n[yiηi−exp⁡(ηi)].\ell(\boldsymbol{\beta}) = \sum_{i=1}^{n} \Big[ y_i \eta_i - \exp(\eta_i) \Big].

Solvers minimize rather than maximize, so flip the sign to get the negative log-likelihood:

NLL(β)=∑i=1n[exp⁡(ηi)−yiηi].\text{NLL}(\boldsymbol{\beta}) = \sum_{i=1}^{n} \Big[ \exp(\eta_i) - y_i \eta_i \Big].

Two terms are doing all the work, and each has a clear job.

  • yiηiy_i \eta_i rewards pushing the linear predictor up when the observed count is large. A big yiy_i makes the model want a big ηi\eta_i.
  • −exp⁡(ηi)-\exp(\eta_i) penalizes large predicted means regardless of what was observed. Even when yiy_i is zero, this term pulls ηi\eta_i down.

The two terms fight, and the balance point is not where each prediction equals its observed count. It is where the feature-weighted residuals balance across the whole dataset. To see why, differentiate the objective with respect to β\boldsymbol{\beta} and set the gradient to zero:

∂ℓ∂β=∑i=1n(yi−exp⁡(ηi))xi=0.\frac{\partial \ell}{\partial \boldsymbol{\beta}} = \sum_{i=1}^{n} \big( y_i - \exp(\eta_i) \big) \mathbf{x}_i = \mathbf{0}.

Read that equation carefully. Each observation contributes a residual yi−λ^iy_i - \hat{\lambda}_i, weighted by its feature vector xi\mathbf{x}_i. At the optimum, those weighted residuals cancel out across the dataset. A single observation with yi=4y_i = 4 and λ^i=3.7\hat{\lambda}_i = 3.7 is not forced to match; it is balanced against every other observation through the shared β\boldsymbol{\beta}. The model trades off errors across all rows to satisfy one aggregate condition per coefficient.

That is the mechanism behind every fitted Poisson model: not per-point matching, but a joint, feature-weighted compromise.

One more property matters in practice: this objective is convex in β\boldsymbol{\beta}. There is a single global optimum and no local-minimum trap. It is a large part of why Poisson GLMs are so easy to fit reliably.

Knowledge check

Check your understanding

Answer this question before you continue.

After dropping terms constant in β, which expression is the negative log-likelihood that a minimization-based solver can optimize?
Single Choice

Focus: Identify the negative log-likelihood optimized for Poisson regression, up to terms constant in the parameters.

Worked Example: One Observation, One Feature

Numbers make the algebra concrete. Take a single observation with y=4y = 4 and x=1x = 1, and a candidate β=(0.5,0.8)\boldsymbol{\beta} = (0.5, 0.8), where the first entry is the intercept.

Compute the linear predictor:

η=0.5+0.8⋅1=1.3.\eta = 0.5 + 0.8 \cdot 1 = 1.3.

Convert to the mean:

λ=exp⁡(1.3)≈3.669.\lambda = \exp(1.3) \approx 3.669.

Evaluate the single-observation log-likelihood term:

yη−exp⁡(η)−log⁡(y!)=4(1.3)−3.669−log⁡(24)≈5.2−3.669−3.178≈−1.647.y\eta - \exp(\eta) - \log(y!) = 4(1.3) - 3.669 - \log(24) \approx 5.2 - 3.669 - 3.178 \approx -1.647.

Change β\boldsymbol{\beta} slightly and the value moves. That surface — one number per candidate parameter vector — is what the solver walks downhill.

You can verify the same quantity in code. The negative log-likelihood for a fitted model is a direct translation of the formula:

import numpy as np
from scipy.special import gammaln

def poisson_nll(beta, X, y):
    eta = X @ beta
    # gammaln(y + 1) == log(y!) but stable for large counts
    return np.sum(np.exp(eta) - y * eta + gammaln(y + 1))

Fit a PoissonRegressor from scikit-learn on a small count dataset, then plug the fitted coef_ and intercept_ into this function. The value you compute by hand should match the model's reported log-likelihood up to sign. When it does, you have confirmed that the derivation and the library are solving the same problem.

Common mistake: Comparing a hand-computed log-likelihood to a library's reported value without checking whether the library includes the log⁡(y!)\log(y!) constant. Some report the full log-likelihood; some report it up to a constant. The difference is a fixed offset, not a modeling error.

Reading the Coefficients: Multiplicative Effects

Because log⁡(λ)=x⊤β\log(\lambda) = \mathbf{x}^\top \boldsymbol{\beta}, a one-unit increase in feature jj adds βj\beta_j to the log of the expected count. Exponentiating, it multiplies the expected count by exp⁡(βj)\exp(\beta_j).

CoefficientEffect on the expected count
βj=0\beta_j = 0No change in the rate
βj>0\beta_j > 0Count grows by a factor of exp⁡(βj)\exp(\beta_j)
βj<0\beta_j < 0Count shrinks by a factor of exp⁡(βj)\exp(\beta_j)

This is the same log-to-odds move you already know from logistic regression, except the quantity being scaled is a rate rather than a probability. In epidemiology and economics, exp⁡(βj)\exp(\beta_j) is often called an incidence rate ratio.

When counts come from windows of different sizes — different observation periods, different populations — you add log⁡(exposure)\log(\text{exposure}) to the linear predictor with a fixed coefficient of 1. This is an offset, and it turns the model into a rate model: you are now predicting events per unit of exposure rather than raw counts.

Assumptions and Where the Model Breaks

The derivation is clean, but it rests on assumptions that real data violates constantly.

Mean equals variance. The Poisson family forces Var(y∣x)=E(y∣x)=λ\text{Var}(y \mid \mathbf{x}) = \mathbb{E}(y \mid \mathbf{x}) = \lambda. Real count data is frequently overdispersed: the variance exceeds the mean. Overdispersion inflates standard errors and makes the model look more confident than the evidence supports. Negative binomial or quasi-Poisson models are the usual escapes.

Independence. The product likelihood assumes observations do not influence each other. Clustered or repeated-measures data violates this, and the standard errors will be wrong even if the coefficients look reasonable.

Multiplicative structure. The log link assumes effects compound multiplicatively on the mean. That is a strong structural claim, not a neutral default. If your domain expects additive effects, the log link will fight you.

Excess zeros. If your data has more zeros than the Poisson predicts, a plain Poisson fit will be biased. Zero-inflated or hurdle models are the right tool.

Warning: Convexity guarantees a unique optimum. It says nothing about whether the Poisson distribution is the right description of your data. A perfectly fitted wrong model is still wrong.

Knowledge check

Check your understanding

Answer this question before you continue.

A count dataset has conditional variance substantially greater than its conditional mean. Which interpretation best follows from the article?
Scenario Interpretation

Focus: Recognize when a violation of the Poisson mean-variance assumption motivates considering an alternative count model.

When to Reach for This Model

Use Poisson regression when the target is a non-negative integer count and the mean-variance relationship is roughly Poisson. Prefer it over OLS on counts whenever negative predictions or heteroscedasticity would distort the fit. Prefer it over log-transformed OLS when zeros are present or when you want a principled likelihood rather than an ad hoc transform.

Move to negative binomial, zero-inflated, or hurdle variants when diagnostics point to overdispersion or excess zeros. The derivation you just worked through is the foundation those models extend.

Next Step

Fit a PoissonRegressor on a small count dataset. Compute the negative log-likelihood by hand for the fitted coefficients using the formula above, and compare it to the model's reported value. Then check the mean-variance relationship in the residuals: if the variance clearly exceeds the mean, that is your signal to look at overdispersed alternatives next.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

In a fitted Poisson regression, a feature’s coefficient is ln(2). Holding other features fixed, what does increasing that feature by one unit do to the expected count?
Question 1 of 2Scenario Interpretation

Focus: Interpret a Poisson regression coefficient as a multiplicative change in expected count for a one-unit feature increase.

A count outcome has many more zeros than a Poisson model predicts, and its variance also exceeds its mean. Which modeling response is most consistent with the article?
Question 2 of 2Comparison Reasoning

Focus: Use count-data diagnostics to distinguish when plain Poisson regression is appropriate from when alternatives may be needed.

References

  1. PoissonRegressor — scikit-learn 1.9.1 documentationscikit-learn.org
  2. Chapter 4 Poisson Regression | Beyond Multiple Linear Regressionbookdown.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.