Skip to content
beginner

Deriving Gradient Updates From a Differentiable Model Objective

You can write theta = theta - lr * grad from memory and still freeze when someone asks where grad came from. That gap is the whole problem. The update rule…

Published 2026-10-02Updated 2026-10-0410 min read
A clear blue sky featuring fluffy white clouds bathed in sunlight.
A clear blue sky featuring fluffy white clouds bathed in sunlight. Photo by Stephen Leonardi on Pexels.

You can write theta = theta - lr * grad from memory and still freeze when someone asks where grad came from. That gap is the whole problem. The update rule is not a magic spell you memorize. It is a short derivation you can rebuild from any differentiable objective.

This article is a derivation drill. We will pick a simple objective, differentiate it, read the update rule off the result, and then watch what the step size does to the path. By the end, you should be able to look at a new objective and derive its update yourself instead of hunting for a formula.

Why the Update Rule Looks Like Magic

Most beginners meet gradient descent in the same order: someone shows a loss function, someone shows an update rule, and the two are connected by a hand wave. You learn that the gradient points "downhill" and that you subtract it. You nod. You move on.

Then you hit a new objective — a weighted loss, a regularized model, a custom penalty — and you have no idea how to update anything. You can recite the rule but you cannot rebuild it.

The fix is to treat the update rule as a consequence, not a definition. Given an objective, the update falls out of a small amount of calculus. That is the durable skill. The specific optimizer you use later is replaceable; the derivation is not.

Before we start, let us be honest about the assumptions. We assume the objective is differentiable in the parameters — meaning we can compute its slope everywhere we care about. We assume we only use first-order information (the gradient, not curvature). And we assume we take small steps, because the math we use is only accurate locally. Hold onto that last one; it explains almost every training failure you will ever debug.

Notation Before Formulas

I want every symbol defined before it appears in an equation. Skipping this step is why derivations feel like fog.

  • Parameters: the numbers the model learns. When there are many, we collect them into a vector θ\theta (theta). For the worked example, we will use a single scalar parameter ww.
  • Objective: a scalar function J(θ)J(\theta) that measures how wrong the model is. Smaller is better. We want to minimize it.
  • Gradient: the vector of partial derivatives, written ∇J(θ)\nabla J(\theta). A partial derivative is a derivative where you hold every other variable constant and ask how the objective changes when only one parameter moves. The gradient stacks all those partials together.
  • Step size (also called the learning rate): a small positive scalar, written η\eta (eta) or lr, that controls how far we move along the descent direction.
  • Iteration index: tt, so θt\theta_t is the parameter value at step tt.

One plain-language translation carries the whole idea: the gradient points in the direction of steepest increase, so its negative points in the direction of steepest decrease. Everything below is just making that sentence precise enough to compute with.

From Objective to Update: The Derivation

Start with the goal. We want the parameter value that minimizes J(θ)J(\theta):

θ∗=arg⁡min⁡θJ(θ)\theta^* = \arg\min_\theta J(\theta)

We cannot solve this directly for most objectives, so we take an iterative approach. Suppose we are sitting at some current value θt\theta_t. We want to choose a small change — call it Δθ\Delta\theta — so that the new objective J(θt+Δθ)J(\theta_t + \Delta\theta) is smaller than J(θt)J(\theta_t).

To reason about small changes, we use a first-order Taylor expansion. This says that for a small step, the objective changes approximately by the gradient dotted with the step:

J(θt+Δθ)≈J(θt)+∇J(θt)⊤ΔθJ(\theta_t + \Delta\theta) \approx J(\theta_t) + \nabla J(\theta_t)^\top \Delta\theta

Read that second term carefully. It is the only part that depends on our choice of Δθ\Delta\theta. To make the objective decrease, we need that term to be negative:

∇J(θt)⊤Δθ<0\nabla J(\theta_t)^\top \Delta\theta < 0

Now ask: what choice of Δθ\Delta\theta makes this most negative? The dot product of two vectors is most negative when they point in opposite directions. So the best descent direction is the one directly opposed to the gradient:

Δθ∝−∇J(θt)\Delta\theta \propto -\nabla J(\theta_t)

We introduce the step size η\eta as the scalar that controls how far we move along that direction:

Δθ=−η ∇J(θt)\Delta\theta = -\eta \, \nabla J(\theta_t)

Substituting this back gives the update rule:

θt+1=θt−η ∇J(θt)\theta_{t+1} = \theta_t - \eta \, \nabla J(\theta_t)

That is the whole derivation. Each piece has a job:

PieceWhat it does
θt\theta_tWhere we are now
∇J(θt)\nabla J(\theta_t)Which way is uphill
The minus signFlip it to go downhill
η\etaHow far to travel

The descent condition is the fine print. This update is guaranteed to decrease JJ only when the step is small enough that the Taylor approximation still holds. It is a local guarantee — it says nothing about whether you will reach the global minimum. That caveat is not pedantry. It is the reason step size matters so much.

Knowledge check

Check your understanding

Answer this question before you continue.

For a small step from \(\theta_t\), which choice of \(\Delta\theta\) makes the first-order change in the objective negative when the gradient is nonzero?
Single Choice

Focus: Choose a parameter change that makes the first-order change in a differentiable objective negative.

A Worked Example You Can Check by Hand

A U-shaped objective curve with marked points w₀ = 0, w₁ = 0.6, and w₂ = 1.08 moving toward the minimum at w* = 3. The successive points sit lower on the curve, and the steps shorten near the minimum.
On this quadratic, each update moves the parameter toward the minimum and lowers the objective; the steps shrink as the slope flattens.

Let us make this concrete with an objective simple enough to differentiate in your head:

J(w)=(w−3)2J(w) = (w - 3)^2

This is a parabola with its minimum at w=3w = 3. The gradient is:

dJdw=2(w−3)\frac{dJ}{dw} = 2(w - 3)

Start at w0=0w_0 = 0 with step size η=0.1\eta = 0.1. Apply the update three times.

Iteration 1. Gradient is 2(0−3)=−62(0 - 3) = -6. Update: w1=0−0.1(−6)=0.6w_1 = 0 - 0.1(-6) = 0.6. Objective: (0.6−3)2=5.76(0.6 - 3)^2 = 5.76.

Iteration 2. Gradient is 2(0.6−3)=−4.82(0.6 - 3) = -4.8. Update: w2=0.6−0.1(−4.8)=1.08w_2 = 0.6 - 0.1(-4.8) = 1.08. Objective: (1.08−3)2=3.69(1.08 - 3)^2 = 3.69.

Iteration 3. Gradient is 2(1.08−3)=−3.842(1.08 - 3) = -3.84. Update: w3=1.08−0.1(−3.84)=1.464w_3 = 1.08 - 0.1(-3.84) = 1.464. Objective: (1.464−3)2=2.36(1.464 - 3)^2 = 2.36.

IterationwwJ(w)J(w)Gradient
00.0009.00-6.00
10.6005.76-4.80
21.0803.69-3.84
31.4642.36-3.07

Two things to notice. First, the parameter marches toward 3 and the objective falls every step. Second, the moves shrink. As ww approaches the minimum, the gradient approaches zero, so each update gets smaller. The geometry is exactly what the algebra promised: the gradient is the slope of the curve, and the update slides down that slope.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, after reaching \(w_1=0.6\), what is the next parameter value using \(\eta=0.1\) and \(dJ/dw=2(w-3)\)?
Output Prediction

Focus: Apply the derived update to calculate one parameter step in the worked quadratic example.

What Step Size Actually Changes

Now rerun the same starting point, w0=0w_0 = 0, with different step sizes. The objective has not changed. Only η\eta has.

For this particular objective, we can do better than guess. Define the distance from the minimum as et=wt−3e_t = w_t - 3. Substitute the update rule and simplify:

et+1=wt+1−3=(wt−3)−η⋅2(wt−3)=(1−2η) ete_{t+1} = w_{t+1} - 3 = (w_t - 3) - \eta \cdot 2(w_t - 3) = (1 - 2\eta) \, e_t

That single number, (1−2η)(1 - 2\eta), is the error multiplier. It tells you everything about the path. If its absolute value is below 1, the distance shrinks and you converge. If it is negative, the sign flips each step and you oscillate. If its absolute value is at least 1, the distance does not shrink and you diverge.

Step size η\etaError multiplier 1−2η1 - 2\etaBehavior
0.010.98Slow, monotone approach
0.10.80Fast, monotone approach
0.50.00Lands exactly on the minimum in one step
0.9-0.80Oscillates, but the distance still shrinks
1.1-1.20Oscillates and grows — diverges

Now the earlier list makes sense, and one entry changes meaning. With η=0.9\eta = 0.9, the parameter does overshoot and the sign does flip each step — but the objective still falls every iteration, because the distance from the minimum shrinks by a factor of 0.8 each time. Overshooting is not the same as failing. Only when ∣1−2η∣≥1|1 - 2\eta| \ge 1 does the objective actually grow.

Common mistake: Beginners see a parameter jump past the minimum and assume training is broken. Watch the objective, not the parameter. Oscillation with a shrinking distance is still convergence.

Why does the general case need the small-step caveat if this example is exact? Because a parabola is a quadratic, and a first-order Taylor expansion of a quadratic is not an approximation at all — it is the function. For a curved, non-quadratic objective, the linear term only describes the surface near your current point. Step too far and you leave the neighborhood where that description holds, so the "descent direction" stops being a descent direction. The parabola gives you an exact, checkable model of the mechanism; real objectives give you the local version.

Common mistake: When loss jumps up or oscillates, beginners blame the model, the data, or the features. Check the step size first. A learning rate that is too high produces exactly this signature, and it is the cheapest thing to fix.

Knowledge check

Check your understanding

Answer this question before you continue.

For \(J(w)=(w-3)^2\), a step size of \(\eta=0.9\) gives error multiplier \(1-2\eta=-0.8\). What behavior follows?
Scenario Interpretation

Focus: Interpret the error multiplier to distinguish convergent oscillation from divergence.

When This Derivation Applies — and When It Breaks

The derivation rests on a few conditions. Know them so you can tell when the update rule is the wrong tool.

It applies when the objective is differentiable in the parameters and you can compute or reasonably approximate the gradient. That covers a large share of classical models: linear regression with squared error, logistic regression with cross-entropy, and many regularized variants.

It weakens or breaks when:

  • The objective is non-differentiable (for example, a hard count or a step function), so no gradient exists.
  • Gradients are noisy or vanishing, so the direction you compute is unreliable.
  • Features are badly scaled. Scaling changes the shape of the objective — it stretches the parabola into a narrow valley — which changes the effective step size in each direction. This is why feature scaling shows up again and again in practical training.

More advanced optimizers exist, and they are worth learning later. But they all rest on the same idea you just derived: compute a gradient, take a step against it. The derivation is the foundation; the optimizer is the variation.

Knowledge check

Check your understanding

Answer this question before you continue.

Which objective feature directly prevents applying this derivation as written?
Misconception Check

Focus: Identify a type of objective for which the differentiable-gradient derivation does not directly apply.

Verify the Derivation in a Few Lines of NumPy

The code here is a check on the math, not a new lesson. It should reproduce the table above.

import numpy as np

def J(w):
    return (w - 3.0) ** 2

def grad_J(w):
    return 2.0 * (w - 3.0)

w = 0.0
lr = 0.1
for t in range(4):
    print(f"t={t}  w={w:.3f}  J(w)={J(w):.3f}  grad={grad_J(w):.3f}")
    w = w - lr * grad_J(w)

If the printed values match the hand-computed table, your derivation is correct. As a second check, compare the analytic gradient to a finite-difference estimate:

eps = 1e-6
w_test = 1.0
numeric = (J(w_test + eps) - J(w_test - eps)) / (2 * eps)
print(numeric, grad_J(w_test))  # should agree closely

When the analytic and numeric gradients agree, you have confirmed the derivative. When they disagree, you have a bug in your math, not in your optimizer.

The Rule to Carry Forward

Before you trust any training run, do two things. First, derive the update from the objective so you know what the code is actually doing. Second, watch whether the loss falls — if it rises or oscillates, suspect the step size before anything else.

The derivation is the durable skill. The specific optimizer is replaceable. Your next step is to apply this same process to a real model objective: take the least-squares loss for a linear model, differentiate it with respect to each coefficient, and read off the update rule. Same drill, more parameters, and now you know exactly where the minus sign comes from.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A learner derives a gradient update using the article's first-order reasoning. Which conclusion is justified by that derivation?
Question 1 of 2Comparison Reasoning

Focus: Distinguish the local decrease guarantee of a small gradient step from a claim of global optimization.

For the article's objective \(J(w)=(w-3)^2\), what value should the analytic gradient return at \(w=1\)?
Question 2 of 2Output Prediction

Focus: Evaluate the worked objective's analytic derivative at a specified parameter value.

References

  1. Gradient Descent, Stochastic Optimization, and Other Talesarxiv.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.