Skip to content
intermediate

KNN Regression as Local Averaging: Neighborhood Size, Bias, and Variance

Fit KNN regression with k=1 and the prediction curve threads through every training point. Set k=50 and the same curve flattens toward a horizontal line.…

Published 2026-10-02Updated 2026-10-0412 min read
Close-up of business analytics charts and graphs on papers and clipboard.
Close-up of business analytics charts and graphs on papers and clipboard. Photo by RDNE Stock project on Pexels.

Fit KNN regression with k=1 and the prediction curve threads through every training point. Set k=50 and the same curve flattens toward a horizontal line. The natural conclusion is that small k is accurate and large k is smooth — that k is an accuracy dial you turn toward "better."

That conclusion is wrong, and it produces bad models. k is not an accuracy dial. It is a bandwidth: the width of a window you slide across the input space, averaging whatever responses fall inside. Widen the window and you smooth away noise, but you also average across regions where the true relationship has already moved. Narrow it and you track the data exactly, including the parts that are pure noise.

This article derives the prediction as a neighborhood average, works a small example by hand, and then reasons carefully about how the window width trades smoothing bias against local variance. The goal is a mechanism you can reason from — not a magic number to copy.

What the Prediction Actually Computes

Start with notation, because every later claim depends on these symbols meaning something concrete.

You have a training set of nn pairs:

(x1,y1),(x2,y2),…,(xn,yn)(x_1, y_1), (x_2, y_2), \dots, (x_n, y_n)

Each xix_i is a feature vector, each yiy_i is a real-valued response. You want to predict at a query point x0x_0 — a point that may or may not be in the training set.

First you need a notion of "near." The default is Euclidean distance:

d(x0,xi)=∑j(x0,j−xi,j)2d(x_0, x_i) = \sqrt{\sum_{j} (x_{0,j} - x_{i,j})^2}

Then you sort the training points by that distance and take the kk closest. Call that set Nk(x0)N_k(x_0) — the neighborhood of x0x_0.

The prediction is the plain average of their responses:

f^(x0)=1k∑xi∈Nk(x0)yi\hat{f}(x_0) = \frac{1}{k} \sum_{x_i \in N_k(x_0)} y_i

That is the entire estimator. Read it once more, because the simplicity hides the design: the weights are 1/k1/k for points inside the neighborhood and 00 for points outside. It is a local averaging estimator with a hard-edged, uniform kernel. Every neighbor counts equally; every non-neighbor counts not at all.

Contrast this with a global linear fit. Linear regression learns one rule — a single set of coefficients — and applies it everywhere. KNN learns nothing in advance. It constructs a fresh rule for every query point, using only the data that happens to sit nearby. One rule for all of space versus a new rule per question.

Two structural choices hide inside the definition, and both matter later:

  • The metric. Euclidean is the default, not a law. Change the metric and you change which points count as neighbors.
  • The weighting scheme. Uniform weights are the simplest choice. Distance-weighted variants exist, but the uniform version is what makes the bias-variance analysis clean.

Knowledge check

Check your understanding

Answer this question before you continue.

For uniform-weight KNN regression, which rule gives the prediction at query point x₀?
Single Choice

Focus: Express the uniform-weight KNN prediction in terms of the responses in the selected neighborhood.

A Worked Example You Can Check by Hand

Small numbers make the mechanism visible. Take a one-dimensional dataset:

xix_i012345
yiy_i132546

Query at x0=2.6x_0 = 2.6. Compute distances:

xix_i231405
dd0.60.41.61.42.62.4
yiy_i253416

Sorted by distance: x=3x=3 (0.4), x=2x=2 (0.6), x=4x=4 (1.4), x=1x=1 (1.6), x=5x=5 (2.4), x=0x=0 (2.6).

Now compute the prediction at three values of kk:

  • k=1k=1: nearest neighbor is x=3x=3, so f^(2.6)=5\hat{f}(2.6) = 5.
  • k=3k=3: neighbors are x=3,2,4x=3, 2, 4, so f^(2.6)=(5+2+4)/3=3.67\hat{f}(2.6) = (5 + 2 + 4)/3 = 3.67.
  • k=5k=5: neighbors are x=3,2,4,1,5x=3, 2, 4, 1, 5, so f^(2.6)=(5+2+4+3+6)/5=4.0\hat{f}(2.6) = (5 + 2 + 4 + 3 + 6)/5 = 4.0.

Same query point, three different answers. Nothing about the data changed — only the window width.

Two boundary cases are worth naming. At k=1k=1, the prediction equals a single training response, so the fitted curve passes exactly through every training point. At k=nk=n, the neighborhood is the entire dataset and the prediction collapses to the global mean of yy — here, 3.53.5. The estimator loses all locality. Between those extremes, kk slides the prediction from "trust one noisy observation" to "trust nothing but the overall average."

Here is the same arithmetic in code, purely to confirm the hand calculation:

import numpy as np

x = np.array([0, 1, 2, 3, 4, 5], dtype=float)
y = np.array([1, 3, 2, 5, 4, 6], dtype=float)
x0 = 2.6

order = np.argsort(np.abs(x - x0))
for k in (1, 3, 5):
    print(k, y[order[:k]].mean())

Run it and you get 5.0, 3.666..., 4.0. The derivation and the output agree, which is the only reason to write the code at all.

Knowledge check

Check your understanding

Answer this question before you continue.

For the article’s query x₀ = 2.6 and k = 3, what prediction results from averaging the three nearest responses?
Output Prediction

Focus: Calculate a KNN regression prediction from the selected neighbors in the article’s hand-worked dataset.

Where the Average Comes From

The estimator looks arbitrary until you derive it from the quantity you actually want.

Regression, stated formally, is the problem of estimating the conditional mean:

f(x)=E[y∣x]f(x) = \mathbb{E}[y \mid x]

If you knew ff exactly, you would never need KNN. You do not know it, so you estimate it from samples. The natural estimate of a conditional mean is a sample mean — but only of samples that are conditioned on being near x0x_0. That is precisely what the neighborhood average does:

f^(x0)=1k∑xi∈Nk(x0)yi≈E[y∣x≈x0]\hat{f}(x_0) = \frac{1}{k} \sum_{x_i \in N_k(x_0)} y_i \approx \mathbb{E}[y \mid x \approx x_0]

The approximation sign is doing real work. It leans on two assumptions:

  1. Smoothness. The true regression function ff is roughly constant inside the neighborhood, so f(xi)≈f(x0)f(x_i) \approx f(x_0) for every neighbor.
  2. Locality. The neighborhood is small enough that its points are genuinely informative about x0x_0, not about some distant region.

Write each response as signal plus noise, yi=f(xi)+ϵiy_i = f(x_i) + \epsilon_i, with noise mean zero and variance σ2\sigma^2. Substitute into the estimator:

f^(x0)=1k∑xi∈Nk(x0)f(xi)+1k∑xi∈Nk(x0)ϵi\hat{f}(x_0) = \frac{1}{k} \sum_{x_i \in N_k(x_0)} f(x_i) + \frac{1}{k} \sum_{x_i \in N_k(x_0)} \epsilon_i

The first term is the average of the true function over the neighborhood. The second is the average of the noise. Now take expectations over the noise. The noise term averages to zero, so the bias is the gap between that first term and the truth:

Bias[f^(x0)]=1k∑xi∈Nk(x0)f(xi)−f(x0)\text{Bias}[\hat{f}(x_0)] = \frac{1}{k} \sum_{x_i \in N_k(x_0)} f(x_i) - f(x_0)

This is the average distance between ff at the neighbors and ff at the query point. If ff is flat across the neighborhood, the bias is zero. If ff is climbing or falling, the neighbors sit above or below x0x_0 on average, and the bias is nonzero. The smoothness assumption is what keeps this term small.

The variance comes from the noise term. Under the simplifying assumption that the noise values are independent across neighbors:

Var[f^(x0)]=σ2k\text{Var}[\hat{f}(x_0)] = \frac{\sigma^2}{k}

Average kk independent noise draws and the variance of the average shrinks by a factor of kk. More neighbors, less jitter.

Note: That σ2/k\sigma^2/k term is a clean idealization, not a promise. It is conditional on the neighbor locations staying fixed and the noise being independent and equally variable across neighbors. Real neighbors are not independent draws — nearby points often share unmeasured structure, so their noise is correlated. When noise is positively correlated across neighbors, the variance falls more slowly than 1/k1/k. Treat the formula as the direction of the effect, not a precise prediction.

Knowledge check

Check your understanding

Answer this question before you continue.

If neighbor locations are fixed and their noise values are independent with equal variance σ², what is the variance contributed by averaging k neighbor noises?
Single Choice

Focus: Identify the variance of the uniform neighborhood average under the article’s fixed-location and independent-noise assumptions.

Why k Is a Bandwidth, Not an Accuracy Dial

A schematic plot with neighborhood size k on the horizontal axis and relative error on the vertical axis. Squared bias rises, variance falls, and their combined error forms a typical U-shaped curve.
A wider neighborhood smooths predictions: variance tends to fall as bias grows, though the total-error curve is not guaranteed to be U-shaped.

Now the two terms can be read against each other, and the tradeoff appears as a consequence rather than a slogan.

Small kk. The neighborhood hugs x0x_0, so f(xi)≈f(x0)f(x_i) \approx f(x_0) and the bias term is tiny. But the prediction rides on very few noisy responses, and σ2/k\sigma^2/k is large. Resample the training set and the prediction at x0x_0 jumps around. Low bias, high variance.

Large kk. Many responses get averaged, so the variance term shrinks. But the neighborhood now spans regions where ff has already moved, so the bias term tends to grow. The prediction is stable and consistently wrong in curved regions. Low variance, high bias.

The two movements usually run in opposite directions:

  • Bias tends to rise with kk — a wider window averages over more of ff's variation.
  • Variance falls roughly as 1/k1/k — more terms in the average means less jitter.

Expected squared error at x0x_0 is the sum of a squared bias term, a variance term, and the irreducible noise σ2\sigma^2:

E[(y−f^(x0))2]=Bias2+σ2k+σ2\mathbb{E}\left[(y - \hat{f}(x_0))^2\right] = \text{Bias}^2 + \frac{\sigma^2}{k} + \sigma^2

One term tends to climb as kk grows. One term falls. Their sum often traces a U-shaped curve, and the minimum of that curve is the neighborhood size that balances the two for this particular problem.

Here is where the tidy story needs a boundary. The assumptions we listed — local smoothness and independent, mean-zero noise — do not guarantee that the bias term increases monotonically as you add neighbors. Adding a neighbor can move the average of ff closer to f(x0)f(x_0), farther from it, or across it, depending on the shape of ff, the geometry of the neighborhood, and the order in which points enter. What the assumptions support is a tendency, not a theorem: wider windows smooth more, and smoothing more usually costs bias in curved regions. The U-shape is a common and useful pattern, not a consequence the current derivation proves pointwise.

Common mistake: Treating the U-shape as a universal law of all models. It is a property of this estimator under these assumptions, and even here it is a tendency rather than a guarantee. The bias-variance tradeoff is well documented for KNN and kernel regression, but it is not guaranteed for every model class, and it should not be assumed without checking.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article’s usual local-smoothing picture, what is the typical tradeoff when k is small?
Comparison Reasoning

Focus: Compare the typical bias and variance effects of using a small neighborhood in KNN regression.

What Breaks the Clean Picture

The derivation above is tidy because it assumes a lot. Each assumption is a place where the analysis stops applying.

Feature scaling. Distances are computed across all features at once. If one feature ranges over thousands and another over single digits, the large-range feature dominates the distance calculation entirely. The "neighborhood" becomes a slice defined by one variable, and the local average is no longer local in any meaningful sense. Standardize features before trusting any of this.

The curse of dimensionality. In high dimensions, the nearest neighbors are barely closer to x0x_0 than randomly chosen points. When everything is roughly equidistant, "near" stops meaning near, the smoothness assumption fails, and the bias argument collapses. This is the single biggest reason KNN struggles as feature count grows.

Non-uniform data density. The same kk produces a wide neighborhood in sparse regions and a narrow one in dense regions. So the effective bandwidth varies across the input space even though kk is fixed. In sparse regions you get high bias; in dense regions you get high variance. One number, two different behaviors.

Boundary effects. Near the edges of the data, neighborhoods are one-sided — all neighbors lie to one side of x0x_0. The average is pulled inward, and predictions at the boundary are systematically biased toward the interior.

Ties and metric choice. Duplicate distances create ambiguity about which points enter the neighborhood. And switching from Euclidean to Manhattan distance changes the neighborhood itself, which changes every downstream number.

From Estimator Analysis to a Practical Choice

The theory tells you the shape of the tradeoff. It does not hand you a number.

The right kk depends on the sample size nn, the noise level σ2\sigma^2, the smoothness of ff, and the local density of your data. Change any of those and the optimal window width moves. That is why there is no universal best kk — any single recommended value is a heuristic, not a result. The same kk that underfits one dataset overfits another.

The practical procedure is straightforward: choose kk by cross-validated error under a fixed evaluation setup, then read the resulting curve as an empirical picture of the bias-variance tradeoff you just derived. A validation curve that dips and rises is the U-shape showing up in your actual data.

One caution before you read that curve too literally. The σ2/k\sigma^2/k result above held the neighbor locations fixed and varied only the noise. Cross-validation does something different: it refits the model on different training subsets, so the neighbor locations themselves change, and the variance you observe includes that resampling effect. The empirical curve and the clean formula are related, not identical. Use the formula for intuition and the curve for decisions.

My rule is to start moderate, inspect the validation curve, and treat kk as a smoothing parameter I can justify rather than a default I inherited. If the curve is flat across a range of kk, the problem is insensitive to the choice and you can pick for convenience. If it has a sharp minimum, the choice matters and you should report how you made it.

This is estimator-level reasoning, not a benchmarking exercise. The goal here is to understand the mechanism — why the average behaves the way it does — not to produce a leaderboard number. Tuning KNN for a specific classification task is a separate workflow with its own evaluation discipline.

The next move is concrete. Take a small one-dimensional dataset, fit KNN regression at k=1,5,25k = 1, 5, 25, and plot all three prediction curves over the same scatter. Watch the k=1k=1 curve thread through every point, watch the k=25k=25 curve flatten, and watch the middle one trade a little of each. The tradeoff stops being a slogan the moment you see it as a shape — and the number you eventually choose will be one you earned through validation, not one you copied from a tutorial.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Why does the article caution that cross-validation variability is related to, but not identical to, the σ²/k calculation?
Question 1 of 2Comparison Reasoning

Focus: Distinguish the fixed-neighbor noise variance calculation from variance observed through cross-validation refits.

A KNN model uses the same k in both a sparse and a dense region. According to the article, which interpretation best describes the resulting neighborhood widths and typical effects?
Question 2 of 2Scenario Interpretation

Focus: Explain how non-uniform data density can make a fixed k behave differently across input regions.

References

  1. Université de Montréal On the Bias-Variance Tradeoffarxiv.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.