Skip to content
intermediate

Test Whether an Interaction Feature Helps a Model

Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset…

Published 2026-10-02Updated 2026-10-048 min read
Students and teachers in a classroom standing and saluting with flags displayed.
Students and teachers in a classroom standing and saluting with flags displayed. Photo by HONG SON on Pexels.

Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset where no interaction existed. The score still went up. That is the trap: a single number cannot tell you whether you found signal or just got lucky.

So we are going to build a test that can tell those two situations apart. The plan is a controlled experiment. You generate data where you know whether an interaction exists, then check whether your workflow detects it when present and stays quiet when absent.

Why a Score Bump Is Not Proof

When a score improves after you add a feature, there are at least three explanations:

  • Real signal. The feature captures a joint effect the model could not represent before.
  • Noise fitting. The model latched onto random structure that will not repeat.
  • Evaluation variance. You got a favorable split, and the same model would score differently on another sample.

A single train/test split cannot separate these. The only way to separate them is to control the data-generating process so you know the ground truth before you fit anything.

That gives us a clear success criterion. The interaction model should win on data with a real interaction and not reliably win on data without one. If your workflow passes both checks, it is measuring something real. If it fails either one, you are reading noise.

We compare two models: an additive baseline using only the main effects, and the same model plus an explicit product feature.

Prerequisites and Setup

You should already know what an interaction is and have run a leakage-safe validation workflow before. If the concept is fuzzy, the mental model is simple: an additive model cannot represent a joint effect, because it treats each feature's contribution as independent. That is exactly why we test instead of assume.

You need Python with NumPy, pandas, and scikit-learn. No external services, no credentials. Fix the random seed so every number here reproduces on your machine.

import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer

RANDOM_STATE = 42

Build a Dataset With a Known Interaction

We define the target as a sum of main effects plus a product term. The coefficient on the product is the knob that controls interaction strength.

def make_data(n=500, interaction_coef=2.0, noise=1.0, seed=42):
    rng = np.random.default_rng(seed)
    x1 = rng.normal(0, 1, n)
    x2 = rng.normal(0, 1, n)
    y = 1.5 * x1 + 0.8 * x2 + interaction_coef * (x1 * x2) + rng.normal(0, noise, n)
    return pd.DataFrame({"x1": x1, "x2": x2, "y": y})

df = make_data(interaction_coef=2.0)
print(df.head())

You get a small frame with x1, x2, and a target whose values depend on both features jointly. When interaction_coef is zero, the target is purely additive.

Knowledge check

Check your understanding

Answer this question before you continue.

In the article's data generator, what does setting `interaction_coef=0.0` make the target relationship?
Single Choice

Focus: Infer how setting the interaction coefficient to zero changes the data-generating relationship.

Fit the Additive Baseline

Split before any fitting, then train a linear model on x1 and x2 only. By construction, this model family cannot represent the product term.

X = df[["x1", "x2"]]
y = df["y"]

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=RANDOM_STATE
)

baseline = LinearRegression().fit(X_train, y_train)
print("baseline R2:", baseline.score(X_test, y_test))

The baseline is misspecified for this target. It is not a bad model; it is the wrong shape for a target that bends with the product of two features. Expect a score clearly below the interaction model.

Add the Interaction Feature

The product feature is a fixed row-wise calculation: x1 * x2 for each row, with no statistics learned from the data. Computing it before the split does not leak anything, because no row's value depends on any other row. That distinction matters, and it is worth stating plainly because the leakage rule gets applied too broadly.

Where the pipeline earns its keep is workflow integrity. It bundles the representation with the estimator so the two travel together, and it protects you the moment you add a preprocessing step that does learn from data, such as scaling or target encoding. Keep the transform inside the pipeline as a habit, not because this particular product needs it.

def add_product(X):
    X = np.asarray(X)
    return np.hstack([X, (X[:, 0] * X[:, 1]).reshape(-1, 1)])

interaction_model = Pipeline([
    ("product", FunctionTransformer(add_product)),
    ("model", LinearRegression()),
])

scores = cross_val_score(interaction_model, X_train, y_train, cv=5, scoring="r2")
print("interaction CV R2:", scores.mean(), "+/-", scores.std())

Common mistake: Assuming every transformation outside a pipeline leaks. A stateless row-wise product does not. A transformation that estimates parameters from the training data — a scaler's mean, a target encoder's category statistics — does, and that is the case the pipeline rule exists to prevent.

Knowledge check

Check your understanding

Answer this question before you continue.

Which statement best matches the article's explanation of the product feature and leakage?
Misconception Check

Focus: Distinguish a stateless row-wise interaction calculation from transformations that learn parameters across training rows.

Compare Both Models on the Same Folds

A two-row comparison shows a positive interaction-minus-baseline score gap for data with a real interaction and a near-zero gap for the no-interaction control; both comparisons use the same cross-validation folds.
A useful interaction feature should help when the joint effect exists, but not reliably help on the no-interaction control.

The baseline above was scored on a held-out test set, while the interaction model was scored by cross-validation on the training set. Those are different samples and different procedures, so the gap between them is not a clean estimate of the feature's benefit. Fix that before drawing any conclusion.

Score both models with the same cross-validation call on the same data:

base_cv = cross_val_score(baseline, X_train, y_train, cv=5, scoring="r2")
inter_cv = cross_val_score(interaction_model, X_train, y_train, cv=5, scoring="r2")

print("baseline CV R2:   ", base_cv.mean(), "+/-", base_cv.std())
print("interaction CV R2:", inter_cv.mean(), "+/-", inter_cv.std())
print("gap:", inter_cv.mean() - base_cv.mean())

Now the two numbers come from the same folds, the same scoring function, and the same rows. The gap is comparable. You want it to be larger than the fold-to-fold variation in either model, not merely positive.

Knowledge check

Check your understanding

Answer this question before you continue.

The baseline score comes from a held-out test set, while the interaction score comes from cross-validation on training rows. What should you change before interpreting their gap?
Debugging

Focus: Choose a comparable evaluation setup for estimating the interaction feature's score change.

Run the Same Test on Data With No Interaction

This is the control. Regenerate the data with interaction_coef=0.0 and keep everything else identical, including the evaluation procedure.

df_null = make_data(interaction_coef=0.0)
X_null = df_null[["x1", "x2"]]
y_null = df_null["y"]

X_tr, X_te, y_tr, y_te = train_test_split(
    X_null, y_null, test_size=0.2, random_state=RANDOM_STATE
)

base_null = LinearRegression().fit(X_tr, y_tr)
base_null_cv = cross_val_score(base_null, X_tr, y_tr, cv=5, scoring="r2")
inter_null_cv = cross_val_score(interaction_model, X_tr, y_tr, cv=5, scoring="r2")

print("baseline CV R2:   ", base_null_cv.mean(), "+/-", base_null_cv.std())
print("interaction CV R2:", inter_null_cv.mean(), "+/-", inter_null_cv.std())
print("gap:", inter_null_cv.mean() - base_null_cv.mean())

Here the two models should perform about the same, and any difference should sit inside the noise band. This is a check, not proof. One run on one seed is a single sample of a noisy process. To treat it as evidence, repeat the whole experiment across several seeds and confirm the null gap stays near zero while the real-interaction gap stays positive.

Knowledge check

Check your understanding

Answer this question before you continue.

On data generated with `interaction_coef=0.0`, the interaction model's cross-validation score is slightly higher than the baseline's. Which interpretation is most consistent with the article?
Scenario Interpretation

Focus: Interpret the expected score relationship when the interaction is absent while accounting for noisy evaluation.

Read the Result Without Overclaiming

A win on one controlled dataset is evidence about that data-generating process, not a universal rule about interaction features. Keep three distinctions in mind:

  • Statistical vs practical significance. A small but consistent gain may not justify the added complexity or the loss of interpretability.
  • Scope. Results depend on sample size, noise level, model family, and how the interaction is encoded.
  • Mechanism. A linear model needs the explicit product. A tree-based model can often capture the same interaction without one.

My decision rule: only add an explicit interaction when the gain survives both the control dataset and a stability check across folds and seeds.

Failure Modes and Debugging Signals

SymptomLikely cause
Score gap vanishes when you change the seedThe effect was noise
Interaction model wins on both datasetsLeakage or an evaluation bug
Both models score near chanceCheck target construction and feature scaling
Folds disagree wildlyDataset too small or too noisy for a stable conclusion

When a gap disappears under a new seed, that is not a failure of your experiment. It is the experiment doing its job.

One Modification to Try Next

Sweep the interaction coefficient from zero upward and plot the score gap against interaction strength. Use the same cross-validation procedure for both models at every value.

for coef in [0.0, 0.5, 1.0, 2.0, 4.0]:
    d = make_data(interaction_coef=coef)
    Xd, yd = d[["x1", "x2"]], d["y"]
    Xtr, Xte, ytr, yte = train_test_split(Xd, yd, test_size=0.2, random_state=RANDOM_STATE)
    base = cross_val_score(LinearRegression(), Xtr, ytr, cv=5, scoring="r2").mean()
    inter = cross_val_score(interaction_model, Xtr, ytr, cv=5, scoring="r2").mean()
    print(f"coef={coef}: baseline={base:.3f} interaction={inter:.3f} gap={inter-base:.3f}")

You will see a threshold where the interaction model starts to reliably win. That turns a single yes/no result into a curve showing how detectability depends on effect size and noise. As an extension, swap the linear model for a tree-based model and note that it can capture the interaction without an explicit product feature.

The reusable rule is this: never trust one score change. Run the controlled test, check the no-interaction control, verify stability across folds and seeds, and only then decide whether the added representation earns its place. Your next step is to run this same workflow inside a real pipeline on your own data, where you do not know the ground truth and the control dataset has to be your discipline instead.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A product feature improves the score on one controlled dataset, but its gain disappears on other seeds and the null-control runs also show gains. What is the best conclusion?
Question 1 of 2Comparison Reasoning

Focus: Apply the article's decision rule for adopting an explicit interaction representation.

The interaction model's score gap vanishes when you change the random seed. Which diagnosis best follows the article's debugging guide?
Question 2 of 2Scenario Interpretation

Focus: Use instability across seeds to diagnose whether an apparent interaction gain may be noise.

References

  1. Partial Dependence and Individual Conditional Expectation Plots — scikit-learn 1.9.0 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.