Skip to content
intermediate

Compare Classical Anomaly Detectors With a Scikit-Learn Experiment

Two detectors, one dataset, two different alert lists — and no labels to referee the disagreement.

Published 2026-10-02Updated 2026-10-049 min read
Detailed view of a video editing software interface showing multi-track timeline and colorful design.
Detailed view of a video editing software interface showing multi-track timeline and colorful design. Photo by Francesco Paggiaro on Pexels.

Two detectors, one dataset, two different alert lists — and no labels to referee the disagreement.

That is the normal condition of anomaly detection. You do not get a clean answer key telling you which flagged point was truly anomalous. You get a scoring function, a threshold, and a decision you have to defend. So the useful skill is not "pick the best algorithm." It is building a small harness where you can see why two detectors disagree, then explaining that disagreement from their assumptions rather than from vibes.

This tutorial builds that harness. You will generate a seeded 2D dataset with a known injected-anomaly mask, split it before fitting, run IsolationForest and LocalOutlierFactor (with novelty=True) on the training portion, and compare what each one flags on held-out data. By the end you will have a settings table, flagged-point plots, an overlap count, and a defensible sentence about the split.

What This Experiment Can and Cannot Prove

Before any code runs, set the boundary.

The synthetic mask is a diagnostic instrument, not ground truth about the world. It tells you whether a detector recovered the points you deliberately injected. It says nothing about whether a detector would find the anomalies that matter in your real data, because in real data you rarely know which points those are.

That gap is the whole reason anomaly detection leans on unsupervised scoring and human review. Precision and recall against an injected mask are useful for debugging your harness. They are not evidence of production performance.

Note: If you already understand outlier detection versus novelty detection and why splitting exists, skip ahead. This article assumes that background and does not re-teach it.

Success criteria for this experiment:

  • Both estimators fit on the training portion.
  • Both produce inspectable flagged cases on held-out data.
  • Every difference between them traces back to an assumption or a setting you can name.

Knowledge check

Check your understanding

Answer this question before you continue.

What can precision and recall against the injected-anomaly mask establish in this experiment?
Misconception Check

Focus: Distinguish post-hoc diagnostic metrics on injected labels from evidence about real-world anomaly performance.

Build the Seeded Dataset and Split Before Fitting

Start with a deterministic dataset. Two dimensions, because you can plot the decision surface and the flagged points — the fastest way to see what a detector is actually doing.

import numpy as np
from sklearn.model_selection import train_test_split

rng = np.random.RandomState(42)

n_inliers = 300
n_anomalies = 20

inliers = rng.normal(loc=[0.0, 0.0], scale=[1.0, 1.0], size=(n_inliers, 2))
anomalies = rng.normal(loc=[5.0, 5.0], scale=[0.5, 0.5], size=(n_anomalies, 2))

X = np.vstack([inliers, anomalies])
mask = np.array([0] * n_inliers + [1] * n_anomalies)

X_train, X_test, mask_train, mask_test = train_test_split(
    X, mask, test_size=0.3, random_state=42, stratify=mask
)

print("train shape:", X_train.shape, "test shape:", X_test.shape)
print("injected in test:", mask_test.sum())

The mask rides through the split so it stays aligned with rows. You will use it only after predictions exist, never as a training signal.

Common mistake: Injecting anomalies, fitting on everything, then "evaluating" on the same rows. The mask becomes a training signal and the numbers turn into decoration. Split first, always.

Knowledge check

Check your understanding

Answer this question before you continue.

A learner fits a detector on all generated rows, then reports its flags against the injected mask on those same rows. What change best fixes this evaluation design?
Scenario Interpretation

Focus: Choose a train/test workflow that keeps the injected mask from becoming a training signal.

Fit IsolationForest and LocalOutlierFactor on the Training Portion

Now fit both detectors. The configuration matters more than the algorithm name.

from sklearn.ensemble import IsolationForest
from sklearn.neighbors import LocalOutlierFactor
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

iso = make_pipeline(
    StandardScaler(),
    IsolationForest(n_estimators=100, contamination=0.05, random_state=42),
)
iso.fit(X_train)

lof = make_pipeline(
    StandardScaler(),
    LocalOutlierFactor(n_neighbors=20, contamination=0.05, novelty=True),
)
lof.fit(X_train)

iso_labels = iso.predict(X_test)
iso_scores = iso.decision_function(X_test)

lof_labels = lof.predict(X_test)
lof_scores = lof.decision_function(X_test)

print("IF alerts:", (iso_labels == -1).sum())
print("LOF alerts:", (lof_labels == -1).sum())

Two things to notice.

IsolationForest isolates points by randomly partitioning the feature space. Points that get isolated in fewer splits are more anomalous. No density model is assumed — just depth in a forest of random trees.

LocalOutlierFactor compares each point's local density to its neighbors' densities. With novelty=True, the estimator unlocks predict and decision_function on new data. Without that flag, you only get fit_predict on the training set, which is a different experiment.

Keep the raw scores, not just the labels. decision_function is what lets you rank and compare detectors instead of only counting flags.

Knowledge check

Check your understanding

Answer this question before you continue.

A pipeline fits LocalOutlierFactor on `X_train`, but calling `predict(X_test)` fails. Which change matches the article's held-out novelty-detection workflow?
Debugging

Focus: Configure LocalOutlierFactor for predictions and decision scores on held-out data.

Debug the Three Failures That Break This Experiment

Most beginner confusion here is a broken harness, not a broken algorithm. Three failures account for nearly all of it.

Scaling. LOF is distance-based. Unscaled features let one axis dominate the neighborhood structure, and the detector quietly answers a different question than you asked. Put scaling in a pipeline and fit it on the training portion only — otherwise you leak test-set statistics into the transform.

Novelty mode. Calling predict on an LOF fitted without novelty=True raises an error. Calling fit_predict on held-out data is not the same experiment — it refits on the test set. If you see that error, you forgot the flag.

Score direction. Both IsolationForest.decision_function and LocalOutlierFactor.decision_function follow the scikit-learn convention: lower means more anomalous. But negative_outlier_factor_ runs the other way. Mixing them silently inverts your ranking, and the top of your list becomes the bottom.

Tip: Sort by score, print the top few rows, and confirm they look like the points you injected. If they do not, suspect direction or scaling before suspecting the algorithm.

Knowledge check

Check your understanding

Answer this question before you continue.

A learner ranks `decision_function` values from highest to lowest and calls the first rows most anomalous. According to the article, what is the correction?
Debugging

Focus: Keep score rankings consistent with the documented direction of the detector scores.

Compare Detectors: Settings Table, Flagged Points, and Overlap

Side-by-side schematic views of the same point cloud. IsolationForest highlights a point isolated far from the main group; LocalOutlierFactor highlights a point in a locally sparse pocket. A small shared-alert marker indicates that some points may be flagged by both.
Equal alert counts can hide different choices: compare which points each detector flags and consider the geometric assumption behind the disagreement.

Turn the two fitted estimators into a comparison you can defend in a conversation.

DetectorKey settingsHeld-out alerts
IsolationForestn_estimators=100, contamination=0.05, random_state=425
LocalOutlierFactorn_neighbors=20, contamination=0.05, novelty=True5

Both detectors flag 5 of the 96 held-out points, which is what contamination=0.05 asks for. The count is identical; the membership is not. That is the point of the next step.

Plot the held-out points with the injected mask drawn as a separate marker. You want to see which alerts land on injected points and which land on ordinary inliers.

import matplotlib.pyplot as plt

fig, axes = plt.subplots(1, 2, figsize=(12, 5))
for ax, labels, title in [
    (axes[0], iso_labels, "IsolationForest"),
    (axes[1], lof_labels, "LocalOutlierFactor"),
]:
    ax.scatter(X_test[:, 0], X_test[:, 1], c="lightgray", label="held-out")
    ax.scatter(
        X_test[mask_test == 1, 0], X_test[mask_test == 1, 1],
        facecolors="none", edgecolors="blue", s=120, label="injected",
    )
    ax.scatter(
        X_test[labels == -1, 0], X_test[labels == -1, 1],
        c="red", marker="x", s=80, label="flagged",
    )
    ax.set_title(title)
    ax.legend()
plt.show()

Now count the overlap.

iso_set = set(np.where(iso_labels == -1)[0])
lof_set = set(np.where(lof_labels == -1)[0])

print("both flagged:", len(iso_set & lof_set))
print("only IF:", len(iso_set - lof_set))
print("only LOF:", len(lof_set - iso_set))

On this run, the two detectors agree on 3 points, IsolationForest flags 2 that LOF misses, and LOF flags 2 that IsolationForest misses. The disagreement set is the interesting part, not the agreement. When the detectors disagree, ask whether the point is far from everything (isolation-friendly) or merely in a locally sparse pocket (density-friendly). That question usually explains the split.

Precision and recall against the mask are a limited diagnostic. Report them, then immediately say what they do not license you to claim — namely, anything about real anomalies.

from sklearn.metrics import precision_score, recall_score

for name, labels in [("IF", iso_labels), ("LOF", lof_labels)]:
    preds = (labels == -1).astype(int)
    print(name, "precision:", precision_score(mask_test, preds, zero_division=0),
          "recall:", recall_score(mask_test, preds, zero_division=0))

Read the two numbers together. A detector can flag the same count as its rival and still score differently, because precision and recall depend on which points it chose, not how many. If IsolationForest catches more injected points but also spends more flags on ordinary inliers, its recall rises and its precision falls. Neither number is a verdict. They are a description of the trade the detector made on this one synthetic split.

Change Contamination and Watch the Alert Set Move

Contamination is a threshold decision with a visible cost, not a parameter to tune until the plot looks nice.

Hold the data, the split, and the random_state fixed. Change only contamination on both estimators. Predict the direction before running: raising contamination should enlarge the alert set, and the newly added alerts should be the next-lowest-scoring points — the ones closest to the current threshold on the anomalous side.

for c in [0.02, 0.05, 0.10]:
    iso = make_pipeline(
        StandardScaler(),
        IsolationForest(n_estimators=100, contamination=c, random_state=42),
    ).fit(X_train)
    lof = make_pipeline(
        StandardScaler(),
        LocalOutlierFactor(n_neighbors=20, contamination=c, novelty=True),
    ).fit(X_train)
    print(c, "IF alerts:", (iso.predict(X_test) == -1).sum(),
          "LOF alerts:", (lof.predict(X_test) == -1).sum())

Watch precision and recall move in opposite directions against the mask. That trade is a property of your threshold, not evidence that one detector is better.

Warning: On a synthetic dataset with a known injection rate, contamination is easy to set correctly. In production you rarely know it. Set contamination from a defensible estimate of the real alert rate or from the review capacity you actually have — not from the default.

What to Carry Into a Real Dataset

Keep the harness. Seeded data generation, split-before-fit, a settings table, and an overlap count are reusable scaffolding for any detector comparison.

Drop the mask. On real data you will not have injected labels, so the comparison shifts to score agreement, stability across seeds, and whether flagged cases survive a domain review.

Treat disagreement as information. Two detectors flagging different points is a signal about geometry and assumptions. It is often more useful than a single averaged score, because it tells you which structural property of the data is driving the decision.

The honest output of this experiment is not "which algorithm won." It is a ranked alert list, a stated contamination assumption, and a written explanation of why the two detectors disagreed.

Your next move: add a third detector or a second dataset shape — multimodal, or with differing densities — and test whether your explanation of the disagreement still holds. If it does, you have a mental model. If it does not, you have found the boundary where the model stops working, which is where the real learning starts.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

On the held-out set, two detectors flag the same number of points but disagree on some of their membership. Which interpretation best follows the article?
Question 1 of 2Comparison Reasoning

Focus: Explain detector disagreement by relating flagged cases to IsolationForest and LOF assumptions.

The data, split, and random state are held fixed while contamination is raised. What should the learner expect and how should changes in precision and recall be interpreted?
Question 2 of 2Scenario Interpretation

Focus: Predict how changing contamination affects alert counts and interpret resulting metric changes cautiously.

References

  1. 2.7. Novelty and Outlier Detection — scikit-learn 1.9.1 documentationscikit-learn.org
  2. Comparing anomaly detection algorithms for outlier detection on toy datasets — scikit-learn 1.9.0 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.