Skip to content
beginner

Fit and Inspect Logistic Regression With Scikit-Learn

You call predict(), and you get back a tidy row of zeros and ones. Clean. Decisive. And completely silent about the thing you actually wanted to see: the…

Published 2026-10-02Updated 2026-10-047 min read
Captivating aerial view of deep blue ocean waves, showcasing nature's dynamic patterns and vibrant colors.
Captivating aerial view of deep blue ocean waves, showcasing nature's dynamic patterns and vibrant colors. Photo by Diego F. Parra on Pexels.

You call predict(), and you get back a tidy row of zeros and ones. Clean. Decisive. And completely silent about the thing you actually wanted to see: the probability.

That silence is where beginners get stuck. The model didn't throw away the probability — it computed one, then applied a rule you never wrote. In this tutorial, you'll fit a logistic-regression classifier, pull the probabilities out into the open, watch them turn into labels, and then move the decision line yourself to see what changes.

If you've read the concept piece on how a linear score becomes a probability through the logistic function, this is the practice half. We won't re-derive the sigmoid. We'll make it run.

What You Need Before You Fit Anything

You need Python 3.8 or newer, plus scikit-learn, NumPy, and pandas. Scikit-learn depends on NumPy and SciPy, so installing it pulls those in. If you're starting fresh:

pip install scikit-learn pandas

We'll use a small built-in dataset so the whole experiment runs in seconds. The point is the model, not the data wrangling — I'd rather you spend your attention on probabilities than on CSV parsing.

Here's the success criterion. By the end of this article you will have:

  • a fitted classifier,
  • a table showing true labels next to predicted probabilities,
  • and two different confusion matrices produced by two different thresholds.

That's it. Three artifacts, all yours.

Note: A quick bridge from the concept article. Logistic regression computes a linear score, pushes it through the logistic function, and gets a number between 0 and 1. That number is a probability. Everything below is about what happens after that number exists.

Split the Data and Fit the Classifier

Load the data and split it. Fix random_state so your output matches mine.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression

X, y = load_breast_cancer(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=42, stratify=y
)

model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)

Two things worth naming.

First, fit() is where learning happens. The model searches for coefficients that minimize log loss — the penalty for assigning low probability to the correct class. This is not least squares. Linear regression minimizes squared error; logistic regression minimizes a different objective because its output is a probability, not a number on a line.

Second, scikit-learn applies L2 regularization by default. That surprises people coming from statistics, where unregularized logistic regression is the norm. In machine learning workflows, regularization is standard because it improves numerical stability and helps generalization. You can control its strength with the C parameter, but don't tune it yet. One thing at a time.

Now sanity-check the fit:

print(model.coef_.shape)       # (1, 30)
print(model.intercept_.shape)  # (1,)

One coefficient per feature, one intercept. If those shapes look wrong, something upstream is off.

Warning: You may see a ConvergenceWarning about max_iter. It means the optimizer hit its iteration limit before settling. It does not mean the model is broken. Raising max_iter is the first fix, and it's usually the only one you need.

Knowledge check

Check your understanding

Answer this question before you continue.

A fit raises a `ConvergenceWarning` about `max_iter`. According to the tutorial, what should you try first?
Debugging

Focus: Interpret a convergence warning from a fitted scikit-learn logistic-regression model and identify the article's first suggested fix.

Read the Probabilities, Not Just the Labels

Here's the part that predict() hides. predict_proba() returns one column per class. Column 0 is P(class = 0); column 1 is P(class = 1). You almost always want the second column:

proba = model.predict_proba(X_test)[:, 1]
preds = model.predict(X_test)

Now build the table you'll keep. This is the artifact that makes the model legible:

import pandas as pd

inspect = pd.DataFrame({
    "true": y_test,
    "proba": proba.round(3),
    "pred": preds,
})
print(inspect.head(10))

You'll see rows where the true label is 1, the probability is 0.51, and the prediction is 1 — and rows where the probability is 0.99 and the prediction is also 1. Same label. Wildly different confidence.

That gap is the whole story. predict() is just predict_proba() plus a 0.5 cutoff. The label is a decision layered on top of the probability, not a separate model output. When you see 0.51 and 0.99 collapse into the same "1," you're watching the threshold do its work — and quietly throwing away information.

Common mistake: Reading predict_proba output as a calibrated probability of being correct. It isn't. It's a model score. A 0.9 doesn't guarantee a 90% hit rate. Calibration is a separate question with its own tools, and we're not touching it here.

Knowledge check

Check your understanding

Answer this question before you continue.

For this binary classifier, which expression selects the probabilities for class 1 from `predict_proba(X_test)`?
Single Choice

Focus: Identify which column of `predict_proba()` contains the probability for class 1 in the tutorial's binary classification example.

Turn Probabilities Into Decisions With a Threshold

The default 0.5 is a convention, not a law. It's the right default only when both kinds of error cost roughly the same. Often they don't.

Converting probabilities to labels by hand takes one line:

threshold = 0.5
manual_preds = (proba >= threshold).astype(int)

Three tokens of code, full control. Compare manual_preds to preds and they'll match — that's the default threshold, written out loud.

Why would you ever move it? Consider two scenarios with the same model:

ScenarioWorse mistakeSensible threshold
Spam filterMarking real mail as spamHigher (be conservative)
Disease screeningMissing a sick patientLower (catch more)

Same probabilities. Different costs. Different decisions. The model doesn't know which world it's in — you do.

This connects directly to precision and recall. Lower the threshold and you flag more positives: recall rises, precision falls. Raise it and the reverse happens. The threshold is a business decision expressed in code.

Knowledge check

Check your understanding

Answer this question before you continue.

In a disease-screening scenario where missing a sick patient is the worse mistake, which change best matches the tutorial's guidance?
Scenario Interpretation

Focus: Choose the threshold direction that prioritizes catching more positive cases when missing a sick patient is the more costly error.

Run the Threshold Experiment

A three-row comparison shows a lower-than-0.5 threshold with false positives increasing and false negatives decreasing, 0.5 as the baseline, and a higher-than-0.5 threshold with false positives decreasing and false negatives increasing.
Keep the model fixed and move the threshold to see how the two error types trade off.

Now the controlled experiment. Hold the model fixed. Change only the threshold. That's what makes it an experiment instead of a fishing trip.

Before you run it, predict the outcome. Lowering the threshold should raise recall and lower precision. Raising it should do the reverse. State that prediction, then check it against the output.

from sklearn.metrics import confusion_matrix, precision_score, recall_score

for t in [0.3, 0.5, 0.7]:
    labels = (proba >= t).astype(int)
    cm = confusion_matrix(y_test, labels)
    print(f"threshold={t}")
    print(cm)
    print(f"  precision={precision_score(y_test, labels):.3f}")
    print(f"  recall={recall_score(y_test, labels):.3f}")

The confusion matrix is the clearest view of what changed. Read it as counts of false positives and false negatives. At 0.3 you'll see more false positives and fewer false negatives. At 0.7, the reverse. The precision and recall numbers just summarize what the matrix already shows.

Debugging signal: If nothing changes across thresholds, your probabilities are probably clustered near 0 and 1. That's worth noticing on its own — it means the model is very confident, and the threshold has little room to maneuver.

Knowledge check

Check your understanding

Answer this question before you continue.

With the model and probabilities held fixed, what does the tutorial say to expect when the threshold is lowered from 0.5 to 0.3?
Output Prediction

Focus: Predict how lowering the classification threshold changes positive predictions and the associated error counts in the tutorial's fixed-model experiment.

Pick a Threshold You Can Defend

Start from the cost of each error type, not from a target accuracy number. Ask which mistake is worse, and by how much. That ratio tells you which direction to move.

Then use the validation set to choose the threshold, and report on the test set once. If you tune the threshold on the test set, you've quietly turned it into training data — the same trap the train/validation/test article warns about, just wearing a different hat.

Two boundaries worth respecting:

  • Threshold tuning cannot fix a model whose probabilities are uninformative. If the ranking is bad, no cutoff saves it. You're just choosing which way to be wrong.
  • If the downstream system consumes the probability directly — a ranking, a risk score, a queue — you may never need a hard label at all. Sometimes the best threshold is no threshold.

Where to Take This Next

One modification to try right now: add a second feature, or swap in your own CSV, and rerun the same threshold loop. Watch how the probability distribution shifts and whether the threshold still has room to move.

The natural next step is wrapping the scaler and classifier into a Pipeline, so the same preprocessing is applied consistently at fit and predict time. That's where this stops being a script and starts being a reusable component.

But the durable idea is simpler. The model produces a ranking. The threshold turns that ranking into a decision. And the decision belongs to you — not to the 0.5 that scikit-learn quietly picked on your behalf.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Which workflow follows the tutorial's advice for threshold selection and final evaluation?
Question 1 of 2Comparison Reasoning

Focus: Distinguish the dataset used to tune a decision threshold from the dataset used for final reporting.

A downstream system uses the model output as a risk score to rank cases, rather than requiring a yes-or-no label. What does the tutorial suggest?
Question 2 of 2Misconception Check

Focus: Recognize when a downstream use may consume predicted probabilities directly rather than requiring a thresholded class label.

References

  1. 1.1. Linear Models — scikit-learn 1.9.1 documentationscikit-learn.org
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.