Skip to content
beginner

Feature Extraction: Turn Raw Text and Records Into Model-Ready Features

You have a list of sentences, a folder of images, or a pile of Python dictionaries. You call .fit() on your model, and it refuses. The error message is not…

Published 2026-10-02Updated 2026-10-049 min read
Dynamic abstract photograph of blue light patterns with a dark backdrop, creating a sense of motion and mystery.
Dynamic abstract photograph of blue light patterns with a dark backdrop, creating a sense of motion and mystery. Photo by Aedrian Salazar on Pexels.

You have a list of sentences, a folder of images, or a pile of Python dictionaries. You call .fit() on your model, and it refuses. The error message is not subtle: the estimator wants numbers, and you handed it words.

That rejection is not a formatting nuisance. It is the moment you discover that classical machine learning models do not see text, images, or records. They see rows of numbers. Feature extraction is the translation step that makes your raw input visible to them.

Why Estimators Refuse Your Raw Data

Most classical estimators in scikit-learn operate on a two-dimensional array with shape (n_samples, n_features). Each row is one example. Each column is one measurable property. Every cell holds a number.

Your raw data rarely arrives in that shape. A sentence has seven words; the next has forty. A dictionary has different keys per record. An image is a grid of pixels with spatial structure. None of these fit a fixed-width numeric grid.

The mismatch is structural. The model has no slot to put a word in. It cannot compute a distance between "dog" and "cat" because neither is a number. It cannot multiply a pixel by a weight if the pixel is still a file path.

You have already met the feature matrix when thinking about sparse data and feature space geometry. Now you are seeing it from the input side: before any model can learn, something must produce that matrix.

Pass a list of strings to a scikit-learn estimator and you get an error. Read the error. It is telling you exactly what is missing: a numerical feature vector for each sample.

Knowledge check

Check your understanding

Answer this question before you continue.

A classical estimator rejects a list of sentences. Which explanation best matches the input problem described in the article?
Misconception Check

Focus: Explain why raw non-numeric inputs must be represented as a fixed-width numerical feature matrix before a classical estimator can use them.

What Feature Extraction Actually Does

A text sample and a dictionary record flow through feature extraction into a numeric matrix. The matrix shows samples as rows and extracted features as columns, with numeric values in its cells.
Feature extraction turns each raw sample into a row in one consistent numeric feature matrix.

Feature extraction maps each raw object to a vector of numbers. The mapping must be the same for every sample. That consistency is what makes the output a matrix instead of a collection of unrelated arrays.

The output has rows and columns. Rows are samples. Columns are extracted features. The number of columns is fixed after extraction, even if the raw inputs had wildly different lengths.

Here is the part that surprises beginners: the feature set is learned from data. When you fit a vectorizer, it builds a vocabulary from the training corpus. That vocabulary becomes a learned parameter. At transform time, new documents are translated using the same vocabulary.

Think of it as a translator with a fixed dictionary built from the training corpus. New documents get translated with the same dictionary. If a word is not in the dictionary, it is either ignored or mapped to a special "unknown" token, depending on the vectorizer.

Where the analogy stops: unlike a translator, extraction can silently drop information. Word order, context, and spatial relationships may vanish. That loss is a design choice, not a bug. You are trading information for a representation the model can actually use.

Text to Numbers: The Bag-of-Words Idea

The most common extraction case is text. You have documents, and you need numerical features. The bag-of-words approach does three things: tokenize, count, and normalize.

Tokenization splits text into words or tokens. Counting records how often each token appears in each document. Normalization adjusts counts so that document length or common terms do not dominate.

CountVectorizer builds a vocabulary from the corpus and counts occurrences per document. The column order is the vocabulary order. Here is a small example:

from sklearn.feature_extraction.text import CountVectorizer

corpus = [
    "the dog bites the man",
    "the man bites the dog",
    "the cat sleeps"
]

vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)

print(vectorizer.get_feature_names_out())
print(X.toarray())
print(X.shape)

The vocabulary might be ['bites', 'cat', 'dog', 'man', 'sleeps', 'the']. The matrix has three rows and six columns. Each cell is a count.

Notice what happened: "dog bites man" and "man bites dog" produce the same vector. Word order is discarded. That is a real limitation with real consequences. If meaning depends on sequence, bag-of-words cannot recover it.

TF-IDF is the next step. It down-weights terms that appear everywhere and up-weights terms that are distinctive. A word like "the" appears in almost every document, so its weight drops. A rare technical term gets more influence.

The resulting matrix is mostly zeros. Most documents use a small subset of the vocabulary. This is the sparse-data situation you have already seen: many possible features, few present. Extraction creates sparsity; it does not fight it.

Knowledge check

Check your understanding

Answer this question before you continue.

With a bag-of-words count vectorizer, what happens to “dog bites man” and “man bites dog” if they contain the same tokens with the same counts?
Output Prediction

Focus: Predict the consequence of bag-of-words extraction for documents with identical token counts but different word order.

Records to Numbers: DictVectorizer and Categorical Fields

Text is not the only raw input. Dictionary-like records are common: user profiles, transaction logs, configuration entries. Each record has keys, and the keys may differ.

DictVectorizer turns a list of Python dictionaries into a NumPy or SciPy matrix usable by scikit-learn estimators. Numeric values pass through as columns. String values become one-hot indicator columns.

from sklearn.feature_extraction import DictVectorizer

records = [
    {'age': 34, 'city': 'Berlin', 'plan': 'premium'},
    {'age': 28, 'city': 'Paris', 'plan': 'basic'},
    {'age': 45, 'city': 'Berlin', 'plan': 'basic'}
]

vec = DictVectorizer(sparse=False)
X = vec.fit_transform(records)

print(vec.get_feature_names_out())
print(X)

The output columns include age, city=Berlin, city=Paris, plan=basic, and plan=premium. Numeric values stay numeric. Categorical strings become binary indicators.

A missing key in one record becomes a zero in that column. That zero means "no value was supplied for this key in this record." Whether that absence means false, unknown, or genuinely zero depends on what the field represents. For a categorical indicator like plan=premium, zero usually means the record is not in that category. For a numeric field like age, a zero could be mistaken for a real age of zero, which would distort the feature. If missingness itself carries meaning, handle it explicitly rather than letting the vectorizer's default zero stand in for it.

The vocabulary is again learned at fit time. Unseen keys at transform time are ignored by default. This keeps the matrix shape stable, but it also means new categories are invisible unless you refit or handle them explicitly.

Knowledge check

Check your understanding

Answer this question before you continue.

A record contains `{'age': 34, 'city': 'Berlin'}`. Which description of its DictVectorizer representation is accurate?
Scenario Interpretation

Focus: Describe how DictVectorizer represents numeric and categorical dictionary fields in its output matrix.

Extraction Is Not Selection, and Not Transformation

Beginners often blur three steps that do different jobs. The cleanest way to separate them is by asking what each step operates on and what it produces.

StepWhat it operates onWhat it producesTypical column-count effect
ExtractionA raw input object (text, dict, image, audio)A model-ready numeric representationOften creates columns where none existed
TransformationAn existing feature representationA modified version of that representationMay keep, expand, or reshape columns
SelectionAn existing feature setA subset of those featuresRemoves columns only

Extraction constructs a representation from something that was not yet a feature matrix. Transformation modifies a representation you already have. Selection keeps a subset of what you already have.

The column-count column is a tendency, not a defining rule. Extraction usually creates columns because the raw input had none, but it can also combine with numeric fields that were already present, as the DictVectorizer example shows. Transformation often preserves column count, but it can expand it—think of adding polynomial terms or one-hot encoding a categorical column. Selection is the only step that strictly removes.

Why does the distinction matter? The leakage and validation rules differ for each step. Fitting a vectorizer on the full dataset leaks information about the test set. Fitting a scaler on the full dataset does the same. Selecting features using the target does too. Each step has its own failure mode, and mixing them up leads to fitting on the wrong data.

Knowledge check

Check your understanding

Answer this question before you continue.

A team converts raw documents into a numeric vocabulary-count matrix. Which step are they performing?
Comparison Reasoning

Focus: Distinguish constructing numeric features from raw input from selecting a subset of existing features.

Fit on Train, Transform Everything Else

The vocabulary, IDF weights, and category set are learned parameters. They must come from training data only.

If you fit a vectorizer on the full dataset, you leak information about the test set into training. The model sees words and categories that only appear in test data. Your evaluation score inflates. In production, the vocabulary is different, and the score collapses.

The correct pattern is simple: fit_transform on train, transform on test. Let a pipeline enforce the order.

from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.feature_extraction.text import TfidfVectorizer

pipeline = Pipeline([
    ('tfidf', TfidfVectorizer()),
    ('clf', LogisticRegression())
])

pipeline.fit(train_texts, train_labels)
score = pipeline.score(test_texts, test_labels)

The pipeline fits the vectorizer on training data only, then transforms the test data with the learned vocabulary. No leakage. No inflated score.

Common mistake: Fitting a vectorizer on all your text before splitting into train and test. The vocabulary then contains words from the test set, and your evaluation is no longer honest.

When Extraction Helps and When It Hurts

Use extraction when your input is genuinely non-numeric and you need a classical estimator to see it. Text, dictionaries, and categorical records are the common cases.

Be cautious when word order, spatial structure, or sequence matters. Bag-of-words throws that away. No amount of tuning brings it back. If sequence is the signal, you need a representation that preserves it.

Watch the dimensionality. A large vocabulary creates a very wide, very sparse matrix. That changes which models behave well. Linear models often handle it; distance-based methods may struggle.

Extraction is not a substitute for domain thinking. It is the mechanical step that follows a decision about what the input even is. You still need to decide what counts as a token, which fields matter, and whether the loss of structure is acceptable.

Deep learning handles raw text and images directly through learned representations. Classical extraction fills that gap by hand. That is not a weakness. It is a tradeoff: more control, more interpretability, less automatic feature discovery.

What to Do Next

Before you touch a model, ask what one row of your data actually is. Is it already a vector of numbers? If not, extraction is the step that makes it one.

Take a small text or dictionary dataset you already have. Run the appropriate vectorizer. Print the shape and a few rows of the resulting matrix. Watch the translation happen. The vocabulary must be learned on training data only, so split first, then fit.

That first matrix is the moment raw data becomes something a model can learn from. Everything after that is modeling. Everything before it was preparation. Extraction is the bridge.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

You are evaluating a text classifier with a held-out test set. Which workflow follows the article's guidance for the vectorizer?
Question 1 of 2Scenario Interpretation

Focus: Choose a leakage-safe workflow for learning a vectorizer's vocabulary and evaluating on held-out data.

A classification task depends on the order of words in a sentence. What limitation should make you cautious about using the article's bag-of-words example as the representation?
Question 2 of 2Misconception Check

Focus: Identify when a simple extraction method may be unsuitable because it discards structure important to the task.

References

  1. 8.2. Feature extraction — scikit-learn 1.9.1 documentationscikit-learn.org
  2. A Survey of Modern Questions and Challenges in Feature ...proceedings.mlr.press
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.