Feature Extraction: Turn Raw Text and Records Into Model-Ready Features
You have a list of sentences, a folder of images, or a pile of Python dictionaries. You call .fit() on your model, and it refuses. The error message is not…

Key topics
You have a list of sentences, a folder of images, or a pile of Python dictionaries. You call .fit() on your model, and it refuses. The error message is not subtle: the estimator wants numbers, and you handed it words.
That rejection is not a formatting nuisance. It is the moment you discover that classical machine learning models do not see text, images, or records. They see rows of numbers. Feature extraction is the translation step that makes your raw input visible to them.
Why Estimators Refuse Your Raw Data
Most classical estimators in scikit-learn operate on a two-dimensional array with shape (n_samples, n_features). Each row is one example. Each column is one measurable property. Every cell holds a number.
Your raw data rarely arrives in that shape. A sentence has seven words; the next has forty. A dictionary has different keys per record. An image is a grid of pixels with spatial structure. None of these fit a fixed-width numeric grid.
The mismatch is structural. The model has no slot to put a word in. It cannot compute a distance between "dog" and "cat" because neither is a number. It cannot multiply a pixel by a weight if the pixel is still a file path.
You have already met the feature matrix when thinking about sparse data and feature space geometry. Now you are seeing it from the input side: before any model can learn, something must produce that matrix.
Pass a list of strings to a scikit-learn estimator and you get an error. Read the error. It is telling you exactly what is missing: a numerical feature vector for each sample.
Knowledge check
Check your understanding
Answer this question before you continue.
What Feature Extraction Actually Does
Feature extraction maps each raw object to a vector of numbers. The mapping must be the same for every sample. That consistency is what makes the output a matrix instead of a collection of unrelated arrays.
The output has rows and columns. Rows are samples. Columns are extracted features. The number of columns is fixed after extraction, even if the raw inputs had wildly different lengths.
Here is the part that surprises beginners: the feature set is learned from data. When you fit a vectorizer, it builds a vocabulary from the training corpus. That vocabulary becomes a learned parameter. At transform time, new documents are translated using the same vocabulary.
Think of it as a translator with a fixed dictionary built from the training corpus. New documents get translated with the same dictionary. If a word is not in the dictionary, it is either ignored or mapped to a special "unknown" token, depending on the vectorizer.
Where the analogy stops: unlike a translator, extraction can silently drop information. Word order, context, and spatial relationships may vanish. That loss is a design choice, not a bug. You are trading information for a representation the model can actually use.
Text to Numbers: The Bag-of-Words Idea
The most common extraction case is text. You have documents, and you need numerical features. The bag-of-words approach does three things: tokenize, count, and normalize.
Tokenization splits text into words or tokens. Counting records how often each token appears in each document. Normalization adjusts counts so that document length or common terms do not dominate.
CountVectorizer builds a vocabulary from the corpus and counts occurrences per document. The column order is the vocabulary order. Here is a small example:
from sklearn.feature_extraction.text import CountVectorizer
corpus = [
"the dog bites the man",
"the man bites the dog",
"the cat sleeps"
]
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)
print(vectorizer.get_feature_names_out())
print(X.toarray())
print(X.shape)
The vocabulary might be ['bites', 'cat', 'dog', 'man', 'sleeps', 'the']. The matrix has three rows and six columns. Each cell is a count.
Notice what happened: "dog bites man" and "man bites dog" produce the same vector. Word order is discarded. That is a real limitation with real consequences. If meaning depends on sequence, bag-of-words cannot recover it.
TF-IDF is the next step. It down-weights terms that appear everywhere and up-weights terms that are distinctive. A word like "the" appears in almost every document, so its weight drops. A rare technical term gets more influence.
The resulting matrix is mostly zeros. Most documents use a small subset of the vocabulary. This is the sparse-data situation you have already seen: many possible features, few present. Extraction creates sparsity; it does not fight it.
Knowledge check
Check your understanding
Answer this question before you continue.
Records to Numbers: DictVectorizer and Categorical Fields
Text is not the only raw input. Dictionary-like records are common: user profiles, transaction logs, configuration entries. Each record has keys, and the keys may differ.
DictVectorizer turns a list of Python dictionaries into a NumPy or SciPy matrix usable by scikit-learn estimators. Numeric values pass through as columns. String values become one-hot indicator columns.
from sklearn.feature_extraction import DictVectorizer
records = [
{'age': 34, 'city': 'Berlin', 'plan': 'premium'},
{'age': 28, 'city': 'Paris', 'plan': 'basic'},
{'age': 45, 'city': 'Berlin', 'plan': 'basic'}
]
vec = DictVectorizer(sparse=False)
X = vec.fit_transform(records)
print(vec.get_feature_names_out())
print(X)
The output columns include age, city=Berlin, city=Paris, plan=basic, and plan=premium. Numeric values stay numeric. Categorical strings become binary indicators.
A missing key in one record becomes a zero in that column. That zero means "no value was supplied for this key in this record." Whether that absence means false, unknown, or genuinely zero depends on what the field represents. For a categorical indicator like plan=premium, zero usually means the record is not in that category. For a numeric field like age, a zero could be mistaken for a real age of zero, which would distort the feature. If missingness itself carries meaning, handle it explicitly rather than letting the vectorizer's default zero stand in for it.
The vocabulary is again learned at fit time. Unseen keys at transform time are ignored by default. This keeps the matrix shape stable, but it also means new categories are invisible unless you refit or handle them explicitly.
Knowledge check
Check your understanding
Answer this question before you continue.
Extraction Is Not Selection, and Not Transformation
Beginners often blur three steps that do different jobs. The cleanest way to separate them is by asking what each step operates on and what it produces.
| Step | What it operates on | What it produces | Typical column-count effect |
|---|---|---|---|
| Extraction | A raw input object (text, dict, image, audio) | A model-ready numeric representation | Often creates columns where none existed |
| Transformation | An existing feature representation | A modified version of that representation | May keep, expand, or reshape columns |
| Selection | An existing feature set | A subset of those features | Removes columns only |
Extraction constructs a representation from something that was not yet a feature matrix. Transformation modifies a representation you already have. Selection keeps a subset of what you already have.
The column-count column is a tendency, not a defining rule. Extraction usually creates columns because the raw input had none, but it can also combine with numeric fields that were already present, as the DictVectorizer example shows. Transformation often preserves column count, but it can expand it—think of adding polynomial terms or one-hot encoding a categorical column. Selection is the only step that strictly removes.
Why does the distinction matter? The leakage and validation rules differ for each step. Fitting a vectorizer on the full dataset leaks information about the test set. Fitting a scaler on the full dataset does the same. Selecting features using the target does too. Each step has its own failure mode, and mixing them up leads to fitting on the wrong data.
Knowledge check
Check your understanding
Answer this question before you continue.
Fit on Train, Transform Everything Else
The vocabulary, IDF weights, and category set are learned parameters. They must come from training data only.
If you fit a vectorizer on the full dataset, you leak information about the test set into training. The model sees words and categories that only appear in test data. Your evaluation score inflates. In production, the vocabulary is different, and the score collapses.
The correct pattern is simple: fit_transform on train, transform on test. Let a pipeline enforce the order.
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.feature_extraction.text import TfidfVectorizer
pipeline = Pipeline([
('tfidf', TfidfVectorizer()),
('clf', LogisticRegression())
])
pipeline.fit(train_texts, train_labels)
score = pipeline.score(test_texts, test_labels)
The pipeline fits the vectorizer on training data only, then transforms the test data with the learned vocabulary. No leakage. No inflated score.
Common mistake: Fitting a vectorizer on all your text before splitting into train and test. The vocabulary then contains words from the test set, and your evaluation is no longer honest.
When Extraction Helps and When It Hurts
Use extraction when your input is genuinely non-numeric and you need a classical estimator to see it. Text, dictionaries, and categorical records are the common cases.
Be cautious when word order, spatial structure, or sequence matters. Bag-of-words throws that away. No amount of tuning brings it back. If sequence is the signal, you need a representation that preserves it.
Watch the dimensionality. A large vocabulary creates a very wide, very sparse matrix. That changes which models behave well. Linear models often handle it; distance-based methods may struggle.
Extraction is not a substitute for domain thinking. It is the mechanical step that follows a decision about what the input even is. You still need to decide what counts as a token, which fields matter, and whether the loss of structure is acceptable.
Deep learning handles raw text and images directly through learned representations. Classical extraction fills that gap by hand. That is not a weakness. It is a tradeoff: more control, more interpretability, less automatic feature discovery.
What to Do Next
Before you touch a model, ask what one row of your data actually is. Is it already a vector of numbers? If not, extraction is the step that makes it one.
Take a small text or dictionary dataset you already have. Run the appropriate vectorizer. Print the shape and a few rows of the resulting matrix. Watch the translation happen. The vocabulary must be learned on training data only, so split first, then fit.
That first matrix is the moment raw data becomes something a model can learn from. Everything after that is modeling. Everything before it was preparation. Extraction is the bridge.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


