
Categorical vs Numerical Features: Choosing Representations That Models Can Use
The real question is not whether a column contains text or numbers. It is whether the model should treat those values as ordered quantities or as separate…
Read tutorialCreating, transforming, or selecting input representations to expose useful signal while preserving prediction-time validity.
Tagged articles
17 articles in this tag.

The real question is not whether a column contains text or numbers. It is whether the model should treat those values as ordered quantities or as separate…
Read tutorial
You've seen the demos. A neural network identifies objects in photos, translates speech in real time, and writes fluent text. Meanwhile, someone keeps…
Read tutorial
You run SelectKBest, get a tidy list of "important" features, then rerun with a different random seed and the list changes. That surprise is not a bug in…
Read tutorial
A feature idea is a hypothesis. Change one thing, measure the same way twice, and let the output decide.
Read tutorial
You've tried different algorithms. You've tuned parameters until your eyes glaze over. And still, your model's score sits at a plateau, stubbornly refusing…
Read tutorial
You have a list of sentences, a folder of images, or a pile of Python dictionaries. You call .fit() on your model, and it refuses. The error message is not…
Read tutorial
You picked 2^18 buckets. You trained the model. The metrics came back slightly worse than your explicit-vocabulary baseline, and now you are staring at the…
Read tutorial
Feature hashing trades a little information and a lot of interpretability for something genuinely rare in machine learning: a categorical encoding that…
Read tutorial
Two representations. One corpus. One split. One classifier. The only honest way to know what hashing costs you is to measure it against the vocabulary you…
Read tutorial
You fit a linear model. The coefficients look sensible. Each feature seems to contribute its own share, and the numbers align with your intuition. Then you…
Read tutorial
You're staring at a wide table of columns, and the model is underperforming. The fix could mean deleting half the columns, reshaping the ones that remain,…
Read tutorial
You fit a model to something like house prices or income. Most values bunch on the left, and a long tail stretches to the right. The model behaves…
Read tutorial
Fit a lasso on a dataset with a dozen features and something strange happens. A few coefficients come back as exactly 0.0. Not 0.0031. Not 0.0007. Zero.…
Read tutorial
You have a wide table of features, a model to build, and a quiet suspicion that half those columns are noise. The instinct is to shrink the dataset. But…
Read tutorial
You fit a straight line to data that clearly curves, and the line misses everything. Not by a little—systematically. It sits above the points in one…
Read tutorial
Target encoding is one of the most seductive traps in feature engineering. You encode a high-cardinality feature like city or merchant_id, watch your…
Read tutorial
Add a product term, watch the validation score tick up, declare victory. I have watched this exact sequence produce a confident conclusion on a dataset…
Read tutorial