Normalization vs standardization: which models need feature scaling?
The two scaling formulas, which models need scaling and which ignore it, which scaler to pick, and the leakage mistake that makes your cross-validation score a lie: one page.
Get the free PDF
One page, print-ready, free to share. No signup needed.
You scaled your data for a random forest. It did nothing. Trees split on order, not magnitude, and the StandardScaler you added out of habit changed nothing except the pipeline length. Meanwhile the KNN next to it, left unscaled, is letting salary drown age. One page on the two formulas, who needs them, and the leak that inflates your score. The print-ready A4 PDF is at the bottom.
The two
- Standardize: (x - mean) / std, a z-score.
- Normalize: (x - min) / (max - min), a 0 to 1 range.
- z-score versus 0-1: the names get swapped, so say which one you mean.
Who needs it
- Distance models: KNN, k-means, SVM.
- Gradient descent: linear, logistic, neural nets.
- Regularized models: L1/L2 penalties punish big scales.
Who does not
- Trees: splits ignore scale.
- Random forest: same.
- Boosting: scaling changes nothing.
Pick one
- Standardize: the default.
- Normalize: nets, images, bounded inputs.
- Robust scaler: outliers, uses the IQR.
Do it right
- Fit on train, transform train and test.
- Inside a Pipeline, so cross-validation stays leak-free.
- Save the scaler: production needs the same one.
The scaler lives inside the Pipeline
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
knn = make_pipeline(StandardScaler(),
KNeighborsClassifier())
knn.fit(X_train, y_train) # scaler fits on train only
knn.score(X_test, y_test) # test is transformed, not fit
# trees: no scaler, same result either way
rf = RandomForestClassifier().fit(X_train, y_train)
Fit on train, transform test, never the other way round. cross_val_score(knn, X, y) now refits the scaler inside every fold. A scaler fitted on all rows first would leak the test mean into training.
Gotchas
- Fit on all data: the test set leaks into the mean.
- One-hot columns: do not scale 0/1.
- Scaled target: unscale the predictions before reporting them.
The trap: three scaling mistakes
| You did | The effect | Instead |
|---|---|---|
| scaled for XGBoost | none, splits ignore scale | skip it, save the step |
| fit scaler on all rows | test mean leaked in | fit on train, in a Pipeline |
| KNN, no scaling | salary drowns age | StandardScaler first |
The quiz
Which model needs the scaler?
A) KNeighborsClassifier()
B) RandomForestClassifier()
A. KNN measures distances, so a feature in thousands drowns one in units. Trees split on order, so B gives the same result scaled or not.
Interview phrasing worth memorizing: I scale for distance-based and gradient-based models, skip it for trees, and always fit the scaler inside the pipeline so CV stays honest.
Frequently asked questions
What is the difference between normalization and standardization?
Which machine learning models need feature scaling?
Do random forests or XGBoost need feature scaling?
How do you scale features without data leakage?
Get the free PDF
One page, print-ready, free to share. No signup needed.