Overfitting Explained: Train, Validation and Test Sets (and Why Models Fool You)

A model that scores 100% during training can still fail the moment it meets real data. It's one of the most common and costly mistakes in machine learning, and it has a name: overfitting. This post shows what it looks like, why it happens, and the simple habits that stop it from fooling you.

In the first scikit-learn tutorial, we split data into a training set and a test set. Here we'll see exactly why that split matters, add a third set, and learn cross-validation, the standard technique professionals use to choose models honestly.

Memorizing is not learning

Imagine two students preparing for an exam using last year's paper. One understands the ideas behind each answer. The other memorizes the answers word for word. On last year's paper, both score 100%. On this year's exam, with new questions, only the first student does well.

Machine learning models can be either student. The goal is always the first kind: learn the underlying pattern so you can handle new cases. Here's what the three possible outcomes look like on the same data:

Underfittingtoo simpleGood fitcaptures the patternOverfittingmemorizes the noise

The same data, three models. Only the middle one will predict new points well.

  • Underfitting: the model is too simple to capture the pattern. It does badly on training data and on new data.
  • Good fit: the model captures the real pattern and ignores the random noise.
  • Overfitting: the model is so flexible that it bends to fit every point, including the noise. It looks perfect on training data but makes wild predictions between and beyond those points.

Why overfitting happens

Real data always contains noise: measurement errors, unusual cases and pure chance. A very flexible model has enough freedom to learn that noise as if it were a real pattern. Overfitting becomes more likely when:

  • the model is very complex, such as a deep decision tree or a huge neural network;
  • the dataset is small, so noise makes up a larger share of what the model sees;
  • the model has many features compared with the number of examples;
  • training runs for too long without checking progress on unseen data.

Watch it happen with real numbers

Let's make overfitting visible. We'll train decision trees on the breast cancer dataset from the previous tutorial, letting each tree grow deeper and therefore more complex. For each depth we measure two things: accuracy on the data it trained on, and accuracy on data it didn't see, using cross-validation (explained below).

90%95%100%123510no limittree depth (model complexity) →training accuracycross-val accuracybest: depth 5

Real results from the code below. Training accuracy keeps climbing to 100%, but accuracy on unseen data peaks at depth 5.

The pattern is the classic signature of overfitting. As the tree gets deeper, training accuracy climbs all the way to a perfect 100%. But performance on unseen data peaks at depth 5 and then drops. The deepest trees have memorized the training data, noise included. If you judged these models on training accuracy alone, you'd pick exactly the wrong one.

The fix, part 1: a three-way split

In the first tutorial, we used a test set to score the final model. But in real projects you make many decisions along the way: which model to use, how deep a tree should be, which features to include. If you make those decisions by checking the test set over and over, you slowly tune your model to that particular test set, and its score stops being honest.

The solution is to add a validation set for making decisions, and keep the test set locked away until the end:

Training60%: fit the modelValidation20%: tuneTest20%: final checkUse the test set only at the very end, and only once.

A three-way split: train on one part, make decisions on another, and keep the last part for a single honest final score.

The golden rule: the test set is for a final exam, not for studying. Look at it once, when every decision has been made. If you then go back and change things, that test set is no longer an honest measure.

The fix, part 2: cross-validation

A single validation set has a weakness: the score depends on which samples happened to land in it, especially with small datasets. K-fold cross-validation fixes this. The training data is divided into k equal parts, called folds. The model is trained k times, each time holding out a different fold for validation, and the k scores are averaged.

Round 1validatetraintraintraintrainRound 2trainvalidatetraintraintrainRound 3traintrainvalidatetraintrainRound 4traintraintrainvalidatetrainRound 5traintraintraintrainvalidatefinal score = average of the 5 validation scores

5-fold cross-validation: every sample gets used for validation exactly once, giving a more reliable score.

With 5 folds, every training sample is used for validation exactly once, and the final score is far more stable than one split. It costs more computation, since the model is trained five times, but for most classic machine learning problems that's well worth it.

The code

Here is the full experiment. It holds back a test set, uses 5-fold cross-validation on the training data to compare tree depths, picks the best depth automatically, and only then looks at the test set, exactly once.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.tree import DecisionTreeClassifier

X, y = load_breast_cancer(return_X_y=True)

# Hold back a test set first, and don't touch it until the very end
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

# 1. Watch overfitting happen: deeper trees memorize more
scores = {}
for depth in [1, 2, 3, 5, 10, None]:
    tree = DecisionTreeClassifier(max_depth=depth, random_state=0)
    tree.fit(X_train, y_train)
    train_acc = tree.score(X_train, y_train)
    cv_acc = cross_val_score(tree, X_train, y_train, cv=5).mean()
    scores[depth] = cv_acc
    print(f"max_depth={str(depth):>4}  train={train_acc:.3f}  cross-val={cv_acc:.3f}")

# 2. Choose the depth with the best cross-validation score
best_depth = max(scores, key=scores.get)
print("Best depth by cross-validation:", best_depth)

# 3. Only now, evaluate once on the untouched test set
final = DecisionTreeClassifier(max_depth=best_depth, random_state=0)
final.fit(X_train, y_train)
print("Final test accuracy:", round(final.score(X_test, y_test), 3))

Running it prints:

max_depth=   1  train=0.923  cross-val=0.897
max_depth=   2  train=0.958  cross-val=0.919
max_depth=   3  train=0.976  cross-val=0.925
max_depth=   5  train=0.993  cross-val=0.938
max_depth=  10  train=1.000  cross-val=0.925
max_depth=None  train=1.000  cross-val=0.925
Best depth by cross-validation: 5
Final test accuracy: 0.93

Notice how the final test accuracy (0.930) is close to the cross-validation estimate (0.938). That's the payoff: cross-validation predicted real-world performance well, while the training score (0.993) was far too optimistic. And for comparison, the simple logistic regression from the previous tutorial scored 0.982 on the same test set. A more complex model isn't automatically a better one.

More ways to prevent overfitting

  • Get more data. The more real examples a model sees, the harder it is to memorize noise.
  • Simplify the model. Limit tree depth, use fewer features, or pick a simpler algorithm.
  • Regularization. Add a penalty for complexity during training, so the model prefers simpler explanations. Most scikit-learn models have a setting for this.
  • Early stopping. For neural networks, stop training when validation performance stops improving.
  • Dropout and data augmentation. In deep learning, randomly switching off neurons or creating varied copies of training images makes memorization harder.

How to spot overfitting quickly

What you seeWhat it usually means
Training score high, validation score much lowerOverfitting: simplify, regularize or get more data
Both scores lowUnderfitting: try a more capable model or better features
Both scores high and close togetherA good fit: confirm once on the test set
Test score much worse than validation scoreThe validation set was overused, or the test data differs from training data

Key terms, decoded

Overfitting
When a model learns noise in the training data and performs poorly on new data.
Underfitting
When a model is too simple to capture the real pattern.
Generalization
A model's ability to perform well on data it has never seen.
Validation set
Data used to compare models and tune settings during development.
Test set
Data kept aside for one final, honest measurement.
Cross-validation
Training and validating several times on different folds of the data, then averaging the scores.
Hyperparameter
A setting you choose rather than the model learning it, such as tree depth.
Regularization
A penalty on model complexity that discourages overfitting.

A high training score proves only that a model can remember. What matters is performance on data it has never seen, and the habits in this post, a locked-away test set, a validation step and cross-validation, are how you measure that honestly. Next in the series: the right metrics for evaluating a model, including precision, recall and ROC curves.

Have you been caught out by an overfit model? Share the story in the comments.

Comments