Decision Trees, Random Forests and Gradient Boosting Explained (with Python)

Deep learning gets the headlines, but when the data lives in rows and columns, such as spreadsheets, databases and business records, a different family of models usually wins: tree-based models. Decision trees, random forests and gradient boosting power everything from credit scoring to demand forecasting. This post explains how each one works, with diagrams and real results in Python.

It builds on the machine learning series, especially overfitting and cross-validation, but you can follow it on its own.

Decision trees: twenty questions for data

A decision tree makes predictions by asking a series of yes-or-no questions about the data, like a game of twenty questions. Here is a real tree, learned from a dataset of 178 Italian wines, that predicts which of three grape cultivars a wine came from using its chemical measurements:

proline ≤ 755?OD280/OD315 ≤ 2.11?(light absorbance ratio)flavanoids ≤ 2.17?yesnoCultivar CCultivar BCultivar CCultivar Ayesnoyesno

A real tree learned from the wine data in the code below: two questions, and you have a prediction.

To classify a wine, start at the top and follow the answers down to a leaf. Is its proline level at most 755? If yes, check a light-absorbance measurement called the OD280/OD315 ratio; if no, check flavanoids. Two questions later, you have a prediction.

How a tree chooses its questions

During training, the algorithm looks at every feature and every possible cut-off, and picks the question that best separates the classes, meaning the groups on each side end up as "pure" as possible. It measures purity with a score such as Gini impurity. Then it repeats the process inside each branch, and keeps going until the groups are pure or a stopping rule, such as a maximum depth, kicks in.

Strengths and weaknesses

  • Easy to explain: you can print the rules and show them to anyone.
  • Little preparation needed: no feature scaling, and they handle a mix of numeric and category columns well.
  • Captures interactions: "if proline is high and flavanoids are low" comes naturally.
  • But they overfit easily. A tree allowed to grow without limits keeps splitting until it memorizes the training data, noise included, exactly the problem covered in the overfitting post. Small changes in the data can also produce a very different tree.

The fix for both weaknesses is to combine many trees. There are two main ways to do it.

Random forests: the wisdom of crowds

A random forest trains hundreds of trees, each on a different random sample of the data, and each considering only a random subset of features at every split. Individually, each tree is a bit different and a bit wrong. Together, they vote, and the majority wins.

Training data178 winesTree 1random samplevotes ATree 2random samplevotes ATree 300random samplevotes B…Majority voteCultivar Aprediction

A random forest trains hundreds of slightly different trees and lets them vote. Their individual mistakes tend to cancel out.

It's the same reason a crowd's average guess of the number of sweets in a jar is often better than most individual guesses: independent errors tend to cancel out. Random forests are robust, hard to overfit badly, and work well with default settings, which makes them an excellent first model for tabular data.

Gradient boosting: learning from mistakes

Gradient boosting also combines many trees, but in sequence rather than in parallel. It starts with one small tree that makes a rough prediction. The next tree is trained specifically on the errors the first one made. The third focuses on what's still wrong after two trees, and so on, often for hundreds of rounds.

Tree 1rough guessTree 2fixes Tree 1's errorsTree 3fixes what's left…hundreds moreremaining erroreach new tree is trained on the mistakes of all the trees before it

Gradient boosting builds trees in sequence. Each one focuses on what the previous trees still get wrong.

Each tree is small and weak on its own, but together they become very accurate. Gradient boosting libraries such as XGBoost, LightGBM and CatBoost, plus scikit-learn's own HistGradientBoostingClassifier, are among the most successful tools for tabular data and are widely used in industry and data science competitions. The trade-off is that boosting has more settings to tune, such as the learning rate and number of trees, and can overfit if pushed too far.

Put them to the test

The program below prints the small tree shown above, compares a single tree against a random forest and gradient boosting using 5-fold cross-validation, and lists the measurements the forest relied on most. The wine dataset ships with scikit-learn, so it runs with no downloads.

from sklearn.datasets import load_wine
from sklearn.model_selection import cross_val_score
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier

# 178 wines, 13 chemical measurements, 3 grape cultivars (class 0, 1, 2)
data = load_wine()
X, y = data.data, data.target

# 1. A small decision tree you can read like a flowchart
small_tree = DecisionTreeClassifier(max_depth=2, random_state=0).fit(X, y)
print(export_text(small_tree, feature_names=list(data.feature_names)))

# 2. Compare a single tree with two ensembles, using 5-fold cross-validation
models = {
    "Decision tree": DecisionTreeClassifier(random_state=0),
    "Random forest": RandomForestClassifier(n_estimators=300, random_state=0),
    "Gradient boosting": HistGradientBoostingClassifier(random_state=0),
}
for name, model in models.items():
    scores = cross_val_score(model, X, y, cv=5)
    print(f"{name:<18} accuracy {scores.mean():.3f}  (+/- {scores.std():.3f})")

# 3. Which measurements matter most? (random forest feature importance)
forest = models["Random forest"].fit(X, y)
ranked = sorted(zip(forest.feature_importances_, data.feature_names), reverse=True)
for importance, feature in ranked[:5]:
    print(f"{feature:<30} {importance:.3f}")

Running it prints:

|--- proline <= 755.00
|   |--- od280/od315_of_diluted_wines <= 2.11
|   |   |--- class: 2
|   |--- od280/od315_of_diluted_wines >  2.11
|   |   |--- class: 1
|--- proline >  755.00
|   |--- flavanoids <= 2.17
|   |   |--- class: 2
|   |--- flavanoids >  2.17
|   |   |--- class: 0

Decision tree      accuracy 0.888  (+/- 0.040)
Random forest      accuracy 0.967  (+/- 0.021)
Gradient boosting  accuracy 0.955  (+/- 0.022)
proline                        0.187
color_intensity                0.168
flavanoids                     0.154
od280/od315_of_diluted_wines   0.121
alcohol                        0.112
85%90%95%100%88.8%Decision tree96.7%Random forest95.5%Gradient boosting5-fold cross-validation · y-axis starts at 80%

Real results from the code below. Both ensembles clearly beat a single tree.

Both ensembles beat the single tree by a wide margin, and they're also more consistent from fold to fold (a smaller +/- spread). On this small, clean dataset the random forest edges out gradient boosting; on larger, messier datasets, well-tuned gradient boosting often comes out ahead. The lesson: try both, and let cross-validation decide.

Which features matter?

Tree ensembles can report feature importance: how much each feature contributed to the splits across all the trees.

proline0.187color intensity0.168flavanoids0.154OD280/OD315 ratio0.121alcohol0.112

Real feature importances from the random forest: which measurements the model relied on most.

Proline, color intensity and flavanoids top the list, which fits the small tree, where proline was the very first question. Treat these scores as a guide rather than proof: importance can be shared or split between correlated features, and it describes what the model uses, not what causes the outcome. For more careful explanations, practitioners often use permutation importance or SHAP values.

Which model should you choose?

ModelBest whenWatch out for
Decision treeYou need rules a person can read and auditOverfitting; limit the depth
Random forestYou want a strong, reliable baseline with little tuningLarger, slower models; less interpretable
Gradient boostingYou want top accuracy on tabular data and can tune itMore settings; can overfit if over-trained

For images, audio and text, neural networks usually win. For tables of rows and columns, tree ensembles remain the model to beat.

Key terms, decoded

Decision tree
A model that predicts by following a sequence of yes/no questions about the features.
Split
One question in the tree, such as "proline ≤ 755?".
Leaf
An end point of the tree that gives the prediction.
Gini impurity
A measure of how mixed the classes are in a group; splits aim to reduce it.
Ensemble
A model that combines many models to get a better result.
Random forest
An ensemble of trees trained on random samples that vote on the answer.
Bagging
Training models on random samples of the data, drawn with replacement, and combining them.
Gradient boosting
An ensemble that adds trees one at a time, each correcting earlier errors.
Feature importance
A score of how much the model relied on each feature.

One tree is easy to understand but fragile. Many trees, voting together in a random forest or correcting each other in gradient boosting, turn that simple idea into some of the most accurate models available for tabular data.

Which algorithm do you reach for first on tabular data? Share it in the comments.

Comments