Your First Machine Learning Model in Python with scikit-learn (Step by Step)
Reading about machine learning only gets you so far. The moment it clicks is when you train a model yourself and watch it make predictions on data it has never seen. This tutorial gets you there in about 30 lines of Python, using scikit-learn, the most popular library for classic machine learning.
We'll build a model that classifies breast tumor measurements as benign or malignant, using a real, well-known medical dataset that ships with scikit-learn. Along the way you'll learn the five-step workflow behind almost every machine learning project.
Educational example only. This is a classic teaching dataset. A model like this is not a diagnostic tool, and real medical AI requires far more data, validation and regulatory review.
The five-step workflow
Almost every machine learning project follows these five steps.
Keep this picture in mind. Whether you're predicting house prices, spotting fraud or reading medical images, the steps are the same: get data, hold some back, train, check honestly, then use the model.
Before you start
You need Python 3 and two libraries. Install them from a terminal:
pip install scikit-learn numpy
If you prefer not to install anything, the same code runs as-is in a free Google Colab notebook, where scikit-learn is already available.
Step 1: Load the data
The Wisconsin breast cancer dataset contains 569 tumor samples. For each one, doctors measured 30 features from a digitized image of a fine-needle biopsy, such as the radius, texture, perimeter and area of the cell nuclei. Each sample is labeled benign (357 samples) or malignant (212 samples). Here are three real rows, showing the first four of the 30 features:
| Mean radius | Mean texture | Mean perimeter | Mean area | … | Label |
|---|---|---|---|---|---|
| 17.99 | 10.38 | 122.80 | 1001.0 | … | malignant |
| 13.54 | 14.36 | 87.46 | 566.3 | … | benign |
| 13.08 | 15.71 | 85.63 | 520.0 | … | benign |
In machine learning terms, each row is a sample, the measurement columns are features (usually called X), and the answer we want to predict is the label or target (usually called y).
Step 2: Split into training and test sets
This is the step beginners most often skip, and it's the most important one. If you test a model on the same data it learned from, it can score perfectly just by memorizing. So we hide 20% of the data and use it only at the very end.
Holding back a test set is how you find out whether a model learned real patterns or just memorized.
Two details in the code matter. random_state=42 makes the split the same every time you run it, so your results are reproducible. stratify=y keeps the same benign-to-malignant ratio in both sets.
Step 3: Choose and train a model
We'll use logistic regression. Despite the name, it's a classification model, and it's a great first choice: fast, reliable and easy to interpret. It learns a weight for each feature, combines them into a single score, and converts that score into a probability.
Logistic regression weighs the measurements, then squeezes the result into a probability between 0 and 1.
Before training, we scale the features with StandardScaler. The raw measurements live on very different scales (area is in the hundreds, smoothness is below 1), and scaling puts them on an equal footing so no feature dominates just because of its units. Wrapping the scaler and the model in a pipeline ensures the same scaling is applied to training and test data automatically.
The complete code
Here is the whole program. Each numbered comment matches a step of the workflow.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, confusion_matrix
# 1. Load data: 569 tumor samples, 30 measurements each, labeled benign or malignant
data = load_breast_cancer()
X, y = data.data, data.target
print("Samples:", X.shape[0], " Features:", X.shape[1])
# 2. Split: keep 20% aside that the model never sees during training
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# 3. Build a model: scale the features, then fit a logistic regression
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
# 4. Train
model.fit(X_train, y_train)
# 5. Evaluate on the held-out test set
y_pred = model.predict(X_test)
print("Test accuracy:", round(accuracy_score(y_test, y_pred), 3))
print("Confusion matrix:\n", confusion_matrix(y_test, y_pred))
# 6. Predict for a new sample (here, the first test sample)
proba = model.predict_proba(X_test[:1])[0]
print("Prediction:", data.target_names[y_pred[0]],
" | P(malignant) =", round(proba[0], 3), " P(benign) =", round(proba[1], 3))
Running it prints:
Samples: 569 Features: 30
Test accuracy: 0.982
Confusion matrix:
[[41 1]
[ 1 71]]
Prediction: malignant | P(malignant) = 1.0 P(benign) = 0.0
Step 4: Evaluate honestly
An accuracy of 0.982 means the model got 112 of the 114 unseen test samples right. That's a strong result, but accuracy alone hides an important question: what kind of mistakes did it make? The confusion matrix answers that.
The confusion matrix from the test set: 112 of 114 correct. The two mistakes are very different kinds of error.
- 41 malignant tumors were correctly flagged, and 71 benign tumors were correctly cleared.
- 1 false alarm: a benign tumor was labeled malignant. In practice, this means an unnecessary follow-up test.
- 1 missed case: a malignant tumor was labeled benign. In medicine, this is the far more serious error.
This is why real-world projects look beyond accuracy, at measures such as recall (how many real positives were caught) and precision (how many flagged cases were real). Those metrics get their own post later in this series.
Step 5: Make predictions
Once trained, the model can score new samples. predict returns the class, while predict_proba returns the probability of each class, which is often more useful because you can choose your own threshold. A cautious screening tool, for example, might flag anything above a 20% chance of malignancy rather than the default 50%.
Try it yourself
- Swap
LogisticRegressionforRandomForestClassifier(import it fromsklearn.ensemble) and compare the scores. - Remove the
StandardScalerand see whether the model still trains as well. - Change
test_sizeto 0.5. Does the score change? Why might a larger test set give a more trustworthy estimate?
Key terms, decoded
- Feature
- An input measurement the model uses, such as tumor radius.
- Label (target)
- The answer the model learns to predict, such as benign or malignant.
- Training set
- The data the model learns from.
- Test set
- Data held back to measure performance on unseen examples.
- Classification
- Predicting a category rather than a number.
- Logistic regression
- A simple, interpretable classification model that outputs probabilities.
- Feature scaling
- Putting features on comparable scales before training.
- Confusion matrix
- A table of correct and incorrect predictions, broken down by class.
You've now completed the full machine learning loop: load, split, train, evaluate and predict. Every advanced technique, from deep learning to large language models, builds on this same foundation. Next up: why a model that looks great in training can fail badly in the real world, and how to catch it early.
Did the code run for you? Share your accuracy, or any error you hit, in the comments.

Comments
Post a Comment