Accuracy Is Not Enough: Precision, Recall, F1 and ROC Curves Explained
"Our model is 95% accurate" sounds impressive. But accurate at what? A model can score high accuracy while completely failing at the one job it was built for. This post explains the metrics professionals actually use, precision, recall, F1 and ROC curves, with real numbers and a clear rule for when to use each.
We'll reuse the breast cancer model from your first scikit-learn model, which scored 98.2% accuracy on its test set. This time, we'll look underneath that number.
The accuracy trap
Our test set has 114 tumors: 72 benign and 42 malignant. Now imagine a "model" that ignores the data completely and always answers "benign".
The accuracy trap: a useless model can still look decent when one class is more common.
It's right 63% of the time, and it's utterly useless: it never catches a single cancer. This is the accuracy trap, and it gets worse as data becomes more imbalanced. For fraud detection, where maybe 1 transaction in 1,000 is fraud, a model that never flags anything scores 99.9% accuracy. Accuracy treats every mistake as equal, but in real life they rarely are.
The four outcomes
To do better, split predictions into four outcomes. We pick the class we care about catching, here malignant tumors, and call it the positive class. "Positive" doesn't mean good; it means "the thing we're looking for".
The four possible outcomes, with the real counts from our model. Every metric in this post is built from these four numbers.
A false negative is a missed case: a cancer the model cleared. A false positive is a false alarm: a healthy patient flagged for follow-up. Which of these is worse depends entirely on the problem, and that choice decides which metric you should care about.
Precision and recall
Here 9 of the 16 flagged dots are real positives (precision 56%), and 9 of the 10 real positives were flagged (recall 90%).
Precision answers: when the model raises an alarm, how often is it right?
Precision = true positives ÷ (true positives + false positives)
Recall (also called sensitivity) answers: of all the real positives, how many did the model catch?
Recall = true positives ÷ (true positives + false negatives)
For our model: precision is 41 ÷ (41 + 1) = 97.6%, and recall is 41 ÷ (41 + 1) = 97.6%. They happen to be equal here because there was exactly one false alarm and one miss.
A memory trick: recall is about recalling every real case, leaving nothing out. Precision is about being precise when you do speak up, with no false alarms.
The trade-off: moving the threshold
A classification model doesn't really output yes or no. It outputs a probability, and a threshold turns that into a decision. By default, anything above 50% is flagged. But you can choose a different threshold, and that choice trades precision against recall.
Real results: lowering the threshold catches every malignant tumor, at the cost of more false alarms. (The y-axis starts at 40%.)
Lowering the threshold to 5% means the model flags any tumor with even a small chance of being malignant. It now catches all 42 malignant tumors (100% recall), but it also flags 18 benign ones, so precision drops to 70%. For a cancer screening tool, that might be exactly the right trade: a false alarm means an extra test, while a miss can be fatal. For a spam filter, the opposite is true: a false alarm hides someone's important email, so you'd favor precision.
F1 score: one number for both
Sometimes you need a single score that balances precision and recall, for example to compare several models. The F1 score is their harmonic mean:
F1 = 2 × precision × recall ÷ (precision + recall)
Unlike a simple average, F1 is dragged down hard if either value is low. A model with 100% recall but 10% precision gets an F1 of only 18%, not 55%. That makes F1 a good default for imbalanced problems when both kinds of error matter.
ROC curves and AUC: judging the model at every threshold
Precision and recall describe the model at one threshold. But what if you haven't chosen a threshold yet, or want to compare models fairly? The ROC curve (receiver operating characteristic) shows performance at every threshold at once. It plots the true positive rate (recall) against the false positive rate (the share of negatives wrongly flagged).
The real ROC curve for our model. The closer the curve gets to the top-left corner, the better the model separates the two classes.
The AUC, the area under that curve, sums it up in one number between 0.5 and 1.0. It has a neat interpretation: the probability that the model ranks a randomly chosen positive higher than a randomly chosen negative. Our model's AUC of 0.995 means it almost always ranks malignant tumors as riskier than benign ones, even though the right threshold still has to be chosen separately.
One caution: when positives are very rare, ROC curves can look flattering. In those cases, many practitioners prefer the precision-recall curve, which focuses on how well the model finds the rare class.
The code
This program trains the same model as the earlier tutorial, then evaluates it at three thresholds, computes the ROC AUC, and exposes the accuracy trap.
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score)
X, y = load_breast_cancer(return_X_y=True) # 0 = malignant, 1 = benign
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_train, y_train)
# We want to catch malignant tumors, so treat class 0 as "positive"
is_malignant = (y_test == 0).astype(int)
p_malignant = model.predict_proba(X_test)[:, 0]
# 1. Compare metrics at different decision thresholds
for threshold in [0.5, 0.2, 0.05]:
flagged = (p_malignant >= threshold).astype(int)
print(f"threshold={threshold:<4} accuracy={accuracy_score(is_malignant, flagged):.3f} "
f"precision={precision_score(is_malignant, flagged):.3f} "
f"recall={recall_score(is_malignant, flagged):.3f} "
f"F1={f1_score(is_malignant, flagged):.3f}")
# 2. One number for the whole ROC curve
print("ROC AUC:", round(roc_auc_score(is_malignant, p_malignant), 3))
# 3. The accuracy trap: a "model" that always says benign
never_flag = [0] * len(is_malignant)
print("Always-benign accuracy:", round(accuracy_score(is_malignant, never_flag), 3),
" recall:", recall_score(is_malignant, never_flag))
Running it prints:
threshold=0.5 accuracy=0.982 precision=0.976 recall=0.976 F1=0.976
threshold=0.2 accuracy=0.939 precision=0.872 recall=0.976 F1=0.921
threshold=0.05 accuracy=0.842 precision=0.700 recall=1.000 F1=0.824
ROC AUC: 0.995
Always-benign accuracy: 0.632 recall: 0.0
Notice something important: lowering the threshold made accuracy worse but recall better. If you only tracked accuracy, you'd never choose the setting that catches every cancer. The right metric depends on the goal, not on which number is highest.
Which metric should you use?
| Situation | Focus on | Example |
|---|---|---|
| Classes are balanced and errors cost about the same | Accuracy | Classifying photos of cats and dogs |
| Missing a positive is very costly | Recall | Cancer screening, fraud detection, safety alerts |
| False alarms are very costly | Precision | Spam filtering, recommending a risky action |
| Both errors matter and classes are imbalanced | F1 score | Detecting defects in manufacturing |
| Comparing models before choosing a threshold | ROC AUC (or precision-recall AUC for rare positives) | Model selection |
Key terms, decoded
- Positive class
- The outcome you're trying to detect, such as a malignant tumor or a fraudulent payment.
- True / false positive
- A correct alarm / a false alarm.
- True / false negative
- A correct all-clear / a missed case.
- Precision
- The share of flagged items that are real positives.
- Recall (sensitivity)
- The share of real positives that were flagged.
- F1 score
- The harmonic mean of precision and recall.
- Threshold
- The probability above which the model flags an item as positive.
- ROC curve and AUC
- A plot of performance across all thresholds, and the area under it.
- Imbalanced data
- A dataset where one class is much rarer than the other.
Accuracy asks "how often is the model right?" The better question is "right about what, and at what cost?" Pick the positive class, decide which mistake hurts more, and choose your metric and threshold to match. Next up in the series: building a question-answering app over your own documents with retrieval-augmented generation (RAG).
Which metric matters most in your work? Share your use case in the comments.

Comments
Post a Comment