How AI Reads Chest X-rays: Multi-Label Image Classification Explained
The chest X-ray is one of the most common medical images in the world. It is quick, cheap and packed with information about the heart, lungs and bones. It's also one of the areas where AI has been studied most intensively. This post explains, in plain language, how an AI model "reads" a chest X-ray, why the task is multi-label, and why a high score in the lab is only the beginning.
Educational content only. Nothing in this post is medical advice. Chest X-ray AI tools are designed to assist qualified clinicians, not to replace them.
One image, many possible findings
Many image-recognition tasks have exactly one right answer: a photo shows a cat or a dog. That's multi-class classification, and the model picks a single winner.
A chest X-ray is different. The same image can show an enlarged heart (cardiomegaly) and fluid around the lungs (effusion) and signs of infection, all at once. Or it can show nothing abnormal. So the model must answer a separate yes-or-no question for every finding. That's multi-label classification.
A chest X-ray can show several findings at once, so it needs a multi-label model that answers yes or no for each one.
This small change has big consequences for how the model is built, trained and evaluated, as we'll see.
How the model sees the image
Chest X-ray models are typically convolutional neural networks (CNNs), or more recently vision transformers. A CNN scans the image with thousands of small learned filters. As covered in how neural networks learn, early layers learn to detect simple patterns such as edges and brightness changes. Deeper layers combine them into textures, shapes and eventually anatomy-level patterns, like the outline of the heart or a hazy region in a lung.
The image goes in; a separate probability comes out for each finding. Values shown are from the example code below.
Most models are not trained from scratch. They start from a network already trained on millions of everyday photos, then are fine-tuned on X-rays. This is called transfer learning: the general skill of seeing edges and textures carries over, so far fewer medical images are needed.
The key design choice: sigmoid, not softmax
At the very end, the network produces one raw score, called a logit, per finding. How those scores become probabilities is where multi-class and multi-label models differ:
- Softmax turns all scores into probabilities that add up to 1. Findings compete: if one goes up, the others must go down. That's right for "cat or dog", wrong for X-rays.
- Sigmoid turns each score into its own probability between 0 and 1, independently. Several findings can all be likely at once.
The same raw scores, two output functions. Softmax forces findings to compete; sigmoid lets several be present at once.
With softmax, the model can't say "both cardiomegaly and effusion are likely": they have to split the probability between them, so cardiomegaly barely scrapes past 0.5 and effusion drops to 0.38 and would be missed. With sigmoid, both come out above 0.85, and edema also crosses the threshold. Training uses a matching loss function, binary cross-entropy, which scores each finding as its own yes-or-no question.
Try it: multi-label in Python
The program below does two things. First, it applies softmax and sigmoid to the same five raw scores, reproducing the chart above. Second, it trains a small multi-label classifier, one yes/no model per label, on synthetic data that stands in for image features, and reports a separate AUC for each label. (The labels are deliberately generic: the synthetic prevalences are not real disease rates.)
import numpy as np
from sklearn.datasets import make_multilabel_classification
from sklearn.model_selection import train_test_split
from sklearn.multiclass import OneVsRestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
findings = ["Cardiomegaly", "Effusion", "Pneumonia", "Nodule", "Edema"]
# 1. Softmax vs sigmoid on the same raw model outputs (logits) for one image
logits = np.array([2.1, 1.8, -1.5, -2.0, 0.4])
softmax = np.exp(logits) / np.exp(logits).sum()
sigmoid = 1 / (1 + np.exp(-logits))
print("finding softmax sigmoid")
for name, s, g in zip(findings, softmax, sigmoid):
print(f"{name:<14} {s:6.2f} {g:6.2f}")
print("softmax sums to", round(softmax.sum(), 2), "| sigmoid sums to", round(sigmoid.sum(), 2))
# 2. Train a multi-label model: one yes/no classifier per finding
# (synthetic data: features stand in for image features from a CNN,
# and labels A-E are generic, not real disease rates)
X, Y = make_multilabel_classification(n_samples=3000, n_features=40, n_classes=5,
n_labels=1, random_state=7)
X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.25, random_state=0)
model = OneVsRestClassifier(LogisticRegression(max_iter=2000)).fit(X_train, Y_train)
probs = model.predict_proba(X_test)
print("\nlabel prevalence AUC")
for i in range(Y.shape[1]):
print(f"Finding {'ABCDE'[i]} {Y_test[:, i].mean():8.0%} {roc_auc_score(Y_test[:, i], probs[:, i]):.3f}")
print("images with 2+ findings:", f"{(Y_test.sum(axis=1) >= 2).mean():.0%}")
Running it prints:
finding softmax sigmoid
Cardiomegaly 0.51 0.89
Effusion 0.38 0.86
Pneumonia 0.01 0.18
Nodule 0.01 0.12
Edema 0.09 0.60
softmax sums to 1.0 | sigmoid sums to 2.65
label prevalence AUC
Finding A 3% 0.885
Finding B 26% 0.927
Finding C 16% 0.900
Finding D 26% 0.935
Finding E 31% 0.919
images with 2+ findings: 28%
Notice two things. The sigmoid probabilities add up to 2.65, which is fine: they're independent answers, not shares of a whole. And each label gets its own score and prevalence. Finding A appears in only 3% of images, a reminder that some findings are rare, which makes them much harder to learn and to evaluate reliably.
For real images, the same idea is built with a deep learning library. This short PyTorch sketch shows the essential changes to a standard image network. It's illustrative and wasn't run for this post, since it needs a labeled X-ray dataset and a GPU to train:
import torch.nn as nn
from torchvision import models
# Start from a network pretrained on everyday photos (transfer learning)
model = models.densenet121(weights="DEFAULT")
# Replace the final layer: 14 outputs, one per finding
model.classifier = nn.Linear(model.classifier.in_features, 14)
# One independent yes/no loss per finding (sigmoid + binary cross-entropy)
loss_fn = nn.BCEWithLogitsLoss()
Where the data comes from
Progress in this field has been driven by large public datasets, released for research:
- ChestX-ray14, released by the U.S. National Institutes of Health in 2017, with more than 100,000 frontal chest X-rays labeled for 14 findings.
- CheXpert, released by Stanford in 2019, with over 220,000 chest X-rays and labels that explicitly mark some findings as uncertain.
- MIMIC-CXR, released in 2019, which pairs a very large collection of chest X-rays with their free-text radiology reports.
A crucial detail: in several of these datasets, labels were extracted automatically from radiology report text using natural language processing, not drawn by radiologists image by image. That makes huge datasets possible, but it also means some labels are wrong or uncertain.
Measuring performance, finding by finding
Because every finding is its own yes/no question, performance is reported per finding, usually as an AUC, the ranking measure explained in the evaluation metrics post. A single average can hide a lot: a model might be excellent at spotting effusion and poor at small nodules. In practice, each finding also needs its own decision threshold, chosen according to whether a miss or a false alarm is more costly for that condition.
Why lab scores aren't the whole story
Four reasons a chest X-ray model that scores well in the lab can disappoint in a real hospital.
- Noisy labels: labels mined from reports can be wrong, and a model can't be more reliable than the labels it's judged against.
- Rare findings: with few positive examples, both learning and evaluation become unstable.
- Shortcut learning: models have been shown to pick up on clues unrelated to the disease itself, such as chest tubes that indicate a condition was already treated, text markers burned into the image, or differences between scanners. The score looks good for the wrong reasons.
- Hospital shift: patients, equipment and imaging practices differ between hospitals, so a model can perform noticeably worse outside the place it was trained.
This is why careful external validation, testing on data from different hospitals, is essential before any clinical use, and why regulators review these tools as medical devices in many countries.
How it's used in practice
In real clinical settings, chest X-ray AI is typically an assistant. It can flag possible findings, highlight suspicious regions with a heatmap, and move urgent-looking cases to the top of a radiologist's worklist. The diagnosis remains the clinician's responsibility.
How chest X-ray AI is typically used: as a second pair of eyes and a triage tool, with a clinician making the final call.
A newer direction is vision-language models, trained on images paired with report text, which can describe findings in words or even draft a report. They are promising but raise the same questions about accuracy, hallucination and validation that apply to large language models generally.
Key terms, decoded
- Multi-label classification
- Predicting any number of labels for one input, each as its own yes/no question.
- Multi-class classification
- Predicting exactly one label out of several.
- Convolutional neural network (CNN)
- A neural network that learns image features using small sliding filters.
- Transfer learning
- Starting from a model trained on another task, then fine-tuning it on new data.
- Logit
- A raw model score before it's turned into a probability.
- Sigmoid
- A function that turns one score into an independent probability between 0 and 1.
- Binary cross-entropy
- The loss function used to train each yes/no output.
- Shortcut learning
- When a model relies on irrelevant clues that happen to correlate with the label.
- External validation
- Testing a model on data from a different source, such as another hospital.
Reading a chest X-ray with AI comes down to a clear recipe: a convolutional network pretrained on everyday images, fine-tuned with one sigmoid output per finding, and evaluated finding by finding. The hard part isn't the architecture; it's the data, the labels, and proving the model works safely in the real world.
Which area of healthcare AI would you like covered next? Let us know in the comments.

Comments
Post a Comment