A Second Opinion, Not a Replacement
A diagnosis-assist model doesn't output a verdict — it outputs a probability, and where you set the threshold for a flag changes what counts as suspicious. Move the slider below and watch that tradeoff happen live.
“AI diagnosed the tumour” is a headline, not what actually happens in a clinic. A diagnosis-assist model does not hand a doctor a verdict — it hands them a probability, and where the system draws the line for what counts as a flag changes what gets caught and what gets missed, every single time.
A second opinion, not a replacement
In actual clinical use, a diagnosis-assist model sits alongside a radiologist or physician, not in place of one. It reviews the same scan or the same chart, surfaces cases it thinks deserve a closer look, and the clinician makes the actual call — the model is there to catch what a tired reviewer on their fortieth case of the day might miss, not to remove the reviewer.
That arrangement is not a temporary stepping stone toward full automation — it is the actual deployed shape almost everywhere one of these tools is in clinical use today. A model that is wrong in a way a busy clinician would catch is an acceptable cost of doing business; a model that is wrong with nobody checking is a different kind of product entirely, and regulators have been explicit that the second one is not what gets approved for anything but the narrowest, lowest-stakes uses.
A model outputs a probability, not a verdict
Underneath the alert a doctor sees, the model itself never outputs “yes” or “no.” It outputs a number between 0 and 1 — its estimate of how likely this case is to be positive, based on everything similar it has seen before. Something has to turn that number into an actual flag, and that something is a threshold, chosen in advance by the people who built and deployed the system.
Two systems can share the exact same underlying number and behave completely differently once a hospital deploys them, because the number itself commits to nothing. A model that says 0.61 has not said “probably positive” or “probably negative” — it has said “more likely than not, by not very much,” and it is entirely up to the people deploying it to decide how much weight that deserves.
Moving the threshold changes what counts as a flag
Move the slider below across forty simulated cases and watch what happens. Lower the threshold and more true cases get caught — but more healthy cases get flagged too, each one a false alarm that costs a clinician's time and a patient's anxiety on a follow-up that finds nothing. Raise it and false alarms drop — at the direct cost of missing real cases that scored just under the line.
Flagged
17/40
Correctly flagged
12
False alarms
5
Missed cases
3
Lower the threshold and false alarms climb before missed cases fall to zero — of these 40 simulated cases, 80% of genuine cases are caught at the current threshold. There is no single threshold that maximises both catching real cases and avoiding false alarms at once.
Sensitivity and specificity are not the number a patient wants
Everything above describes properties of the test, not properties of your result. Sensitivity is the share of people who actually have the condition that the model correctly flags. Specificity is the share of people who don't have it that the model correctly clears. Both are honest, useful numbers — and neither one answers the question a patient is actually holding when a result comes back positive: given that I was flagged, how likely am I to actually have this?
That number has a name — positive predictive value, or PPV — and it depends on something sensitivity and specificity don't capture at all: how common the condition actually is in the population being tested. A test can be excellent by every measure a manufacturer would put on a spec sheet and still produce a flag that is wrong far more often than it is right, once the condition it is looking for is rare enough.
The arithmetic that collapses at low prevalence
Work it out with real numbers rather than taking that on faith. Screen 10,000 people for a condition that actually affects 1% of them, using a test that is 90% sensitive and 90% specific — numbers a manufacturer would happily print on a box.
100 people actually have the condition. 9,900 don't. Sensitivity 90% Specificity 90% 90 true positives 990 false positives 10 false negatives 8,910 true negatives Flagged positive: 90 + 990 = 1,080Of those, 90 actually have the condition. Positive predictive value: 90 / 1,080 ≈ 8%A test that is right nine times out of ten in both directions still produces a positive flag that is wrong about eleven times out of twelve, because at 1% prevalence the healthy population is so much larger than the sick one that even a small false-positive rate applied to it swamps the true positives. That is not the test failing. That is arithmetic — sensitivity and specificity stay the same 90% and 90% no matter how rare the condition is; prevalence is the variable actually doing the damage.
Testing everyone, prevalence 1% — positive predictive value lands around 8%. Most flags are false alarms.
Testing only people already showing signs, prevalence 30% — the same test's positive predictive value climbs to around 79%. Same model, same threshold, a completely different number.
This is why a screening programme aimed at an entire healthy population and a diagnostic test ordered for someone already showing symptoms are different products wearing the same underlying model, and why moving the threshold in the section above changes sensitivity and specificity without ever touching the deeper problem. A model tuned to catch more true cases at a low prevalence still returns mostly false alarms — it just returns a different mix of them.
Key takeaways
- A diagnosis-assist model is deployed as a second opinion a clinician reviews, not a replacement for their judgement — and regulators have kept it that way.
- The model itself outputs a probability between 0 and 1, never a direct yes-or-no verdict; a threshold set by the people deploying it turns that number into a flag.
- Sensitivity and specificity describe the test, not your result — the number a patient actually wants is positive predictive value, and it depends on how common the condition is.
- At 1% prevalence, a test that's 90% sensitive and 90% specific still returns a positive flag that's wrong about eleven times out of twelve.
- No single threshold eliminates both missed cases and false alarms — choosing one is a values decision about which failure is more acceptable, not just a technical setting.
Quick check
Answer these to unlock the next chapter — 3 of 4 to pass. You can retake it anytime.
Answer every question to check.
Make a free account to read on
Every chapter is free — an account is how your progress, XP, and streak follow you from your laptop to your phone, and how you show up on the leaderboard. No payment, no trial.