Card 05 of 13 · clinical decision
copilot

Diagnosis and differential

"Can you use AI for diagnosis?" is the wrong question. The right one is: pointed which way? Aimed at confirming, the tool amplifies your bias with fluency. Aimed at challenging, it is one of the best defenses against premature closure medicine has ever had.

The case

You have a working diagnosis and about ninety seconds. You open the chat, describe the case the way you already see it, and ask what it thinks. Whatever comes back will feel like a second opinion. Whether it actually is one depends entirely on how you asked.

Guiding question: is the tool pointed at confirming me, or at challenging me?

Broadening the differentialCopilot
Naming the discriminatorsCopilot
Testing your hypothesisCopilot
Establishing the diagnosisNot delegable
01 · What it does, and what it does not
What it does well

Widens a differential that a tired memory has narrowed, remembers the rare thing you have seen twice, audits your history-taking for the question you did not ask, and articulates what argues for and against each possibility.

What it does not do

Know what it does not know about your case. It will not ask what is missing unless you tell it to, and it will not notice that you only told it the half that supports you.

Clinical design position · the intern rule

Treat the tool as the most dedicated intern in the world. It has read everything and never gets tired. But it is still an intern, with two differences that are the whole lesson. An intern learns from your correction; the model makes the same mistake tomorrow. An intern calls you at 3 a.m. to say "I don't know"; the model is exactly as fluent when it is wrong as when it is right. So supervision here is structural. It is not trust that accumulates over time.

02 · Where it fails here
Deference to your anchor

The hypothesis you state arrives with an unfair advantage: models tuned on human feedback weight the user's assertion, and agreeing is the path of least resistance. Even the order in which you tell the story organizes the differential around what you said first.

Premature closure at scale

A machine built on "most likely" converges on the common, confidently. The atypical presentation and the dangerous rare thing are exactly what falls off the list, and exactly what the differential existed to catch.

Decorative differentials

Ten possibilities in a list look like rigor. Without the discriminator for each one, the question, finding, or test that separates it from the others, it is rigor as decoration: a list that helps you decide nothing.

These three link to the site's failure taxonomy, where the transcripts live.

03 · The ego in the room
The failure mode that is not the model's

There is a use this card would be dishonest to skip: the "just to confirm" use. The clinician has already decided, presents only the points that support the decision, and pushes back until the model agrees.

Models do yield under pressure; that is documented and exploitable. But manufactured agreement confirms nothing. If you pushed until it agreed, the only thing you confirmed is that it yields.

The invitation

Could you be right anyway? Sure. You could also be watching your own bias come back to you in the voice of a second opinion.

The exercise this card invites is the cheapest one the tool has made possible: could I be wrong? Reading "maybe" from a machine asks something of the reader. That is exactly why it is worth doing.

04 · The safer workflow, in two passes
00Your hypothesis, written down.One line is enough. It is the precommitment, a barrier against anchoring: without it, the model's answer becomes your opinion and you never notice the swap.
01Pass one, blind.Give the case without revealing your hypothesis and ask for its differential. Since you already committed on paper, comparing the two lists is the most honest test available.
02Pass two, directed audit.Now reveal your leading diagnosis and ask it to be examined: what argues for, what argues against, what you did not consider, and which questions and exam maneuvers are missing.
03Disagreement is a trigger, not a verdict.If the model insists on something you ruled out, it obliges you to articulate why. If you cannot articulate it, that is information: go look.
04Record the differential you excluded, and fly manual on purpose.What was considered and ruled out, with the reason, in the chart: patient safety and one of the best forms of documentation protection there is. And run some cases without the tool, deliberately. Skill atrophy under continuous AI use is documented, not theoretical.
Console · exercise

Same case, two directions

Synthetic case. 34-year-old woman, chest pain for two days. "I worked out the day before yesterday and overdid it, so it's probably muscular." The pain is worse with arm movement and tender to palpation. Also in the case, not emphasized: combined oral contraceptive, an eleven-hour flight last week, HR 96, and shortness of breath on stairs she attributes to being out of shape.

Direction one · "probably muscular, right?"

This is consistent with musculoskeletal chest pain after exertion, likely costochondritis or a strain. Worsening with movement and tenderness to palpation support that. Plan: anti-inflammatory, relative rest, local heat; reassess if it persists.

Polite, plausible, and notice what did not happen: not one question about what the case did not say.

Direction two · "what can I not afford to miss?"

Before the differential. Data are missing that would change the probabilities. Estrogen use? Recent immobilization or long travel? Any shortness of breath, even mild? Heart rate?

With a combined contraceptive, an eleven-hour flight a week ago, HR 96, and new exertional dyspnea, chest pain reproducible on palpation does not rule out pulmonary embolism: reproducibility lowers the probability, it does not zero it.

Position. The musculoskeletal hypothesis is plausible, but what I cannot afford to miss here is PE. Discriminator: pretest probability by a validated score before any imaging.

What actually changed

Same patient, same model, same two days of chest pain. The user supplied only what supported the anchor, and the model handed the anchor back validated. The ego in the room, in action.

What changed in the second pass was the direction of the question, not the quality of the tool.

Where it goes next

The second answer does not diagnose anything. It names the thing that must be excluded and points at the instrument that decides the next step.

That score, and how easily it goes wrong, is Card 06.

05 · Before it changes your management

0 of 5

A guideline preference that never changes lives in your persistent instructions (Card 01); a choice specific to this case goes in the prompt.

The prompt · pass one, blind
Case: [objective description, no hypothesis from me,
no identifiers].
Build the differential, from most likely to what I cannot
afford to miss. For each possibility, give the discriminator:
the question, finding, or test that separates it from the
others.
Before the list, tell me what is missing from the case that
would change your probabilities.
Pass two · directed audit
My leading diagnosis is [hypothesis]. Reference
guideline, if applicable: [guideline].
It is fine to disagree with me. Evaluate this technically; do
not validate out of politeness.
Answer in exactly this structure:
1) Position: agree or disagree, and why.
2) What in the case argues for it.
3) What argues against it.
4) Possibilities I did not consider, each with its
discriminator.
5) Missing history questions and exam maneuvers.
6) What I cannot afford to miss in this presentation.

The fixed structure is not fussiness. Numbered sections tell you what you are reading and where to look, and a wall of text with no addresses gets skimmed. Skimming is trusting out of fatigue.

The time a case deserves

Honesty this site owes you: not every visit has room for the full two-pass ritual, and productivity pressure is real, not a choice you are making. The ritual is for the case that earns it. When it will not fit in the visit, the invitation still stands for later, and reviewing the interesting cases of the week during study time is the perfect place for "could I have been wrong?" It has never cost so little to ask. And the principle underneath belongs to this whole page: patients should have a right to an unhurried visit, and clinicians should have a right to enough time to build the reasoning that produces good care. The tools on this map exist to give minutes back to that time. Not to compress it further.

06 · Evidence

Exam performance and clinical performance are different measurements, models lose ground when they have to run the case rather than receive it, and continuous use has a measurable cost to the clinician's own skill. All three support the same conclusion: structural supervision, not accumulated trust.

  1. [article] BRIDGE: benchmarking large language models for understanding real-world clinical practice texts. Nature Biomedical Engineering, 2026. 44.8% on real clinical text tasks, against 90+ on exams. doi.org/10.1038/s41551-026-01719-2
  2. [systematic review] Knowledge-practice performance gap in clinical large language models: systematic review of 39 benchmarks. J Med Internet Res, 2025;27:e84120. 84 to 90% on knowledge exams versus 45 to 69% on practical tasks. PROSPERO CRD420251139729. DOI to confirm
  3. [article] Hager P, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine, 2024;30:2613-2622. doi.org/10.1038/s41591-024-03097-1
  4. [observational] Budzyń K, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. Lancet Gastroenterology & Hepatology, 2025;10:896-903. Adenoma detection without AI fell from 28.4% to 22.4% after continuous exposure. doi.org/10.1016/S2468-1253(25)00133-5

Further reading · Sharma M, et al. Towards understanding sycophancy in language models. ICLR 2024. Why agreement is the path of least resistance: arxiv.org/abs/2310.13548

Educational content; synthetic case. Not a substitute for clinical judgment or for the rules that apply where you practice.