"Can you use AI for diagnosis?" is the wrong question. The right one is: pointed which way? Aimed at confirming, the tool amplifies your bias with fluency. Aimed at challenging, it is one of the best defenses against premature closure medicine has ever had.
You have a working diagnosis and about ninety seconds. You open the chat, describe the case the way you already see it, and ask what it thinks. Whatever comes back will feel like a second opinion. Whether it actually is one depends entirely on how you asked.
Guiding question: is the tool pointed at confirming me, or at challenging me?
Widens a differential that a tired memory has narrowed, remembers the rare thing you have seen twice, audits your history-taking for the question you did not ask, and articulates what argues for and against each possibility.
Know what it does not know about your case. It will not ask what is missing unless you tell it to, and it will not notice that you only told it the half that supports you.
Treat the tool as the most dedicated intern in the world. It has read everything and never gets tired. But it is still an intern, with two differences that are the whole lesson. An intern learns from your correction; the model makes the same mistake tomorrow. An intern calls you at 3 a.m. to say "I don't know"; the model is exactly as fluent when it is wrong as when it is right. So supervision here is structural. It is not trust that accumulates over time.
The hypothesis you state arrives with an unfair advantage: models tuned on human feedback weight the user's assertion, and agreeing is the path of least resistance. Even the order in which you tell the story organizes the differential around what you said first.
A machine built on "most likely" converges on the common, confidently. The atypical presentation and the dangerous rare thing are exactly what falls off the list, and exactly what the differential existed to catch.
Ten possibilities in a list look like rigor. Without the discriminator for each one, the question, finding, or test that separates it from the others, it is rigor as decoration: a list that helps you decide nothing.
These three link to the site's failure taxonomy, where the transcripts live.
There is a use this card would be dishonest to skip: the "just to confirm" use. The clinician has already decided, presents only the points that support the decision, and pushes back until the model agrees.
Models do yield under pressure; that is documented and exploitable. But manufactured agreement confirms nothing. If you pushed until it agreed, the only thing you confirmed is that it yields.
Could you be right anyway? Sure. You could also be watching your own bias come back to you in the voice of a second opinion.
The exercise this card invites is the cheapest one the tool has made possible: could I be wrong? Reading "maybe" from a machine asks something of the reader. That is exactly why it is worth doing.
Synthetic case. 34-year-old woman, chest pain for two days. "I worked out the day before yesterday and overdid it, so it's probably muscular." The pain is worse with arm movement and tender to palpation. Also in the case, not emphasized: combined oral contraceptive, an eleven-hour flight last week, HR 96, and shortness of breath on stairs she attributes to being out of shape.
This is consistent with musculoskeletal chest pain after exertion, likely costochondritis or a strain. Worsening with movement and tenderness to palpation support that. Plan: anti-inflammatory, relative rest, local heat; reassess if it persists.
Polite, plausible, and notice what did not happen: not one question about what the case did not say.
Before the differential. Data are missing that would change the probabilities. Estrogen use? Recent immobilization or long travel? Any shortness of breath, even mild? Heart rate?
With a combined contraceptive, an eleven-hour flight a week ago, HR 96, and new exertional dyspnea, chest pain reproducible on palpation does not rule out pulmonary embolism: reproducibility lowers the probability, it does not zero it.
Position. The musculoskeletal hypothesis is plausible, but what I cannot afford to miss here is PE. Discriminator: pretest probability by a validated score before any imaging.
Same patient, same model, same two days of chest pain. The user supplied only what supported the anchor, and the model handed the anchor back validated. The ego in the room, in action.
What changed in the second pass was the direction of the question, not the quality of the tool.
The second answer does not diagnose anything. It names the thing that must be excluded and points at the instrument that decides the next step.
That score, and how easily it goes wrong, is Card 06.
0 of 5
A guideline preference that never changes lives in your persistent instructions (Card 01); a choice specific to this case goes in the prompt.
Case: [objective description, no hypothesis from me,
no identifiers].
Build the differential, from most likely to what I cannot
afford to miss. For each possibility, give the discriminator:
the question, finding, or test that separates it from the
others.
Before the list, tell me what is missing from the case that
would change your probabilities.
My leading diagnosis is [hypothesis]. Reference guideline, if applicable: [guideline]. It is fine to disagree with me. Evaluate this technically; do not validate out of politeness. Answer in exactly this structure: 1) Position: agree or disagree, and why. 2) What in the case argues for it. 3) What argues against it. 4) Possibilities I did not consider, each with its discriminator. 5) Missing history questions and exam maneuvers. 6) What I cannot afford to miss in this presentation.
The fixed structure is not fussiness. Numbered sections tell you what you are reading and where to look, and a wall of text with no addresses gets skimmed. Skimming is trusting out of fatigue.
Honesty this site owes you: not every visit has room for the full two-pass ritual, and productivity pressure is real, not a choice you are making. The ritual is for the case that earns it. When it will not fit in the visit, the invitation still stands for later, and reviewing the interesting cases of the week during study time is the perfect place for "could I have been wrong?" It has never cost so little to ask. And the principle underneath belongs to this whole page: patients should have a right to an unhurried visit, and clinicians should have a right to enough time to build the reasoning that produces good care. The tools on this map exist to give minutes back to that time. Not to compress it further.
Exam performance and clinical performance are different measurements, models lose ground when they have to run the case rather than receive it, and continuous use has a measurable cost to the clinician's own skill. All three support the same conclusion: structural supervision, not accumulated trust.
Further reading · Sharma M, et al. Towards understanding sycophancy in language models. ICLR 2024. Why agreement is the path of least resistance: arxiv.org/abs/2310.13548