Card 13 of 13 · clinical decision
outside the spectrum

Limits: images, ECGs, and tracings

Every honest spectrum has an endpoint, and it is the endpoint that makes the rest of it credible. This map can say “delegate with review” on twelve cards because it knows where to say “do not delegate.”

The case

None, on purpose. This card also has no exercise and no copy-paste prompt. There is no prompt to copy because the problem is not the prompt: it is the input, the validation, and the verification, none of which meet the intended clinical use. The absence is the content.

Guiding question: If I cannot read this myself, how would I check the answer?

Reading an image in a general chatOutside the spectrum
ECGs and tracings pasted as photosOutside the spectrum
Working from the written reportCopilot · Card 03
A second opinion on imagingAnother human, not another tab
01 · Why not
Performance does not support it

On unambiguous radiologic findings, a frontier multimodal model reached single-digit diagnostic accuracy without clinical context, and under a third with it. And it comes in the worst possible wrapper: the model identifies the modality perfectly, knows this is an ECG, a CT, a chest film, and gets the diagnosis wrong in the vocabulary of a radiology report. It looks like a reading. It is a caption.

What you paste is not the study

It is a compressed photo of a screen: no window and level, no full series, no comparison with the prior study, at whatever resolution the messaging app left behind. No radiologist would report from that artifact, and the problem starts before the model does. The input was already below diagnostic standard.

There is no verification path

On every other card the output can be checked against a source: the list, the report, the document, the case. Here the source is the image itself. Either you are qualified to read it, in which case you do not need the chat, or you are not, in which case you cannot check anything it tells you.

Images carry identity inside them

DICOM headers, a name burned into the corner of a screen capture, a face or identifying marks in a clinical photograph. A black box drawn on top does not remove what is inside the file (Card 02), and there is no "paste less" here: the identifiable image is the entire product.

The rule this card protects

What is delegable is what you can verify. Without a verification path there is no copilot, only blind outsourcing. What cannot be verified is not delegated. It is referred.

02 · The right boundary
This card is not against AI in imaging

Dedicated tools exist, validated, cleared as devices, running inside an institutional workflow with a named responsible party. This card is not about them.

It is about a general-purpose chat receiving a photo from your phone. The boundary is not technology versus medicine. It is device versus chat window.

The three-part test

Published clinical validation. Regulatory clearance. A workflow with an accountable owner, for the specific use in question.

An impressive demo is none of the three, and a model that scores well on a public benchmark is still none of the three.

03 · What you can still do
01Work from the report. That is Card 03, and the evidence below shows why: it is the written description that carries diagnostic accuracy, not the pixels.
02Prepare the question before ordering. What you want the radiologist to answer, and the clinical context that lets them answer it.
03Discuss a described finding. A reported finding can enter your differential like any other piece of text (Card 05).
04Structure the referral of the question. If it needs another set of eyes, the letter is Card 10, and the eyes belong to a person.
Console · the asymmetry

Right modality, wrong diagnosis

There is no exercise on this card, so this console shows the numbers instead. They describe the same failure from three directions, and together they explain why confidence is not a signal here.

What the studies measured
100%Modality recognition. The model knows it is looking at an ECG, a CT, a chest film. Every time.
8.3%Diagnostic accuracy from the image alone, across 206 studies with unambiguous findings from a university hospital.
29.1%With clinical context added. Better, and still not a reading you could act on.
ρ ≈ 0Self-reported confidence did not correlate with accuracy. The tone of certainty carries no information at all.
Why that combination is the dangerous one

A tool that failed visibly would be safe: you would stop using it after the first attempt.

This one succeeds at the part you can check and fails at the part you cannot. You confirm it identified the study correctly, the vocabulary is right, the structure looks like a report, and none of that predicts whether the finding is real.

And the last line closes the loop with the fourth principle of this map: verification follows risk, not the confidence you perceive in the answer.

The text keeps winning

In work testing what actually drives diagnostic accuracy in multimodal cases, the written description of the findings contributed far more than the images themselves.

The safe path was never a workaround. It was the better-performing one.

When this changes

This is the card with the shortest shelf life on the map, and it says so out loud. The test for changing position is not a product launch. It is the three-part boundary above, for the specific use in question.

When that exists, this card changes, with a date. Until then, the endpoint stands, and it is what holds up every "yes, with review" on the rest of the page.

04 · Evidence

Three findings, one conclusion: the model recognizes what it is looking at, does not reliably read it, and gives no usable signal about which of the two is happening in front of you.

  1. [article] Huppertz MS, et al. Revolution or risk? Assessing the potential and challenges of GPT-4V in radiologic image interpretation. European Radiology, 2025;35:1111-1121. 206 studies with unambiguous findings; 8.3% diagnostic accuracy from the image alone, 29.1% with clinical context; self-reported confidence did not correlate with accuracy. doi.org/10.1007/s00330-024-11115-6
  2. [article] Effectiveness of GPT-4o in interpreting electrocardiogram images, 2025. Modality recognition reaches 100% while diagnostic interpretation remains a separate problem. citation to complete
  3. [preprint] Impact of multimodal prompt elements on diagnostic performance in challenging brain MRI cases, 2024. The textual description of the findings is by far the largest contributor to diagnostic accuracy. doi.org/10.1101/2024.03.05.24303767