Every honest spectrum has an endpoint, and it is the endpoint that makes the rest of it credible. This map can say “delegate with review” on twelve cards because it knows where to say “do not delegate.”
None, on purpose. This card also has no exercise and no copy-paste prompt. There is no prompt to copy because the problem is not the prompt: it is the input, the validation, and the verification, none of which meet the intended clinical use. The absence is the content.
Guiding question: If I cannot read this myself, how would I check the answer?
On unambiguous radiologic findings, a frontier multimodal model reached single-digit diagnostic accuracy without clinical context, and under a third with it. And it comes in the worst possible wrapper: the model identifies the modality perfectly, knows this is an ECG, a CT, a chest film, and gets the diagnosis wrong in the vocabulary of a radiology report. It looks like a reading. It is a caption.
It is a compressed photo of a screen: no window and level, no full series, no comparison with the prior study, at whatever resolution the messaging app left behind. No radiologist would report from that artifact, and the problem starts before the model does. The input was already below diagnostic standard.
On every other card the output can be checked against a source: the list, the report, the document, the case. Here the source is the image itself. Either you are qualified to read it, in which case you do not need the chat, or you are not, in which case you cannot check anything it tells you.
DICOM headers, a name burned into the corner of a screen capture, a face or identifying marks in a clinical photograph. A black box drawn on top does not remove what is inside the file (Card 02), and there is no "paste less" here: the identifiable image is the entire product.
What is delegable is what you can verify. Without a verification path there is no copilot, only blind outsourcing. What cannot be verified is not delegated. It is referred.
Dedicated tools exist, validated, cleared as devices, running inside an institutional workflow with a named responsible party. This card is not about them.
It is about a general-purpose chat receiving a photo from your phone. The boundary is not technology versus medicine. It is device versus chat window.
Published clinical validation. Regulatory clearance. A workflow with an accountable owner, for the specific use in question.
An impressive demo is none of the three, and a model that scores well on a public benchmark is still none of the three.
There is no exercise on this card, so this console shows the numbers instead. They describe the same failure from three directions, and together they explain why confidence is not a signal here.
A tool that failed visibly would be safe: you would stop using it after the first attempt.
This one succeeds at the part you can check and fails at the part you cannot. You confirm it identified the study correctly, the vocabulary is right, the structure looks like a report, and none of that predicts whether the finding is real.
And the last line closes the loop with the fourth principle of this map: verification follows risk, not the confidence you perceive in the answer.
In work testing what actually drives diagnostic accuracy in multimodal cases, the written description of the findings contributed far more than the images themselves.
The safe path was never a workaround. It was the better-performing one.
This is the card with the shortest shelf life on the map, and it says so out loud. The test for changing position is not a product launch. It is the three-part boundary above, for the specific use in question.
When that exists, this card changes, with a date. Until then, the endpoint stands, and it is what holds up every "yes, with review" on the rest of the page.
Three findings, one conclusion: the model recognizes what it is looking at, does not reliably read it, and gives no usable signal about which of the two is happening in front of you.