Tests are not interpreted. Questions are answered. Every test was ordered to settle something, and when you hand over only the numbers, the model has no pretest probability at all: it reads the panel against a population instead of against your patient.
A panel comes back the morning of the visit. You paste it in with no context, and you get back a tidy paragraph that connects a borderline TSH, a mildly elevated ESR, and a high platelet count into one story, with four follow-up tests at the end. One complaint has become three workups, and none of them were yours.
Guiding question: What question was this test ordered to answer?
Organizes a large panel, explains what a finding usually means, compares against prior values when you provide them, and drafts the language you will use to explain the result to the patient.
Know why the test was ordered. Without the clinical question there is no pretest probability, so every deviation looks equally important.
The input on this card is the question first and the numbers second. That single reordering does most of the safety work here: it restores the prior that turns a list of deviations back into a clinical answer, and it keeps incidental findings in their own lane instead of letting them recruit the whole visit.
The model is trained to produce a coherent, complete answer, and a loose finding looks like an unfinished one. So it stitches. Each association sounds defensible alone; together they turn a visit into a systemic workup. The most expensive failure here, because it does not look like an error. It looks like diligence.
A range is a population; medicine is a delta. Potassium of 5.4 means one thing with normal kidneys and another in CKD on spironolactone. Hemoglobin of 11.5 is "mild anemia" until you learn it was 14 two months ago.
A radiologist's hedging carries information. "Cannot be excluded," "correlate clinically," "recommend follow-up in 6 months" all survive poorly in paraphrase, and an incidental finding with a follow-up recommendation becomes "no significant abnormalities."
These three link to the site's failure taxonomy, where the transcripts live.
Synthetic case. 38-year-old woman, three months of fatigue, heavier menstrual bleeding. Panel: Hb 10.9 (13.8 eight months ago), MCV 74, ferritin 8, platelets 480,000, TSH 4.8.
Microcytic anemia with thrombocytosis, a finding that may indicate chronic inflammation or, less commonly, a myeloproliferative disorder; suggest ESR, CRP, and hematology evaluation if persistent. Borderline TSH raises the possibility of thyroid dysfunction contributing to fatigue; suggest anti-TPO and free T4. Also recommend evaluation for celiac disease as a cause of iron malabsorption.
Coherent, complete, and every sentence defensible on its own. One complaint, three workups.
Main question. Yes. Microcytic anemia with a ferritin of 8 confirms iron deficiency, and the drop from 13.8 to 10.9 over eight months fits the reported menstrual loss. The thrombocytosis is an expected reactive finding in iron deficiency and tends to normalize with repletion; it does not call for a separate workup right now.
Outside the question. TSH 4.8, mildly above range, no prior value for comparison. Typically warrants repeat testing at your discretion; it does not explain the hematologic picture.
Same numbers, different medicine.
In the first reading it is a clue to a new disease. In the second it is a consequence of the main finding, explained by the very thing the test was ordered to confirm.
What changed was not the model. It was the question.
It does not disappear in the second answer, and it should not: an incidental finding still deserves a line. What it loses is protagonism, and the sentence that would have turned it into a workup.
"No prior value for comparison" is the declared absence of this card, doing its job in one clause.
The tool organizes; you read the numbers. That division is the whole verification on this card.
This test was ordered to answer the following clinical question: [the question that prompted it]. Case context: [age, relevant conditions, medications that affect interpretation]. Prior values, if any: [test: value and date]. Structure your answer in two parts. 1) Main question: answer it using the relevant data, interpreting against the prior values I gave you. Where there is no prior value, write "no prior value for comparison." 2) Findings outside the question: list them separately, saying what each one is and whether it typically warrants follow-up. Do not integrate them into the main picture. If you propose a connection between findings, label it as speculation and say what would confirm it. For reports, preserve hedges and recommendations in the original wording. [paste results with identifiers removed]copied
The gap between exam and practice applies here with force: models above 90 on knowledge exams score 44.8% on tasks built from real clinical text1, and interpreting a result in the context of one patient is exactly that kind of task.