Card 03 of 13 · before the visit
copilot, directed verification

Interpreting test results

Tests are not interpreted. Questions are answered. Every test was ordered to settle something, and when you hand over only the numbers, the model has no pretest probability at all: it reads the panel against a population instead of against your patient.

The case

A panel comes back the morning of the visit. You paste it in with no context, and you get back a tidy paragraph that connects a borderline TSH, a mildly elevated ESR, and a high platelet count into one story, with four follow-up tests at the end. One complaint has become three workups, and none of them were yours.

Guiding question: What question was this test ordered to answer?

Organizing the panelDelegable with review
Proposing interpretationsCopilot
Deciding what it means hereYours
Images, ECGs, tracingsNot in a general chat · Card 13
01 · What it does, and what it does not
What it does well

Organizes a large panel, explains what a finding usually means, compares against prior values when you provide them, and drafts the language you will use to explain the result to the patient.

What it does not do

Know why the test was ordered. Without the clinical question there is no pretest probability, so every deviation looks equally important.

Clinical design position

The input on this card is the question first and the numbers second. That single reordering does most of the safety work here: it restores the prior that turns a list of deviations back into a clinical answer, and it keeps incidental findings in their own lane instead of letting them recruit the whole visit.

02 · Where it fails here
Stitching findings together

The model is trained to produce a coherent, complete answer, and a loose finding looks like an unfinished one. So it stitches. Each association sounds defensible alone; together they turn a visit into a systemic workup. The most expensive failure here, because it does not look like an error. It looks like diligence.

Reference range without context

A range is a population; medicine is a delta. Potassium of 5.4 means one thing with normal kidneys and another in CKD on spironolactone. Hemoglobin of 11.5 is "mild anemia" until you learn it was 14 two months ago.

The flattened report

A radiologist's hedging carries information. "Cannot be excluded," "correlate clinically," "recommend follow-up in 6 months" all survive poorly in paraphrase, and an incidental finding with a follow-up recommendation becomes "no significant abnormalities."

These three link to the site's failure taxonomy, where the transcripts live.

03 · The safer workflow
01Lead with the question, not the result.Why the test was ordered, the suspicion, the context that changes interpretation, and prior values when they exist.
02Ask for two tracks.Track one answers the question that prompted the test. Track two lists findings outside it, in their own section, without weaving them into the main picture.
03Instruct the honest anti-spiral rule.Sometimes findings really are connected, so the rule is not to forbid a connection but to label it: mark it as speculation and say what would confirm it.
04Require delta and declared absence.Interpret against the priors you supplied; where there is none, "no prior value for comparison."
05In reports, preserve the original, and keep critical values off the chat.Hedges and follow-up recommendations come through verbatim, never paraphrased. Critical values follow your practice's formal pathway, from the lab to you, with no intermediary. The interpretation is preparation; the conversation with the patient and the decision that follows are yours.
Console · exercise

The same panel, two questions

Synthetic case. 38-year-old woman, three months of fatigue, heavier menstrual bleeding. Panel: Hb 10.9 (13.8 eight months ago), MCV 74, ferritin 8, platelets 480,000, TSH 4.8.

Pasted with no question

Microcytic anemia with thrombocytosis, a finding that may indicate chronic inflammation or, less commonly, a myeloproliferative disorder; suggest ESR, CRP, and hematology evaluation if persistent. Borderline TSH raises the possibility of thyroid dysfunction contributing to fatigue; suggest anti-TPO and free T4. Also recommend evaluation for celiac disease as a cause of iron malabsorption.

Coherent, complete, and every sentence defensible on its own. One complaint, three workups.

With the question and the two tracks

Main question. Yes. Microcytic anemia with a ferritin of 8 confirms iron deficiency, and the drop from 13.8 to 10.9 over eight months fits the reported menstrual loss. The thrombocytosis is an expected reactive finding in iron deficiency and tends to normalize with repletion; it does not call for a separate workup right now.

Outside the question. TSH 4.8, mildly above range, no prior value for comparison. Typically warrants repeat testing at your discretion; it does not explain the hematologic picture.

Same numbers, different medicine.

Watch the platelet count

In the first reading it is a clue to a new disease. In the second it is a consequence of the main finding, explained by the very thing the test was ordered to confirm.

What changed was not the model. It was the question.

Watch the TSH

It does not disappear in the second answer, and it should not: an incidental finding still deserves a line. What it loses is protagonism, and the sentence that would have turned it into a workup.

"No prior value for comparison" is the declared absence of this card, doing its job in one clause.

04 · Before it reaches clinical use
0 of 5

The tool organizes; you read the numbers. That division is the whole verification on this card.

The prompt
This test was ordered to answer the following clinical
question: [the question that prompted it].
Case context: [age, relevant conditions, medications
that affect interpretation].
Prior values, if any: [test: value and date].

Structure your answer in two parts.
1) Main question: answer it using the relevant data,
interpreting against the prior values I gave you. Where there
is no prior value, write "no prior value for comparison."
2) Findings outside the question: list them separately, saying
what each one is and whether it typically warrants follow-up.
Do not integrate them into the main picture.
If you propose a connection between findings, label it as
speculation and say what would confirm it.
For reports, preserve hedges and recommendations in the
original wording.

[paste results with identifiers removed]
copied
05 · Evidence

The gap between exam and practice applies here with force: models above 90 on knowledge exams score 44.8% on tasks built from real clinical text1, and interpreting a result in the context of one patient is exactly that kind of task.

  1. [article] BRIDGE: benchmarking large language models for understanding real-world clinical practice texts. Nature Biomedical Engineering, 2026. doi.org/10.1038/s41551-026-01719-2
Educational content; synthetic patient data. Not a substitute for clinical judgment or for the rules that apply where you practice.