A number carries the look of objectivity. A score with a decimal point seems verified by nature, and that visual authority is exactly what hides the problem: not the arithmetic, but the path you cannot see and the criteria that were filled in before it.
The patient from Card 05 is still in front of you, and the next step depends on a score. The model returns one instantly, with a decimal point and a recommendation attached. It looks like the most objective thing on the screen.
Guiding question: Can I see how this number was built?
Remembers which score fits your question, lists the criteria with the exact definition of each item, names the version and the population it was validated in, interprets the result against a guideline, and flags where the score is weak for this particular patient. It can also calculate with the work shown, which is what makes a number auditable.
Show you how it got there unless you ask. And it will not tell you that a criterion was filled in by inference rather than by evidence.
The problem here is not that models cannot add; current models usually can. The problem is that a chat answer hides the path, and most of the error enters before any arithmetic, in how the criteria were filled. So the division of labor is: the tool remembers, structures, and interprets; the arithmetic is either shown openly or reborn in a validated calculator; the judgment criteria are yours.
A chat answer does not show whether it came from real computation or from pattern completion, and that varies by product and version. A closed number cannot be trusted or doubted; it cannot be evaluated at all.
Training freezes; guidelines move. An outdated score, criteria from one Wells score bleeding into another, a threshold swapped, a cutoff invented with the look of a citation.
The error that enters before the arithmetic. Objective items misextracted: a heart rate of 96 marked as "tachycardia," a long flight marked as "immobilization" against the strict definition. And judgment items filled in by inference, which is the sneakier one.
These three link to the site's failure taxonomy, where the transcripts live.
Instruments that depend on what was discussed with the patient, such as severity scales in mental health, get scored from a note that does not contain everything the conversation contained. A note is compression (Card 04); a score computed on top of it inherits that compression and hands it back wearing numerical authority.
Worse, some instruments are self-report by design. Filling them in by inference from documentation is not a shortcut. It is a different measurement, without validity.
A score that depends on clinical reasoning and on a conversation is completed by the clinician, in full, without the tool. The result classifies severity, and severity drives the plan: a condition that becomes "severe" on paper changes the entire treatment path.
If the tool participates anyway, the safeguards in the workflow below are the minimum, and they are harm reduction, not endorsement.
The patient from Card 05 carries over: chest pain, combined contraceptive, an eleven-hour flight last week, HR 96. Check the answer item by item against the definitions, not against the sum.
Wells score for PE: PE is the most likely diagnosis (+3); elevated heart rate (+1.5); recent immobilization from prolonged travel (+1.5). Total: 6 points, high probability. Proceed directly to CT pulmonary angiography, skip the D-dimer.
One decimal point, one recommendation, no visible path.
None of the three failures was the arithmetic. They all entered through the criteria and the cutoff, which is exactly where this card tells you to look, and exactly what a validated calculator would not have caught either: feed it the same wrong criteria and it returns the same wrong score, faster.
Had the answer arrived with the formula named and the values substituted, every one of these three would have been visible in seconds, without any pharmacology or clinical reasoning at all.
That is the whole argument for asking the tool to show its work: it moves the error from invisible to obvious.
0 of 6
Validated calculators for the second pass: MDCalc, QxMD Calculate, or your institution's own. Narrow margins, and anything going into the chart, always get one.
Clinical question: [what decision this score will support]. 1) Name the appropriate score, with version and the population it was validated in. 2) List the criteria with the exact definition of each item. Do not calculate yet. [you fill in the criteria] 3) With the criteria I provide: if you calculate, show the work, with the formula named, values substituted, steps, and units. 4) For any item you filled in or assumed, list item by item what you used and where it came from. Items with no information are "not assessable," never assumed. 5) Interpret the result against [guideline, if any], citing the cutoff with a source, and name this score's limitations in this case.
There is a dedicated benchmark for this, and what it measures is precisely the three-part split this card is built on: knowing the rule, extracting the parameters, and doing the arithmetic. Model scores on it have improved since release, and the principle does not depend on this year's number: an auditable path, human criteria, and a second pass in narrow margins remain the rule.