Card 06 of 13 · clinical decision
copilot, directed verification

Clinical scores and calculations

A number carries the look of objectivity. A score with a decimal point seems verified by nature, and that visual authority is exactly what hides the problem: not the arithmetic, but the path you cannot see and the criteria that were filled in before it.

The case

The patient from Card 05 is still in front of you, and the next step depends on a score. The model returns one instantly, with a decimal point and a recommendation attached. It looks like the most objective thing on the screen.

Guiding question: Can I see how this number was built?

Selecting the scoreCopilot
Filling in the criteriaYours
Running the arithmeticOpen in chat, or a validated calculator
Acting on the resultNot delegable
01 · What it does, and what it does not
What it does well

Remembers which score fits your question, lists the criteria with the exact definition of each item, names the version and the population it was validated in, interprets the result against a guideline, and flags where the score is weak for this particular patient. It can also calculate with the work shown, which is what makes a number auditable.

What it does not do

Show you how it got there unless you ask. And it will not tell you that a criterion was filled in by inference rather than by evidence.

Clinical design position

The problem here is not that models cannot add; current models usually can. The problem is that a chat answer hides the path, and most of the error enters before any arithmetic, in how the criteria were filled. So the division of labor is: the tool remembers, structures, and interprets; the arithmetic is either shown openly or reborn in a validated calculator; the judgment criteria are yours.

02 · Where it fails here
A number with no receipt

A chat answer does not show whether it came from real computation or from pattern completion, and that varies by product and version. A closed number cannot be trusted or doubted; it cannot be evaluated at all.

Version and blending

Training freezes; guidelines move. An outdated score, criteria from one Wells score bleeding into another, a threshold swapped, a cutoff invented with the look of a citation.

Inferred criteria

The error that enters before the arithmetic. Objective items misextracted: a heart rate of 96 marked as "tachycardia," a long flight marked as "immobilization" against the strict definition. And judgment items filled in by inference, which is the sneakier one.

These three link to the site's failure taxonomy, where the transcripts live.

03 · Scores that depend on the conversation
Why they are different

Instruments that depend on what was discussed with the patient, such as severity scales in mental health, get scored from a note that does not contain everything the conversation contained. A note is compression (Card 04); a score computed on top of it inherits that compression and hands it back wearing numerical authority.

Worse, some instruments are self-report by design. Filling them in by inference from documentation is not a shortcut. It is a different measurement, without validity.

The position of this card

A score that depends on clinical reasoning and on a conversation is completed by the clinician, in full, without the tool. The result classifies severity, and severity drives the plan: a condition that becomes "severe" on paper changes the entire treatment path.

If the tool participates anyway, the safeguards in the workflow below are the minimum, and they are harm reduction, not endorsement.

04 · The safer workflow
01The tool selects and sets up, without calculating.The score that fits your question, with version, validation population, and the criteria list with each item's exact definition.
02You fill in the criteria.Filling in a criterion is clinical judgment, not arithmetic. Any item with no information goes in as "not assessable," never assumed.
03The arithmetic, with a receipt.Open in the chat when stakes are low: formula named, values substituted, steps, units. In narrow margins, a validated calculator as an independent second pass.
04If the tool filled in any item, it declares it.Item by item: what it used and where it came from. Your check is against the case, not against the sum.
05The tool comes back to interpret, and a spontaneous number stays unverified.The result against the declared guideline, the cutoff with a source, and the score's limitations in this patient. And when a score appears mid-answer, as it can in a Card 05 conversation, it was produced with no check at all: it does not travel into management or the chart without going through the steps above.
Console · exercise

Did the model get it right?

The patient from Card 05 carries over: chest pain, combined contraceptive, an eleven-hour flight last week, HR 96. Check the answer item by item against the definitions, not against the sum.

The model's answer

Wells score for PE: PE is the most likely diagnosis (+3); elevated heart rate (+1.5); recent immobilization from prolonged travel (+1.5). Total: 6 points, high probability. Proceed directly to CT pulmonary angiography, skip the D-dimer.

One decimal point, one recommendation, no visible path.

Item by item, against the definitions
+1.5HR 96 scored as tachycardia. The item requires a rate above 100. Does not score.
+1.5The flight scored as immobilization. The strict definition is immobilization for three days or more, or surgery within four weeks. A long flight is a clinical risk factor, but it does not meet this criterion in this score. Does not score.
cutoff"High probability" at 6 points. In the three-tier reading of Wells, low is under 2, moderate is 2 to 6, and high is above 6. Six is moderate. The label came from blending two classification schemes.
redone3 points, moderate probability, D-dimer first. Three small slips, an entirely different course of action, with radiation and contrast in between.
What this exercise does not contain

None of the three failures was the arithmetic. They all entered through the criteria and the cutoff, which is exactly where this card tells you to look, and exactly what a validated calculator would not have caught either: feed it the same wrong criteria and it returns the same wrong score, faster.

Why the receipt matters

Had the answer arrived with the formula named and the values substituted, every one of these three would have been visible in seconds, without any pharmacology or clinical reasoning at all.

That is the whole argument for asking the tool to show its work: it moves the error from invisible to obvious.

05 · Before it changes your management

0 of 6

Validated calculators for the second pass: MDCalc, QxMD Calculate, or your institution's own. Narrow margins, and anything going into the chart, always get one.

The prompt
Clinical question: [what decision this score will
support].
1) Name the appropriate score, with version and the population
it was validated in.
2) List the criteria with the exact definition of each item.
Do not calculate yet.
[you fill in the criteria]
3) With the criteria I provide: if you calculate, show the
work, with the formula named, values substituted, steps, and
units.
4) For any item you filled in or assumed, list item by item
what you used and where it came from. Items with no
information are "not assessable," never assumed.
5) Interpret the result against [guideline, if any],
citing the cutoff with a source, and name this score's
limitations in this case.
06 · Evidence

There is a dedicated benchmark for this, and what it measures is precisely the three-part split this card is built on: knowing the rule, extracting the parameters, and doing the arithmetic. Model scores on it have improved since release, and the principle does not depend on this year's number: an auditable path, human criteria, and a second pass in narrow margins remain the rule.

  1. [benchmark] MedCalc-Bench: evaluating large language models for medical calculations. NeurIPS 2024, Datasets & Benchmarks. 1,047 instances across 55 widely used clinical calculators; the best result at release was around 51%, with errors spread across the three steps. arxiv.org/abs/2406.12036
  2. [preprint] From scores to steps: diagnosing and improving LLM performance in evidence-based medical calculations, 2025. Scoring formula, extraction, and arithmetic separately drops measured accuracy from 62.7% to 43.6%, because tolerance in the final answer masks errors along the way. arxiv.org/abs/2509.16584
Educational content; synthetic case. Not a substitute for clinical judgment or for the rules that apply where you practice.