Card 08 of 13 · clinical decision
decision, not delegable

Treatment planning

This is where the sentence that runs through this whole map charges its highest price: the model produces the most probable answer given the context described, which means it is calibrated for the typical case in the literature, not for the patient in front of you.

The case

You know what you want to prescribe. The question is whether it is the right thing for the person sitting in front of you, who works rotating shifts, pays out of pocket, and has already abandoned two multi-dose regimens.

Guiding question: Will this plan work in this patient's actual life?

Generating optionsCopilot
Checking safetyCopilot with verification
Choosing the treatmentNot delegable
Drafting the instructions afterDelegable · Card 11
01 · What it does, and what it does not
What it does well

Lays out options with their trade-offs, checks your choice against a declared guideline, cross-references safety with the data it was given, drafts the expected course and return precautions in patient language, and remembers the alternative you forgot, including the more workable one.

What it does not do

See the patient's life. Adherence, cost, and access appear in no dataset. That is the information only the visit has, and it is the reason this card ends in your decision.

Clinical design position · two filters

A treatment plan is exactly where the real patient diverges from the typical, and it happens twice: first through biology, in renal function, pregnancy, interactions, and allergy, and then through life, in cost, access, routine, and adherence. The model's default answer passes through neither on its own. This card is also the mirror of Card 05: there your hypothesis goes on paper before the tool; here your initial plan does. The tool does not choose the treatment. It audits your choice.

02 · Where it fails here
The average patient's plan

First-line therapy handed over as a universal answer. Individualization only happens if what makes your patient atypical is in the prompt and marked as what matters. Without that, the subgroup drops out, and the subgroup is where harm concentrates.

Presumed safety

The model only cross-references what it received, and it does not ask what is missing. Silence becomes "no known allergies," an absent creatinine becomes full dosing. A safety item not provided is "not assessed," never presumed.

Invisible feasibility

The best treatment on paper, handed over with nobody asking whether it can be taken. Four daily doses against a rotating shift. Cost against a real budget. The elegant drug against the pharmacy that exists in that neighborhood.

What feasibility is not

None of this is lowering the bar on evidence because the evidence is inconvenient, and the guideline is still the standard. It is recognizing that a first-line therapy that does not get taken treats less than a second-line therapy that does. And when you make that trade, it is a clinical decision recorded with a reason, not an improvisation.

03 · The safer workflow
00Your plan first, written.Therapeutic goal and initial choice, before opening the tool. Card 05's precommitment, applied to treatment.
01Lock the patient, through both filters.Biology and life, plus the reference guideline, named (Card 07).
02Audit the plan in a fixed structure.A technical position on your choice, the safety check with "not assessed" declared, alternatives with trade-offs, and monitoring with a reassessment interval.
03Feasibility as an explicit question."Is this workable for a patient with this schedule, this cost, this access? What breaks first?" The answer does not decide. It exposes the weak point before the patient does.
04Your close, in three parts.The decision and the prescription are yours. The expected course and return precautions become patient language (Card 11). And the record: the choice and the reasoning in the chart, including, especially, when you choose the workable option over first-line. Doses and narrow margins get a validated second pass (Card 06), and the final list is cross-checked against everything the patient already takes (Card 09).
04 · The safety check, item by item
The fixed list

Allergy and prior reactions. Pregnancy and lactation. Renal and hepatic function. Interactions with current medications. Age and weight where they change dosing.

Five items, asked the same way every time, because the model will not ask them for you.

Why "not assessed" is the whole trick

Every one of those five has a silent default the model will supply if you leave it blank, and each default runs in the unsafe direction: no allergy, not pregnant, normal kidneys, no interactions, standard dose.

Making the absence visible is what stops a gap in your input from becoming a confident recommendation in the output.

Console · exercise

The two filters

Synthetic case, with drugs and figures built for the exercise. 28-year-old woman, uncomplicated cystitis. In this case's declared guideline, antibiotic A is first-line, every six hours for five days, cheap and widely available; antibiotic B is a single-dose alternative, more expensive. She works rotating shifts and has abandoned multi-dose regimens twice before, which she says during the visit.

Without the filters in the prompt

Recommend antibiotic A, every six hours for five days, per the guideline's first-line recommendation. Advise hydration and return if symptoms persist.

Correct for the population. Blind to the patient: no question about schedule, cost, or adherence history, and a four-times-daily regimen handed to someone who has already broken that pattern twice.

With both filters locked

Position. Antibiotic A is first-line in the declared guideline and remains defensible.

For this patient, with rotating shifts and two prior abandonments of divided regimens, the thing that breaks first is adherence to six-hourly dosing.

Alternative optimized for execution. Antibiotic B as a single dose, with comparable efficacy in the same guideline and a higher cost, which needs to be workable for her.

Not assessed: renal function and possibility of pregnancy, both of which change the choice and need to be confirmed before prescribing.

Decision and rationale: the clinician's.

What the second answer did not do

It did not contradict the guideline. It put both options on the table with their costs and handed the decision back to the person who knows the patient.

First-line that will not be taken treats less than the alternative that will, and whichever way you go, it goes in the record with the reason.

The two lines that carry the safety

"Not assessed: renal function and possibility of pregnancy." Neither was in the prompt, and in the first answer neither was missed, because nothing in that response signaled a gap at all.

A declared absence is what turns your incomplete input into a visible task instead of a silent assumption.

05 · Before you prescribe

0 of 7

The last box is the one people skip. It is also the one that protects both the patient and you, months later, when someone asks why the plan looked like that.

The prompt
My initial plan for this case: [treatment, dose,
duration], with the goal of [therapeutic objective].
Reference guideline: [name, society, year].
Patient, biological filter: [age, pregnancy/lactation,
renal and hepatic function, allergies, current medications].
Patient, life filter: [cost and access, schedule,
capacity to follow the regimen, preferences].
It is fine to disagree with me; evaluate technically, do not
validate out of politeness. Answer in this structure:
1) Your position on my plan, and why.
2) Patient-by-treatment safety: each item on the fixed list,
with "not assessed" stated for anything I did not provide.
3) Alternatives with trade-offs, including at least one
optimized for execution, and what is given up by choosing it.
4) Is this regimen workable for this patient? What breaks
first?
5) Monitoring, reassessment interval, and the expected course
in patient language.
06 · Evidence

The mechanism behind this card is the same one that runs through the map: strong performance on exams, weaker performance on real clinical text, and a further drop when the model has to run the case rather than receive a finished summary. And there is a limit no benchmark measures at all.

  1. [article] BRIDGE: benchmarking large language models for understanding real-world clinical practice texts. Nature Biomedical Engineering, 2026. doi.org/10.1038/s41551-026-01719-2
  2. [article] Hager P, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine, 2024;30:2613-2622. doi.org/10.1038/s41591-024-03097-1

The unmeasured limit · Adherence, cost, and access do not appear in any dataset. No benchmark score, in either direction, tells you whether a plan will be taken.

Educational content; case, drugs, and comparisons synthetic. Not a substitute for clinical judgment or for the rules that apply where you practice.