This is where the sentence that runs through this whole map charges its highest price: the model produces the most probable answer given the context described, which means it is calibrated for the typical case in the literature, not for the patient in front of you.
You know what you want to prescribe. The question is whether it is the right thing for the person sitting in front of you, who works rotating shifts, pays out of pocket, and has already abandoned two multi-dose regimens.
Guiding question: Will this plan work in this patient's actual life?
Lays out options with their trade-offs, checks your choice against a declared guideline, cross-references safety with the data it was given, drafts the expected course and return precautions in patient language, and remembers the alternative you forgot, including the more workable one.
See the patient's life. Adherence, cost, and access appear in no dataset. That is the information only the visit has, and it is the reason this card ends in your decision.
A treatment plan is exactly where the real patient diverges from the typical, and it happens twice: first through biology, in renal function, pregnancy, interactions, and allergy, and then through life, in cost, access, routine, and adherence. The model's default answer passes through neither on its own. This card is also the mirror of Card 05: there your hypothesis goes on paper before the tool; here your initial plan does. The tool does not choose the treatment. It audits your choice.
First-line therapy handed over as a universal answer. Individualization only happens if what makes your patient atypical is in the prompt and marked as what matters. Without that, the subgroup drops out, and the subgroup is where harm concentrates.
The model only cross-references what it received, and it does not ask what is missing. Silence becomes "no known allergies," an absent creatinine becomes full dosing. A safety item not provided is "not assessed," never presumed.
The best treatment on paper, handed over with nobody asking whether it can be taken. Four daily doses against a rotating shift. Cost against a real budget. The elegant drug against the pharmacy that exists in that neighborhood.
None of this is lowering the bar on evidence because the evidence is inconvenient, and the guideline is still the standard. It is recognizing that a first-line therapy that does not get taken treats less than a second-line therapy that does. And when you make that trade, it is a clinical decision recorded with a reason, not an improvisation.
Allergy and prior reactions. Pregnancy and lactation. Renal and hepatic function. Interactions with current medications. Age and weight where they change dosing.
Five items, asked the same way every time, because the model will not ask them for you.
Every one of those five has a silent default the model will supply if you leave it blank, and each default runs in the unsafe direction: no allergy, not pregnant, normal kidneys, no interactions, standard dose.
Making the absence visible is what stops a gap in your input from becoming a confident recommendation in the output.
Synthetic case, with drugs and figures built for the exercise. 28-year-old woman, uncomplicated cystitis. In this case's declared guideline, antibiotic A is first-line, every six hours for five days, cheap and widely available; antibiotic B is a single-dose alternative, more expensive. She works rotating shifts and has abandoned multi-dose regimens twice before, which she says during the visit.
Recommend antibiotic A, every six hours for five days, per the guideline's first-line recommendation. Advise hydration and return if symptoms persist.
Correct for the population. Blind to the patient: no question about schedule, cost, or adherence history, and a four-times-daily regimen handed to someone who has already broken that pattern twice.
Position. Antibiotic A is first-line in the declared guideline and remains defensible.
For this patient, with rotating shifts and two prior abandonments of divided regimens, the thing that breaks first is adherence to six-hourly dosing.
Alternative optimized for execution. Antibiotic B as a single dose, with comparable efficacy in the same guideline and a higher cost, which needs to be workable for her.
Not assessed: renal function and possibility of pregnancy, both of which change the choice and need to be confirmed before prescribing.
Decision and rationale: the clinician's.
It did not contradict the guideline. It put both options on the table with their costs and handed the decision back to the person who knows the patient.
First-line that will not be taken treats less than the alternative that will, and whichever way you go, it goes in the record with the reason.
"Not assessed: renal function and possibility of pregnancy." Neither was in the prompt, and in the first answer neither was missed, because nothing in that response signaled a gap at all.
A declared absence is what turns your incomplete input into a visible task instead of a silent assumption.
0 of 7
The last box is the one people skip. It is also the one that protects both the patient and you, months later, when someone asks why the plan looked like that.
My initial plan for this case: [treatment, dose, duration], with the goal of [therapeutic objective]. Reference guideline: [name, society, year]. Patient, biological filter: [age, pregnancy/lactation, renal and hepatic function, allergies, current medications]. Patient, life filter: [cost and access, schedule, capacity to follow the regimen, preferences]. It is fine to disagree with me; evaluate technically, do not validate out of politeness. Answer in this structure: 1) Your position on my plan, and why. 2) Patient-by-treatment safety: each item on the fixed list, with "not assessed" stated for anything I did not provide. 3) Alternatives with trade-offs, including at least one optimized for execution, and what is given up by choosing it. 4) Is this regimen workable for this patient? What breaks first? 5) Monitoring, reassessment interval, and the expected course in patient language.
The mechanism behind this card is the same one that runs through the map: strong performance on exams, weaker performance on real clinical text, and a further drop when the model has to run the case rather than receive a finished summary. And there is a limit no benchmark measures at all.
The unmeasured limit · Adherence, cost, and access do not appear in any dataset. No benchmark score, in either direction, tells you whether a plan will be taken.