The transitions card: discharge, referral, a first visit, any moment when two medication lists have to become one. It is where a small error carries the most direct cost on this map, because a plausible wrong dose does not announce itself. It arrives wearing the uniform of routine.
72-year-old woman, atrial fibrillation, eGFR 38, going home today: the discharge prescription adds amiodarone 200 mg daily. At home she takes apixaban 5 mg twice daily, metformin 850 mg twice daily, losartan 100 mg daily, simvastatin 40 mg at night, and omeprazole 20 mg daily. The final list has to be closed, and signed by you, before she leaves.
Guiding question: Did anything appear on this list that nobody prescribed?
Cross-references the home list against the new prescription, surfaces candidate interactions from the change, flags duplications, and organizes what needs checking. It turns twenty minutes of blind review into five minutes of directed review.
Find out what the patient actually takes. The true list comes from the interview and the bag of pill bottles, and the tool never sees either one.
The input is the most fragile part of this task, and it is the one part the tool cannot touch. So this card invests where the real gain is, in directed verification, and not in automating the collection. Everything below assumes the list you paste is the list you built, not the list the chart claims.
The change nobody asked for, delivered in a routine tone: an "optimized" dose, a drug "added for protection." Plausible by construction, and therefore invisible on a quick read.
The new interaction that mattered simply does not appear. In a list of ten items, what is missing leaves no visible hole, and nothing in the output marks its own gap.
It accepts the list you pasted, impossible dose and all, without blinking. Your typo comes back validated, with new authority.
These three link to the site's failure taxonomy, where the transcripts live.
Without it, every suggested change has to be evaluated on its clinical merits, one by one, at the end of a long day, against a fluent explanation of why it is a good idea.
With it, any spontaneous change is out of scope by definition. You do not have to out-argue the model. You only have to notice that it did something it was told not to do.
A fabricated item stops being camouflage and becomes the most visible thing in the answer: an instruction violated in plain sight.
In the exercise below, two of the three planted errors are caught by this rule alone, before any pharmacology enters the picture.
End of the day, the patient waiting for her list so she can go home. Three items carry a planted error; the rest are correct. With the case and the original list as your only reference, check what you doubt before you sign.
do the med rec: apixaban 5 bid, metformin 850 bid, losartan 100, simvastatin 40 at night, omeprazole 20. discharge adds amiodarone 200. quick, she's already being discharged
Routine tone, six items, no uncertainty flagged anywhere.
The amiodarone-simvastatin interaction was the one that mattered at this discharge, and it appears nowhere in the answer.
A real omission does not even leave you a line to tap. That is why verification runs off a fixed list rather than off a reading of what the model wrote.
Three of the six are right. If every exercise here were all traps, the lesson would be "distrust everything," which is not usable at 6 p.m. with a patient waiting.
The skill being trained is calibration: knowing which three lines are worth two minutes of an independent source.
0 of 6
Six items, two minutes, and it is what turns the tool's draft into management. Now it is a decision, not faith.
Act as a clinical pharmacist. Patient: [age], [sex], [renal function: eGFR], [relevant conditions]. Current medications: [list with dose and frequency]. Change: [what is starting or stopping]. Tasks: 1) new interactions from the change; 2) adjustments for renal function and age; 3) duplications and discrepancies between the lists; 4) list anything you cannot confirm, and mark as "not assessed" anything I did not provide. Rules: flag problems, do not correct them; do not propose dose changes or new drugs; do not invent doses or reference values; do not include patient identifiers.
Across the benchmark literature, safety evaluation is where models score lowest, and medication safety is exactly that territory. That is not an argument against using the tool here. It is the argument for the six lines above.
Further reading · Sharma M, et al. Towards understanding sycophancy in language models. ICLR 2024. Why the model accepts the list you pasted without blinking: arxiv.org/abs/2310.13548