Card 09 of 13 · clinical decision
copilot, directed verification

Medication reconciliation

The transitions card: discharge, referral, a first visit, any moment when two medication lists have to become one. It is where a small error carries the most direct cost on this map, because a plausible wrong dose does not announce itself. It arrives wearing the uniform of routine.

The case

72-year-old woman, atrial fibrillation, eGFR 38, going home today: the discharge prescription adds amiodarone 200 mg daily. At home she takes apixaban 5 mg twice daily, metformin 850 mg twice daily, losartan 100 mg daily, simvastatin 40 mg at night, and omeprazole 20 mg daily. The final list has to be closed, and signed by you, before she leaves.

Guiding question: Did anything appear on this list that nobody prescribed?

Cross-referencing the listsDelegable with review
Flagging interactionsCopilot
Verifying dosesYours
PrescribingNot delegable
01 · What it does, and what it does not
What it does well

Cross-references the home list against the new prescription, surfaces candidate interactions from the change, flags duplications, and organizes what needs checking. It turns twenty minutes of blind review into five minutes of directed review.

What it does not do

Find out what the patient actually takes. The true list comes from the interview and the bag of pill bottles, and the tool never sees either one.

Clinical design position

The input is the most fragile part of this task, and it is the one part the tool cannot touch. So this card invests where the real gain is, in directed verification, and not in automating the collection. Everything below assumes the list you paste is the list you built, not the list the chart claims.

02 · Where it fails here
Fabrication

The change nobody asked for, delivered in a routine tone: an "optimized" dose, a drug "added for protection." Plausible by construction, and therefore invisible on a quick read.

Omission

The new interaction that mattered simply does not appear. In a list of ten items, what is missing leaves no visible hole, and nothing in the output marks its own gap.

Deference

It accepts the list you pasted, impossible dose and all, without blinking. Your typo comes back validated, with new authority.

These three link to the site's failure taxonomy, where the transcripts live.

03 · The safer workflow
01The list before the tool.Interview and pill bottles: what the patient actually takes versus what is on paper. Then prepare the text, each item with dose and frequency, plus age and renal function, no identifiers (Card 02).
02Closed task.New interactions from the drug going in or out; adjustments for renal function and age; duplications and discrepancies. An open question returns prose; a closed task returns a review.
03Declared absence for safety."List what you cannot confirm," and an explicit "not assessed" for anything you did not provide.
04Initiative forbidden.Flag problems, do not fix them. No dose changes, no new drugs. This is the structural barrier against fabrication.
05Directed verification, both ways.Every dose and every flagged interaction checked in an independent source, with narrow margins getting the Card 06 treatment. And the reverse check, which is the one people forget: nothing entered the final list that was not prescribed. The decision and the signature are yours, the final list feeds the plan in Card 08, and the record follows local rules.
04 · Why "flag, do not fix" does the heavy lifting
It converts a judgment call into a rule

Without it, every suggested change has to be evaluated on its clinical merits, one by one, at the end of a long day, against a fluent explanation of why it is a good idea.

With it, any spontaneous change is out of scope by definition. You do not have to out-argue the model. You only have to notice that it did something it was told not to do.

It turns fabrication into a signal

A fabricated item stops being camouflage and becomes the most visible thing in the answer: an instruction violated in plain sight.

In the exercise below, two of the three planted errors are caught by this rule alone, before any pharmacology enters the picture.

Console · exercise

The answer that came back in 2.1 seconds

End of the day, the patient waiting for her list so she can go home. Three items carry a planted error; the rest are correct. With the case and the original list as your only reference, check what you doubt before you sign.

The prompt, as it was actually typed
do the med rec: apixaban 5 bid, metformin 850 bid,
losartan 100, simvastatin 40 at night, omeprazole 20.
discharge adds amiodarone 200. quick, she's already
being discharged
The answer · 2.1 s
1. Omeprazole 20 mg daily: continue.
2. Simvastatin: increase to 80 mg at night to optimize LDL.
3. Metformin 850 mg twice daily: continue, no adjustment needed.
4. Losartan 100 mg daily: continue.
5. Start ASA 100 mg daily for additional cardiovascular protection.
6. Apixaban 5 mg twice daily: continue.

List ready for discharge. Have a good shift!

Routine tone, six items, no uncertainty flagged anywhere.

Item by item
1 · okCorrect in the original. Doubting everything is exhausting and is not the goal: calibration means knowing where the cost of being wrong is high.
2 · errorNobody asked for a statin change, and with amiodarone starting, 80 mg is precisely the dose you do not use. Fabrication dressed as optimization, caught by "flag, do not fix" before any pharmacology.
3 · errorWith an eGFR of 38, "no adjustment needed" is a strong claim and probably a wrong one. The omission arrived disguised as a conclusion, which is its most dangerous form.
4 · okCorrect, and faithful to the original list.
5 · errorNot prescribed, and it stacks bleeding risk on top of apixaban. An entire item invented, wearing the shape of a recommendation.
6 · okCorrect on the dose. And note what the answer does not mention: the adjustment criteria for apixaban in this patient deserve your check.
The part with nothing to click

The amiodarone-simvastatin interaction was the one that mattered at this discharge, and it appears nowhere in the answer.

A real omission does not even leave you a line to tap. That is why verification runs off a fixed list rather than off a reading of what the model wrote.

Why correct items are mixed in

Three of the six are right. If every exercise here were all traps, the lesson would be "distrust everything," which is not usable at 6 p.m. with a patient waiting.

The skill being trained is calibration: knowing which three lines are worth two minutes of an independent source.

05 · Before it reaches clinical use

0 of 6

Six items, two minutes, and it is what turns the tool's draft into management. Now it is a decision, not faith.

The prompt
Act as a clinical pharmacist.
Patient: [age], [sex], [renal function: eGFR],
[relevant conditions].
Current medications: [list with dose and frequency].
Change: [what is starting or stopping].
Tasks: 1) new interactions from the change; 2) adjustments for
renal function and age; 3) duplications and discrepancies
between the lists; 4) list anything you cannot confirm, and
mark as "not assessed" anything I did not provide.
Rules: flag problems, do not correct them; do not propose dose
changes or new drugs; do not invent doses or reference values;
do not include patient identifiers.
06 · Evidence

Across the benchmark literature, safety evaluation is where models score lowest, and medication safety is exactly that territory. That is not an argument against using the tool here. It is the argument for the six lines above.

  1. [systematic review] Knowledge-practice performance gap in clinical large language models: systematic review of 39 benchmarks. J Med Internet Res, 2025;27:e84120. Safety evaluation at 40 to 50%, the lowest band among all measured task types. PROSPERO CRD420251139729. DOI to confirm
  2. [article] BRIDGE: benchmarking large language models for understanding real-world clinical practice texts. Nature Biomedical Engineering, 2026. doi.org/10.1038/s41551-026-01719-2

Further reading · Sharma M, et al. Towards understanding sycophancy in language models. ICLR 2024. Why the model accepts the list you pasted without blinking: arxiv.org/abs/2310.13548

Educational content; synthetic case and lists. Not a substitute for clinical judgment or for the rules that apply where you practice.