Card 07 of 13 · clinical decision
copilot, directed verification

Guideline retrieval

The right question is never "what does the guideline say?" It is: what does guideline X, version Y, say for population Z? An honest calibration first: today’s tools, with live search, are far better than the era of invented citations. The risk did not disappear. It moved.

The case

You need the current recommendation, and you need it before the patient finishes getting dressed. What comes back looks exactly like a guideline: class of recommendation, level of evidence, section number. All of it can be reconstructed.

Guiding question: Which document, which version, which population?

Locating the recommendationCopilot
Comparing societiesCopilot
Reading it at the sourceYours
Deciding which rule governsNot delegable
01 · What it does, and what it does not
What it does well

Finds the right recommendation in minutes, translates it into decision language, compares societies when you ask, and points to exactly where each statement lives. With live search and sources required, it is an excellent librarian. What it hands you is an address, not a reading.

What it does not do

Decide which guidance governs your patient, or notice that your patient belongs to a subgroup with a rule of its own.

Clinical design position

This card does not pretend it is 2023. Fabrication rates fell sharply between model generations, and the failure that replaced them is subtler and harder to catch: substituting the document and blending populations. Retrieval that ends in the chat answer is not retrieval. It ends in the PDF.

02 · Where it fails here
The averaged guideline

Answering from memory, the model returns an average of everything it read: old versions fused with new, societies mixed, an adult recommendation leaking into pregnancy. It sounds like a guideline and is none of them, with class, level, and section number attached. Live search reduces this. It does not eliminate it.

The substitute source

What comes back is often not the guideline but something about it: an article that cites it and concludes differently, a review that qualifies it, a site that summarizes it. The answer inherits the document's authority with the intermediary's content.

Wrong granularity

Guidelines carry general recommendations and subgroup rules, and fast reading errs both ways: applying a subgroup threshold to a patient outside it, or applying the general rule to a patient who had a rule of her own by age, pregnancy, or renal function.

These three link to the site's failure taxonomy, where the transcripts live.

The variant that sits between the second and the third

A real citation with swapped content: the document exists, the section exists, and the recommendation attributed to it is wrong. It survives the lazy check of "it exists, so it's fine," which is why the verification on this card is reading the sentence rather than confirming the reference.

03 · The safer workflow
01Ask with a surname.Name the guideline, society, and year when you know them; when you do not, ask for the most recent from that society. Either way, require the answer to declare which document and which year it is using.
02Require the passage.For each recommendation, the excerpt and the section it lives in. Plus the honesty instruction: if you do not have the exact text, say so instead of reconstructing from memory.
03Ask what kind of source this is.From the guideline document, or from something citing it? A source about the guideline does not replace the guideline, however good it is.
04Lock the population.State your patient and ask for both layers: the general recommendation and any applicable subgroup rule, with the model saying which layer this patient falls into, and why.
05Read anything that changes management at the source.A recommendation that alters a prescription, a workup, or follow-up gets read in the official document, every time. The excerpt from step two is the address for that check, not a substitute for it. And local context lives once, in your persistent instructions (Card 01): societies disagree with each other, and which one governs in your setting is a matter of context, not retrieval.
Console · exercise

Which of these would you trust?

Four guideline-shaped answers about anticoagulation in atrial fibrillation. The statements are synthetic, built for this exercise around a fictional "Society X, 2024." Mark the ones you would carry into management, then read what separates them.

Four answers
APer the Society X (2024) guideline, section 7.2, anticoagulation is recommended above the score threshold. Excerpt: [verbatim text from section 7.2]. Source: the official guideline document.
BAccording to the Society X (2024) guideline, as cited in a 2025 review, the benefit of anticoagulation in this profile is questionable and observation alone may be considered.
CThe Society X (2024) guideline, section 7.2, sets the threshold one point higher than the document actually states, with class and level cited normally.
DApplying the Society X (2024) guideline: use the reduced dose of the anticoagulant, which was the rule for the subgroup with renal impairment, and this patient does not belong to it.
The reveal

Only A holds up, and notice what makes it trustworthy: document declared, verbatim excerpt, section, and the nature of the source stated out loud.

B is the substitute source. A review speaking in the guideline's name, with a conclusion of its own.

C is the real citation with swapped content. The most dangerous of the four, because it survives the "does it exist" check.

D is wrong granularity. A subgroup rule applied to someone outside the subgroup, and the reverse error is just as common.

One check catches B, C, and D: opening the document at the address the answer gave you.

Why "it exists" is not verification

Checking that a reference is real catches only the crudest failure. C is real, findable, correctly formatted, and wrong about the one thing you needed.

The check that works is not bibliographic. It is reading the sentence in the document.

The topic effect

Fabrication is not evenly distributed. In simulated reviews with a frontier model, fabricated citations ran around 20% overall, rising to 28-29% on less visible topics against roughly 6% on the most familiar one.

Rarer question, higher risk. Which is exactly when you are most likely to be relying on the tool.

04 · Before it changes your management

0 of 5

Retrieval tools that cite sources, such as UpToDate, OpenEvidence, or your specialty society's own repository, reduce the version problem. The application problem stays yours, and so does the verification.

The prompt
Clinical question: [the decision this guideline will
support].
Use the guideline [name, society, year], or if I have
not specified, the most recent from [society], stating
which document and year you are using.
Tell me whether your answer comes from the guideline document
or from a source citing it. An intermediary source does not
replace the document.
My patient: [age, pregnancy status, renal function,
relevant comorbidities]. Give me the general recommendation
and any applicable subgroup rule, stating which layer this
patient falls into.
For each recommendation, include the excerpt and the section
where it appears, so I can verify at the source.
If you do not have access to the exact text, say so instead of
reconstructing from memory.
05 · Evidence

Fabrication rates fell sharply between model generations, and they did not fall to zero. What matters clinically is where the remaining errors concentrate: on less common topics, and in well-formatted citations that pass a superficial check.

  1. [study] Walters WH, Wilder EI, 2023. 55% of GPT-3.5 citations and 18% of GPT-4 citations entirely fabricated across 636 references and 42 topics, often pairing real authors with fictitious titles. journal to confirm
  2. [study] Influence of topic familiarity and prompt specificity on citation fabrication, JMIR Mental Health, 2025. Around 20% of citations entirely fabricated by a frontier model, rising to 28-29% on less visible topics against roughly 6% on the most familiar one; DOI was the field with the most errors. authors and DOI to confirm

Further reading · A 2026 reference-integrity audit across 2.5 million biomedical papers, showing well-formatted fabrications with real authors and plausible dates entering the published literature. The strongest argument for reading at the source rather than checking that a reference exists.

Educational content; the four statements in the exercise are synthetic. Not a substitute for clinical judgment or for the rules that apply where you practice.