The right question is never "what does the guideline say?" It is: what does guideline X, version Y, say for population Z? An honest calibration first: today’s tools, with live search, are far better than the era of invented citations. The risk did not disappear. It moved.
You need the current recommendation, and you need it before the patient finishes getting dressed. What comes back looks exactly like a guideline: class of recommendation, level of evidence, section number. All of it can be reconstructed.
Guiding question: Which document, which version, which population?
Finds the right recommendation in minutes, translates it into decision language, compares societies when you ask, and points to exactly where each statement lives. With live search and sources required, it is an excellent librarian. What it hands you is an address, not a reading.
Decide which guidance governs your patient, or notice that your patient belongs to a subgroup with a rule of its own.
This card does not pretend it is 2023. Fabrication rates fell sharply between model generations, and the failure that replaced them is subtler and harder to catch: substituting the document and blending populations. Retrieval that ends in the chat answer is not retrieval. It ends in the PDF.
Answering from memory, the model returns an average of everything it read: old versions fused with new, societies mixed, an adult recommendation leaking into pregnancy. It sounds like a guideline and is none of them, with class, level, and section number attached. Live search reduces this. It does not eliminate it.
What comes back is often not the guideline but something about it: an article that cites it and concludes differently, a review that qualifies it, a site that summarizes it. The answer inherits the document's authority with the intermediary's content.
Guidelines carry general recommendations and subgroup rules, and fast reading errs both ways: applying a subgroup threshold to a patient outside it, or applying the general rule to a patient who had a rule of her own by age, pregnancy, or renal function.
These three link to the site's failure taxonomy, where the transcripts live.
A real citation with swapped content: the document exists, the section exists, and the recommendation attributed to it is wrong. It survives the lazy check of "it exists, so it's fine," which is why the verification on this card is reading the sentence rather than confirming the reference.
Four guideline-shaped answers about anticoagulation in atrial fibrillation. The statements are synthetic, built for this exercise around a fictional "Society X, 2024." Mark the ones you would carry into management, then read what separates them.
Only A holds up, and notice what makes it trustworthy: document declared, verbatim excerpt, section, and the nature of the source stated out loud.
B is the substitute source. A review speaking in the guideline's name, with a conclusion of its own.
C is the real citation with swapped content. The most dangerous of the four, because it survives the "does it exist" check.
D is wrong granularity. A subgroup rule applied to someone outside the subgroup, and the reverse error is just as common.
One check catches B, C, and D: opening the document at the address the answer gave you.
Checking that a reference is real catches only the crudest failure. C is real, findable, correctly formatted, and wrong about the one thing you needed.
The check that works is not bibliographic. It is reading the sentence in the document.
Fabrication is not evenly distributed. In simulated reviews with a frontier model, fabricated citations ran around 20% overall, rising to 28-29% on less visible topics against roughly 6% on the most familiar one.
Rarer question, higher risk. Which is exactly when you are most likely to be relying on the tool.
0 of 5
Retrieval tools that cite sources, such as UpToDate, OpenEvidence, or your specialty society's own repository, reduce the version problem. The application problem stays yours, and so does the verification.
Clinical question: [the decision this guideline will support]. Use the guideline [name, society, year], or if I have not specified, the most recent from [society], stating which document and year you are using. Tell me whether your answer comes from the guideline document or from a source citing it. An intermediary source does not replace the document. My patient: [age, pregnancy status, renal function, relevant comorbidities]. Give me the general recommendation and any applicable subgroup rule, stating which layer this patient falls into. For each recommendation, include the excerpt and the section where it appears, so I can verify at the source. If you do not have access to the exact text, say so instead of reconstructing from memory.
Fabrication rates fell sharply between model generations, and they did not fall to zero. What matters clinically is where the remaining errors concentrate: on less common topics, and in well-formatted citations that pass a superficial check.
Further reading · A 2026 reference-integrity audit across 2.5 million biomedical papers, showing well-formatted fabrications with real authors and plausible dates entering the published literature. The strongest argument for reading at the source rather than checking that a reference exists.