Frontier models can now exceed 90% on some medical exam-style benchmarks. But clinical safety is a different question. The question is how they fail, and who is qualified to find it before a patient does.
These are not isolated bugs to be patched one at a time. They are ten systematic failure modes that recur across models, and many conventional medical benchmarks fail to capture these behaviors. Each one takes several forms. Select any category to explore how it appears, from the patient's side and the clinician's, and why it matters.
These are the failure modes I have found most consequential in my own testing and reading. The list is not exhaustive, and it grows. As I document new patterns, they will join it here.
Naming a risk is not the same as knowing how often it occurs or how severe it is. The work is turning what is plausible into what is measured, through a repeatable process.
Each scenario is built deliberately to probe a single failure mode. Knowing what an invented detail would have looked like had it been real requires clinical judgment, which is precisely what a generic benchmark cannot supply.
The model's behavior is documented directly before any conclusion is drawn. Nothing here is assumed or inferred. This is the purpose of independent field testing.
The clinically plausible actions for a case are enumerated in advance, and each is judged for appropriateness and for the harm of omitting or committing it. This is what makes an evaluation systematic rather than anecdotal.
Harm is assessed as immediate, short term, and long term, rather than compressed into a single score. A failure that appears minor in the moment can carry serious consequences later, and the reverse is also true.
Evaluation considers the entire exchange, not one message in isolation. The failures that emerge only across a conversation are invisible to any test that grades a single turn.
The patterns described here are not personal opinion. A body of peer reviewed work is beginning to measure how medical AI fails, and to argue that safety must be evaluated on its own terms. A selection of the studies that inform this work.
One of the first large non-creator red teaming efforts in healthcare. Clinicians and technical reviewers stress tested models with real clinical cases and categorized failures along safety, privacy, hallucination and accuracy, and bias, the same axes used throughout this site.
20.1% of prompts produced inappropriate responsesThe clearest evidence that accuracy and safety are different measurements. Applying frontier model recommendations directly carried potential for severe harm in a meaningful share of cases, and most severe errors were omissions rather than commissions.
Up to 24.6% of cases carried potential for severe harm; over 80% of severe errors were omissionsMakes the case at the heart of this work: a model can reach a correct final answer through unsafe reasoning, and grading only the final response hides it. Argues for examining the intermediate steps, especially in ethically charged clinical scenarios.
Directly supports evaluating the path, not only the answerIntroduces a safety protocol that evaluates the same model from three points of view, patient, clinician, and general user. Reinforces that a failure can look entirely different depending on who is asking and how.
First protocol to define safety across patient, clinician, and general-user perspectivesMoves medical AI evaluation beyond licensing-exam questions toward the diversity of real clinical work, including clinical decision support, documentation, patient communication, medical research, and administrative workflows. Shows that strong benchmark performance does not translate uniformly across the tasks models may actually encounter in healthcare.
121 clinical tasks across 22 subcategories; performance varied substantially by taskProposes a formal, repeatable red teaming framework for clinical decision support, covering adversarial prompts, hallucination measurement, and structured scoring. A published counterpart to the method described here.
A structured, four-phase red teaming framework for clinical AIThis site reflects my independent work in clinical AI evaluation. Cases come from real field tests. Model responses are reproduced verbatim. Identifying information and visual formatting may be modified. External studies and resources are cited for educational purposes, and I hold no financial ties to the works, authors, or companies referenced.

I am Nathalia Müller, a physician and clinical researcher who moved from practice into the backstage of medical AI. I started working as an annotator and medical expert, and today, much of my work is in healthcare AI red teaming: designing the cases that reveal where models fail, and judging how much each failure genuinely matters. It is a deliberately different place to stand. Not the clinician using an AI tool, but the clinician testing it from the inside, before it reaches anyone.
Part of that work is confidential. This site gathers what I can show openly, drawn from independent field tests and the published literature. What you see here is a condensed slice of a larger framework I work from.
New cases, patterns, and discussions land first on LinkedIn, where I publish field tests as I document them. If you want to follow the work as it develops, or join the conversation around clinical AI safety, that is the place.