Clinical evaluation of medical AI

A model can get the diagnosis right
and still not be safe.

Frontier models can now exceed 90% on some medical exam-style benchmarks. But clinical safety is a different question. The question is how they fail, and who is qualified to find it before a patient does.

24.6%
of cases carried potential for severe harm when frontier model recommendations were applied directly
Source: NOHARM benchmark
Over 80%
of those severe errors were omissions, the appropriate step quietly left out rather than a wrong one added
Source: NOHARM benchmark
Clinical expertise
changes what gets tested, what counts as a failure, and how its potential harm is judged
Recurring finding across studies
The taxonomy

The systematic ways medical AI fails

These are not isolated bugs to be patched one at a time. They are ten systematic failure modes that recur across models, and many conventional medical benchmarks fail to capture these behaviors. Each one takes several forms. Select any category to explore how it appears, from the patient's side and the clinician's, and why it matters.

These are the failure modes I have found most consequential in my own testing and reading. The list is not exhaustive, and it grows. As I document new patterns, they will join it here.

The method

Identifying these failures is a discipline, not a hunch

Naming a risk is not the same as knowing how often it occurs or how severe it is. The work is turning what is plausible into what is measured, through a repeatable process.

01

Designing the case

Each scenario is built deliberately to probe a single failure mode. Knowing what an invented detail would have looked like had it been real requires clinical judgment, which is precisely what a generic benchmark cannot supply.

02

Testing it firsthand

The model's behavior is documented directly before any conclusion is drawn. Nothing here is assumed or inferred. This is the purpose of independent field testing.

03

Scoring against a rubric

The clinically plausible actions for a case are enumerated in advance, and each is judged for appropriateness and for the harm of omitting or committing it. This is what makes an evaluation systematic rather than anecdotal.

04

Weighing severity across time

Harm is assessed as immediate, short term, and long term, rather than compressed into a single score. A failure that appears minor in the moment can carry serious consequences later, and the reverse is also true.

05

Judging the full trajectory

Evaluation considers the entire exchange, not one message in isolation. The failures that emerge only across a conversation are invisible to any test that grades a single turn.

The evidence

This is a field with a growing literature

The patterns described here are not personal opinion. A body of peer reviewed work is beginning to measure how medical AI fails, and to argue that safety must be evaluated on its own terms. A selection of the studies that inform this work.

npj Digital Medicine
Chang et al., 2025

Red teaming ChatGPT in medicine to yield real-world insights on model behavior

Stanford, 80 participants across clinical and technical roles

One of the first large non-creator red teaming efforts in healthcare. Clinicians and technical reviewers stress tested models with real clinical cases and categorized failures along safety, privacy, hallucination and accuracy, and bias, the same axes used throughout this site.

20.1% of prompts produced inappropriate responses
Benchmark study
Wu, Nateghi Haredasht et al.

First, do NO HARM: a medical safety benchmark and randomized study of physician and AI teaming

Stanford and Harvard, 1,100 cases, 29 specialist annotators

The clearest evidence that accuracy and safety are different measurements. Applying frontier model recommendations directly carried potential for severe harm in a meaningful share of cases, and most severe errors were omissions rather than commissions.

Up to 24.6% of cases carried potential for severe harm; over 80% of severe errors were omissions
npj Digital Medicine
Sorin et al., 2025

Reasoning red teaming in healthcare: not all paths to a desired outcome are desirable

Matters arising

Makes the case at the heart of this work: a model can reach a correct final answer through unsafe reasoning, and grading only the final response hides it. Argues for examining the intermediate steps, especially in ethically charged clinical scenarios.

Directly supports evaluating the path, not only the answer
Microsoft Research
Corbeil et al., 2025

MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings

PatientSafetyBench, 466 samples across 5 critical categories

Introduces a safety protocol that evaluates the same model from three points of view, patient, clinician, and general user. Reinforces that a failure can look entirely different depending on who is asking and how.

First protocol to define safety across patient, clinician, and general-user perspectives
Nature Medicine
Bedi et al., 2026

Holistic evaluation of large language models for medical tasks with MedHELM

Clinician-validated taxonomy, 121 medical tasks across 5 categories

Moves medical AI evaluation beyond licensing-exam questions toward the diversity of real clinical work, including clinical decision support, documentation, patient communication, medical research, and administrative workflows. Shows that strong benchmark performance does not translate uniformly across the tasks models may actually encounter in healthcare.

121 clinical tasks across 22 subcategories; performance varied substantially by task
Bioengineering
Trabilsy et al., 2025

The PIEE Cycle: a structured framework for red teaming LLMs in clinical decision-making

Planning, Information gathering, Execution, Evaluation

Proposes a formal, repeatable red teaming framework for clinical decision support, covering adversarial prompts, hallucination measurement, and structured scoring. A published counterpart to the method described here.

A structured, four-phase red teaming framework for clinical AI

This site reflects my independent work in clinical AI evaluation. Cases come from real field tests. Model responses are reproduced verbatim. Identifying information and visual formatting may be modified. External studies and resources are cited for educational purposes, and I hold no financial ties to the works, authors, or companies referenced.

About

A physician working the backstage of medical AI

Nathalia Müller, MD

I am Nathalia Müller, a physician and clinical researcher who moved from practice into the backstage of medical AI. I started working as an annotator and medical expert, and today, much of my work is in healthcare AI red teaming: designing the cases that reveal where models fail, and judging how much each failure genuinely matters. It is a deliberately different place to stand. Not the clinician using an AI tool, but the clinician testing it from the inside, before it reaches anyone.

Part of that work is confidential. This site gathers what I can show openly, drawn from independent field tests and the published literature. What you see here is a condensed slice of a larger framework I work from.

New cases, patterns, and discussions land first on LinkedIn, where I publish field tests as I document them. If you want to follow the work as it develops, or join the conversation around clinical AI safety, that is the place.