Leading AI systems produced unsupported medical diagnoses in nearly one in five tests where the required scan had been intentionally omitted, according to a Carnegie Mellon University study that raises fresh concerns about patients relying on chatbots for medical advice.
The research examined Claude Opus 4.7, OpenAI’s GPT-5.4 and Google’s Gemini 3.1 Pro across 11,700 prompts involving chest X-rays, brain MRI scans and images of skin moles.
Although each prompt referred to an image, none was actually provided, allowing researchers to test whether the AI models would acknowledge the missing evidence or generate an answer regardless.
Across the experiment, the models correctly declined to offer a diagnosis in about 82% of responses, but supplied a disease in the remaining 18%, despite having no medical image to examine.
Demographics and word choice changed the answers
Siddharth Vohra, a master’s student at Carnegie Mellon’s Robotics Institute and the study’s sole author, tested 12 simulated patient profiles using different combinations of age, sex and race, alongside a control prompt containing no demographic details.
Rather than failing consistently, the three systems showed different patterns of behavior. GPT-5.4 generated diagnoses across all 36 patient-and-scan combinations, while Claude refused more frequently but repeatedly produced the same diseases in certain demographic scenarios. Gemini declined most often, although the diagnoses it did provide still changed when patient characteristics were altered.
In one test, Claude identified melanoma in 94% of responses involving a 65-year-old white man asking about a missing image of a skin mole. GPT-5.4, meanwhile, named sarcoidosis in 77 out of 100 chest X-ray prompts involving a young Black patient, even though no X-ray was attached.
The findings were also highly sensitive to small changes in wording. When researchers replaced “mole” with “lesion” for the same patient profile, Claude shifted from identifying melanoma in 94% of responses to refusing the task every time, while GPT-5.4’s behavior changed little.
Refusal language did not always stop a diagnosis
The study also found that some models appeared to refuse a diagnosis in their written explanation while still placing a disease name inside a structured field that could be read by hospital software or another automated system.
In Claude’s weakest-performing test, 62 of 94 invented diagnoses were accompanied by language resembling a refusal, creating a risk that a human reader might believe the model had safely declined while an unsupported diagnosis continued through the technical workflow.
Study highlights limits of fluent medical answers
The study relied on simulated patient profiles rather than real clinical cases. However, its findings still exposed a key medical AI risk: fluent, authoritative answers can appear clinically sound even when the model lacks the evidence needed to support them.
Vohra said healthcare providers should test AI systems for demographic sensitivity and examine how they behave when patient information is incomplete before introducing them into clinical settings.
Suggested safeguards include forcing diagnostic fields to remain blank when an image is missing and comparing model responses generated with and without the underlying scan.




