What the study found
The study found that publicly available large language models, or LLMs, sometimes gave problematic and unsafe answers to patient medical questions. The rate of problematic responses differed across chatbots, and the authors report that some responses had the potential to cause serious patient harm.
Why the authors say this matters
The authors conclude that millions of patients may be using LLM chatbots for medical advice, so the safety of these tools matters for patient care. The study suggests that further work is needed to improve the clinical safety of these chatbots.
What the researchers tested
A physician-led red-teaming study compared four publicly available chatbots: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B. The researchers evaluated 888 responses to 222 patient-posed advice-seeking medical questions using a new dataset called HealthAdvice and an evaluation framework for quantitative and qualitative analysis.
What worked and what didn't
The study found statistically significant differences between chatbots. Problematic responses ranged from 21.6% for Claude to 43.2% for Llama, while unsafe responses ranged from 5% for Claude to 13% for GPT-4o and Llama.
What to keep in mind
The abstract does not describe detailed limitations beyond the fact that the study used a new dataset and focused on primary care topics in internal medicine, women's health, and pediatrics. The summary provided here is limited to the title and abstract only.
Key points
- Four public chatbots were tested: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B.
- The study evaluated 888 responses to 222 patient-posed medical questions.
- Problematic response rates ranged from 21.6% to 43.2% across systems.
- Unsafe response rates ranged from 5% to 13% across systems.
- The authors say some responses had the potential to lead to serious patient harm.
Disclosure
- Research title:
- Public chatbots gave unsafe medical advice in many responses
- Authors:
- Rachel Lea Draelos, Samina Afreen, Barbara Blasko, Tiffany L. Brazile, Natasha Chase, Dimple Patel Desai, Jessica Evert, Heather Gardner, Lauren Herrmann, Aswathy Vaikom House, Stephanie Kass, Marianne Kavan, Kirshma Khemani, Amanda Koire, Lauren M. McDonald, Zahraa Rabeeah, Amy Shah
- Institutions:
- Brigham and Women's Hospital, Brownsville Public Library, California Wellness Foundation, Cooper University Hospital, Emory University, Inova Fairfax Hospital, Institute of Glass, Northwell Health, Piedmont HealthCare, Riverside Community Hospital, University of California, San Francisco, University of Louisville, University of Oklahoma Health Sciences Center, University of Teacher Education Zug, University of Virginia
- Publication date:
- 2026-02-13
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.