AI Summary of Scholarly Research

This page presents an AI-generated summary of a published research paper. The original authors did not write or review this article. [See full disclosure ↓]

Public chatbots gave unsafe medical advice in many responses

Research area:health-policy-serviceshealth-services-mgmt

What the study found

The study found that publicly available large language models, or LLMs, sometimes gave problematic and unsafe answers to patient medical questions. The rate of problematic responses differed across chatbots, and the authors report that some responses had the potential to cause serious patient harm.

Why the authors say this matters

The authors conclude that millions of patients may be using LLM chatbots for medical advice, so the safety of these tools matters for patient care. The study suggests that further work is needed to improve the clinical safety of these chatbots.

What the researchers tested

A physician-led red-teaming study compared four publicly available chatbots: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B. The researchers evaluated 888 responses to 222 patient-posed advice-seeking medical questions using a new dataset called HealthAdvice and an evaluation framework for quantitative and qualitative analysis.

What worked and what didn't

The study found statistically significant differences between chatbots. Problematic responses ranged from 21.6% for Claude to 43.2% for Llama, while unsafe responses ranged from 5% for Claude to 13% for GPT-4o and Llama.

What to keep in mind

The abstract does not describe detailed limitations beyond the fact that the study used a new dataset and focused on primary care topics in internal medicine, women's health, and pediatrics. The summary provided here is limited to the title and abstract only.

Key points

  • Four public chatbots were tested: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B.
  • The study evaluated 888 responses to 222 patient-posed medical questions.
  • Problematic response rates ranged from 21.6% to 43.2% across systems.
  • Unsafe response rates ranged from 5% to 13% across systems.
  • The authors say some responses had the potential to lead to serious patient harm.

Disclosure

Research title:
Public chatbots gave unsafe medical advice in many responses
Authors:
Rachel Lea Draelos, Samina Afreen, Barbara Blasko, Tiffany L. Brazile, Natasha Chase, Dimple Patel Desai, Jessica Evert, Heather Gardner, Lauren Herrmann, Aswathy Vaikom House, Stephanie Kass, Marianne Kavan, Kirshma Khemani, Amanda Koire, Lauren M. McDonald, Zahraa Rabeeah, Amy Shah
Institutions:
Brigham and Women's Hospital, Brownsville Public Library, California Wellness Foundation, Cooper University Hospital, Emory University, Inova Fairfax Hospital, Institute of Glass, Northwell Health, Piedmont HealthCare, Riverside Community Hospital, University of California, San Francisco, University of Louisville, University of Oklahoma Health Sciences Center, University of Teacher Education Zug, University of Virginia
Publication date:
2026-02-13
OpenAlex record:
View
AI provenance: This post was generated by gpt-5.4-mini (OpenAI). The original authors did not write or review this post.