What the study found
The study found that publicly available large language models, or LLMs, sometimes gave problematic and unsafe answers to patient medical questions. The rate of problematic responses differed across chatbots, and the authors report that some responses had the potential to cause serious patient harm.
Why the authors say this matters
The authors conclude that millions of patients may be using LLM chatbots for medical advice, so the safety of these tools matters for patient care. The study suggests that further work is needed to improve the clinical safety of these chatbots.
What the researchers tested
A physician-led red-teaming study compared four publicly available chatbots: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B. The researchers evaluated 888 responses to 222 patient-posed advice-seeking medical questions using a new dataset called HealthAdvice and an evaluation framework for quantitative and qualitative analysis.
What worked and what didn't
The study found statistically significant differences between chatbots. Problematic responses ranged from 21.6% for Claude to 43.2% for Llama, while unsafe responses ranged from 5% for Claude to 13% for GPT-4o and Llama.
What to keep in mind
The abstract does not describe detailed limitations beyond the fact that the study used a new dataset and focused on primary care topics in internal medicine, women's health, and pediatrics. The summary provided here is limited to the title and abstract only.
- Four public chatbots were tested: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B.
- The study evaluated 888 responses to 222 patient-posed medical questions.
- Problematic response rates ranged from 21.6% to 43.2% across systems.
- Unsafe response rates ranged from 5% to 13% across systems.
- The authors say some responses had the potential to lead to serious patient harm.