What the study found
GPT-o1 performed better than Llama-3.2-8b-instruct on clinically grounded laboratory test scenarios that examined causal reasoning. The models were tested across association, intervention, and counterfactual reasoning, and GPT-o1 had higher scores in each area.
Why the authors say this matters
The authors conclude that GPT-o1 offers more consistent causal reasoning. They also state that further refinement is needed before high-stakes clinical deployment.
What the researchers tested
The researchers evaluated 99 laboratory test scenarios mapped to Pearl's Ladder of Causation, a framework with three levels: association, intervention, and counterfactual reasoning. The scenarios focused on common tests such as hemoglobin A1c (HbA1c), creatinine, and vitamin D, paired with factors like age, gender, obesity, and smoking. Two large language models, GPT-o1 and Llama-3.2-8b-instruct, were rated by four medically trained human experts.
What worked and what didn't
GPT-o1 had higher overall discriminative performance, with AUROC (area under the receiver operating characteristic curve, a measure of classification performance) of 0.80 ± 0.12 versus 0.73 ± 0.15 for Llama-3.2-8b-instruct. It also scored higher on association, intervention, and counterfactual reasoning, and had higher sensitivity and specificity. Both models performed best on intervention questions and worst on counterfactual questions, especially altered outcome scenarios.
What to keep in mind
The abstract does not describe limitations beyond noting that further refinement is needed before high-stakes clinical deployment. The study summary provided here is limited to the tested scenarios, the two models, and the expert ratings reported in the abstract.
Key points
- The study compared GPT-o1 and Llama-3.2-8b-instruct on 99 clinical laboratory test scenarios.
- GPT-o1 scored higher than Llama-3.2-8b-instruct on overall AUROC, sensitivity, and specificity.
- Both models did best on intervention questions and worst on counterfactual questions.
- The scenarios used common lab tests including HbA1c, creatinine, and vitamin D.
- The authors say more refinement is needed before high-stakes clinical deployment.
Disclosure
- Research title:
- GPT-o1 outperformed Llama-3.2-8b-instruct in causal reasoning
- Authors:
- Balu Bhasuran, Mattia Prosperi, Karim Hanna, John Petrilli, Caretia JeLayne Washington, Zhe He
- Institutions:
- Florida A&M University – Florida State University College of Engineering, Florida State University, Florida State University, University of Florida, University of Florida, University of South Florida, University of South Florida
- Publication date:
- 2026-04-23
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.