What the study found
The study found that reasoning-based large language models could predict 12-week remission in patients with depressive disorder receiving antidepressant monotherapy. The best-performing model was Claude 3.7 Sonnet with 32,000 reasoning tokens and a referencing of deep research prompt.
Why the authors say this matters
The authors conclude that these models show promise as interpretable adjunctive tools in depressive disorder treatment planning. They also say prospective validation in real-world clinical settings remains essential.
What the researchers tested
The researchers analyzed data from 390 patients in the MAKE Biomarker discovery study who were taking first-step antidepressant monotherapy. They tested three large language models — ChatGPT o1, o3-mini, and Claude 3.7 Sonnet — using prompting strategies including zero-shot chain-of-thought, atom-of-thoughts, and a novel referencing of deep research prompt. Three psychiatrists independently rated the model outputs for clinical validity on 5-point Likert scales.
What worked and what didn't
Claude 3.7 Sonnet with 32,000 reasoning tokens and the referencing of deep research prompt achieved the highest performance, with balanced accuracy of 0.6697, sensitivity of 0.7183, and specificity of 0.6210. Medication-specific analysis showed negative predictive values of 0.75 or higher across major antidepressants, suggesting stronger performance for identifying likely nonresponders. Psychiatrists gave favorable mean ratings for correctness, consistency, specificity, helpfulness, and human likeness.
What to keep in mind
The study used retrospective data from a single biomarker discovery dataset after excluding patients with uncommon medications or missing biomarker data. The abstract does not describe longer-term follow-up beyond 12 weeks, and it states that prospective real-world validation is still needed.
Key points
- The best model was Claude 3.7 Sonnet with 32,000 reasoning tokens and a referencing of deep research prompt.
- Balanced accuracy for the top model was 0.6697, with sensitivity of 0.7183 and specificity of 0.6210.
- Negative predictive values were 0.75 or higher across major antidepressants in medication-specific analysis.
- Three psychiatrists rated the model outputs favorably on correctness, consistency, specificity, helpfulness, and human likeness.
- The authors say prospective validation in real-world clinical settings remains essential.
Disclosure
- Research title:
- Reasoning-based LLMs predicted 12-week antidepressant remission
- Authors:
- Jin-Hyun Park, Hee-Ju Kang, Ji Hyeon Jeon, S W Kang, Ju-Wan Kim, Jae-Min Kim, Hwamin Lee
- Institutions:
- Chonnam National University, Chonnam National University, Chonnam National University, Chonnam National University, Chonnam National University, Korea University, Korea University
- Publication date:
- 2026-01-23
- DOI:
- 10.2196/83352
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.