What the study found
The study found that a human-supervised pipeline using multiple large language models (LLMs) and a consensus scheme can reduce the manual effort needed to filter papers for systematic literature reviews. The authors report that it also achieved lower error rates than single human annotators.
Why the authors say this matters
The authors say this matters because systematic literature reviews are important for understanding a research field and guiding future research, but the literature screening step is time-consuming and labor-intensive. They conclude that responsible human-AI collaboration can accelerate and improve this workflow.
What the researchers tested
The researchers proposed a pipeline that classifies papers using descriptive prompts across multiple LLMs, then combines the model outputs with a consensus scheme. The process was human-supervised and controlled through an open-source visual analytics web interface called LLMSurver, which allowed real-time inspection and modification of outputs.
What worked and what didn't
Using ground-truth data from a recent systematic literature review with 8,323 candidate papers, the pipeline reduced manual effort and showed lower error rates than single human annotators. The abstract says that modern open-source models were sufficient for the task and that the approach was cost-effective and accessible.
What to keep in mind
The abstract does not describe detailed failure cases, specific error measurements, or how performance varied across individual models. It also does not provide limitations beyond noting that the process remains human-supervised and interactively controlled.
Key points
- The study proposes a semi-automatic pipeline for filtering papers in systematic literature reviews.
- Multiple LLMs classify papers, and their outputs are combined using a consensus scheme.
- A visual analytics interface, LLMSurver, lets users inspect and modify model outputs in real time.
- The evaluation used ground-truth data from 8,323 candidate papers from a recent systematic literature review.
- The abstract reports lower error rates than single human annotators and lower manual effort.
Disclosure
- Research title:
- LLM consensus pipeline reduced manual effort in literature screening
- Authors:
- Lucas Joos, Daniel A. Keim, Maximilian T. Fischer
- Institutions:
- University of Konstanz, University of Konstanz, University of Konstanz
- Publication date:
- 2026-02-16
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.