AI Summary of Scholarly Research

This page presents an AI-generated summary of a published research paper. The original authors did not write or review this article. [See full disclosure ↓]

LLM consensus pipeline reduced manual effort in literature screening

Research area:business-management

What the study found

The study found that a human-supervised pipeline using multiple large language models (LLMs) and a consensus scheme can reduce the manual effort needed to filter papers for systematic literature reviews. The authors report that it also achieved lower error rates than single human annotators.

Why the authors say this matters

The authors say this matters because systematic literature reviews are important for understanding a research field and guiding future research, but the literature screening step is time-consuming and labor-intensive. They conclude that responsible human-AI collaboration can accelerate and improve this workflow.

What the researchers tested

The researchers proposed a pipeline that classifies papers using descriptive prompts across multiple LLMs, then combines the model outputs with a consensus scheme. The process was human-supervised and controlled through an open-source visual analytics web interface called LLMSurver, which allowed real-time inspection and modification of outputs.

What worked and what didn't

Using ground-truth data from a recent systematic literature review with 8,323 candidate papers, the pipeline reduced manual effort and showed lower error rates than single human annotators. The abstract says that modern open-source models were sufficient for the task and that the approach was cost-effective and accessible.

What to keep in mind

The abstract does not describe detailed failure cases, specific error measurements, or how performance varied across individual models. It also does not provide limitations beyond noting that the process remains human-supervised and interactively controlled.

Key points

  • The study proposes a semi-automatic pipeline for filtering papers in systematic literature reviews.
  • Multiple LLMs classify papers, and their outputs are combined using a consensus scheme.
  • A visual analytics interface, LLMSurver, lets users inspect and modify model outputs in real time.
  • The evaluation used ground-truth data from 8,323 candidate papers from a recent systematic literature review.
  • The abstract reports lower error rates than single human annotators and lower manual effort.

Disclosure

Research title:
LLM consensus pipeline reduced manual effort in literature screening
Authors:
Lucas Joos, Daniel A. Keim, Maximilian T. Fischer
Institutions:
University of Konstanz, University of Konstanz, University of Konstanz
Publication date:
2026-02-16
OpenAlex record:
View
AI provenance: This post was generated by gpt-5.4-mini (OpenAI). The original authors did not write or review this post.