AI Summary of Scholarly Research

This page presents an AI-generated summary of a published research paper. The original authors did not write or review this article. [See full disclosure ↓]

Author labels bias LLM evaluation scores

Research area:computer-science-aitext-data-analysis

What the study found

The study found that large language model (LLM) judgments of text quality were strongly influenced by which model was said to have written the text. Posts labeled as "Claude" were rated higher, while posts labeled as "Gemini" were rated lower, even when the content was the same.

Why the authors say this matters

The authors conclude that LLM-based assessment may not reliably separate content quality from author labels. They say this raises concerns for benchmarking, content moderation, and automated review pipelines, and they suggest blind assessment, multi-model consensus scoring, and statistical safeguards to detect label-induced bias.

What the researchers tested

The researchers generated blog posts with three LLMs: Chat-GPT, Gemini, and Claude. Each model then evaluated every post under three conditions: with no author label, with the correct author label, and with intentionally incorrect author labels.

What worked and what didn't

The results showed substantial bias tied to perceived authorship rather than actual content quality. False labels sometimes changed absolute scores and even reversed preference rankings, with shifts reported as large as 50 percentage points; the effects also appeared across coherence, informativeness, and conciseness.

What to keep in mind

The abstract does not describe limitations beyond the scope of the tested setup. The reported findings come from blog posts evaluated by three specific LLMs under the conditions described.

Key points

  • LLM evaluations changed when the same text was paired with different author labels.
  • "Claude" labels were consistently scored more favorably, while "Gemini" labels were downgraded.
  • Incorrect labels sometimes reversed preference rankings and shifted scores by as much as 50 percentage points.
  • Bias appeared in overall preferences and in quality dimensions such as coherence, informativeness, and conciseness.
  • The authors say blind assessment and other safeguards may help detect label-induced bias.

Disclosure

Research title:
Author labels bias LLM evaluation scores
Authors:
Muskan Saraf, Sajjad Rezvani Boroujeni, Justin Beaudry, Hossein Abedi, Tom Bush
Institutions:
Virtual Reality Medical Center, Virtual Reality Medical Center, Virtual Reality Medical Center, Virtual Reality Medical Center, Virtual Reality Medical Center
Publication date:
2026-06-30
OpenAlex record:
View
AI provenance: This post was generated by gpt-5.4-mini (OpenAI). The original authors did not write or review this post.