Tag: Language & Text Computing

  • Author labels bias LLM evaluation scores

    What the study found

    The study found that large language model (LLM) judgments of text quality were strongly influenced by which model was said to have written the text. Posts labeled as "Claude" were rated higher, while posts labeled as "Gemini" were rated lower, even when the content was the same.

    Why the authors say this matters

    The authors conclude that LLM-based assessment may not reliably separate content quality from author labels. They say this raises concerns for benchmarking, content moderation, and automated review pipelines, and they suggest blind assessment, multi-model consensus scoring, and statistical safeguards to detect label-induced bias.

    What the researchers tested

    The researchers generated blog posts with three LLMs: Chat-GPT, Gemini, and Claude. Each model then evaluated every post under three conditions: with no author label, with the correct author label, and with intentionally incorrect author labels.

    What worked and what didn't

    The results showed substantial bias tied to perceived authorship rather than actual content quality. False labels sometimes changed absolute scores and even reversed preference rankings, with shifts reported as large as 50 percentage points; the effects also appeared across coherence, informativeness, and conciseness.

    What to keep in mind

    The abstract does not describe limitations beyond the scope of the tested setup. The reported findings come from blog posts evaluated by three specific LLMs under the conditions described.

    • LLM evaluations changed when the same text was paired with different author labels.
    • "Claude" labels were consistently scored more favorably, while "Gemini" labels were downgraded.
    • Incorrect labels sometimes reversed preference rankings and shifted scores by as much as 50 percentage points.
    • Bias appeared in overall preferences and in quality dimensions such as coherence, informativeness, and conciseness.
    • The authors say blind assessment and other safeguards may help detect label-induced bias.
  • WorldView-Bench measures cultural bias in LLMs

    What the study found

    The study found that WorldView-Bench, a benchmark for Global Cultural Inclusivity in LLMs, can measure cultural bias through free-form generative evaluation. The reported results show higher perspective diversity and a shift toward more positive sentiment when multiplex-aware approaches were used.

    Why the authors say this matters

    The authors conclude that cultural bias in LLMs can be meaningfully measured and mitigated through structured worldview diversity. They suggest this may support more inclusive, globally representative, and ethically aligned AI systems.

    What the researchers tested

    The researchers introduced WorldView-Bench, which is grounded in the Multiplex Worldview framework. They compared a baseline with two intervention strategies: Contextually-Implemented Multiplex LLMs, which use system prompts to embed multiplexity principles, and Multi-Agent System-Implemented Multiplex LLMs, where multiple LLM agents representing distinct cultural perspectives generate responses together.

    What worked and what didn't

    The abstract reports a rise in Perspectives Distribution Score entropy from 13% at baseline to 94% with Multi-Agent System-Implemented Multiplex LLMs. It also reports a shift toward positive sentiment, at 67.7%, and enhanced cultural balance. The abstract does not give a detailed comparison of which intervention worked better beyond these reported results.

    What to keep in mind

    The abstract provides summary results only, without detailed experimental settings, dataset information, or broader validation details. It also does not describe any limitations beyond the general scope of the benchmark and interventions.

    • WorldView-Bench is presented as a benchmark for evaluating Global Cultural Inclusivity in LLMs.
    • The benchmark uses free-form generative evaluation rather than closed-form categorical testing.
    • The paper reports a rise in Perspectives Distribution Score entropy from 13% at baseline to 94% with Multi-Agent System-Implemented Multiplex LLMs.
    • The reported results also show 67.7% positive sentiment and improved cultural balance.
    • The authors say cultural bias in LLMs can be measured and mitigated through structured worldview diversity.
  • People’s Daily and CCTV News used different sentiment strategies on Douyin

    What the study found

    The study found that two Chinese state media outlets used different sentiment patterns when presenting rural revitalization on Douyin, a short-video platform. People’s Daily was dominated by positive emotions across topics, while CCTV News used a more varied emotional design.

    Why the authors say this matters

    The authors suggest these findings show how institutional identity shapes digital storytelling strategies. They conclude that the party newspaper, People’s Daily, emphasizes ideological reinforcement, while the State Television Station, CCTV, balances political content with emotional resonance.

    What the researchers tested

    The researchers examined 445 rural revitalization videos posted by People’s Daily and CCTV News on Douyin. They used a computational approach combining Latent Dirichlet Allocation, a topic modeling method for finding themes, and StructBERT sentiment analysis, which estimates emotional tone from text.

    What worked and what didn't

    The analysis identified different communication styles in the two outlets. People’s Daily was dominated by positive sentiment across all topics, while CCTV News showed a more differentiated sentiment pattern, especially for poverty alleviation and rural livelihood issues.

    What to keep in mind

    The summary only describes two media outlets, one policy topic, and one platform, so the findings are limited to that scope. The abstract does not describe additional limitations beyond the study design and sample used.

    • The study analyzed 445 rural revitalization videos on Douyin.
    • People’s Daily showed positive sentiment across all topics.
    • CCTV News used a more differentiated emotional approach.
    • The results varied especially for poverty alleviation and rural livelihood issues.
    • The authors link the pattern to institutional identity and digital storytelling strategy.
  • UN Security Council transcript dataset spans 1946 to 2024

    What the study found

    The article presents a new machine-readable dataset of public United Nations Security Council transcripts from 1946 to 2024. It includes more than 160,000 speeches and over 87 million words, with speaker identity, affiliation, and speaking order preserved.

    Why the authors say this matters

    The authors say the dataset offers unprecedented historical depth and can support research on global security norms, institutional discourse, and the relationship between language and international policy. The study suggests it can help analyze how different actors express security concepts over time.

    What the researchers tested

    The researchers built a dataset from every available public transcript of the Security Council and organized it in a machine-readable format. They demonstrated its analytical potential with three illustrative applications using traditional text analysis and transformer-based text analysis, a type of machine learning used to process language.

    What worked and what didn't

    The dataset appears to support detailed analysis across nearly eight decades of deliberations, including Cold War and post-Cold War periods. The examples shown examine the evolution of sovereignty from right to responsibility, the transformation of human rights discourse after the Cold War, and the identification of institutional champions of the humanitarian turn.

    What to keep in mind

    The abstract describes illustrative applications, so the results presented are examples of the dataset's analytical potential rather than a full evaluation of every possible use. Limitations are not described in the available summary.

    • The article introduces a machine-readable dataset of UN Security Council public transcripts from 1946 to 2024.
    • The dataset contains more than 160,000 speeches and over 87 million words.
    • It preserves speaker identity, affiliation, and exact speaking order.
    • The authors demonstrate three illustrative text-analysis applications involving sovereignty, human rights, and humanitarian discourse.
    • The study says the resource can support research on global security norms and institutional discourse.