Author: editor@focalinterest.com

  • Public chatbots gave unsafe medical advice in many responses

    What the study found

    The study found that publicly available large language models, or LLMs, sometimes gave problematic and unsafe answers to patient medical questions. The rate of problematic responses differed across chatbots, and the authors report that some responses had the potential to cause serious patient harm.

    Why the authors say this matters

    The authors conclude that millions of patients may be using LLM chatbots for medical advice, so the safety of these tools matters for patient care. The study suggests that further work is needed to improve the clinical safety of these chatbots.

    What the researchers tested

    A physician-led red-teaming study compared four publicly available chatbots: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B. The researchers evaluated 888 responses to 222 patient-posed advice-seeking medical questions using a new dataset called HealthAdvice and an evaluation framework for quantitative and qualitative analysis.

    What worked and what didn't

    The study found statistically significant differences between chatbots. Problematic responses ranged from 21.6% for Claude to 43.2% for Llama, while unsafe responses ranged from 5% for Claude to 13% for GPT-4o and Llama.

    What to keep in mind

    The abstract does not describe detailed limitations beyond the fact that the study used a new dataset and focused on primary care topics in internal medicine, women's health, and pediatrics. The summary provided here is limited to the title and abstract only.

    • Four public chatbots were tested: Claude, Gemini, GPT-4o, and Llama-3.0/3.1-70B.
    • The study evaluated 888 responses to 222 patient-posed medical questions.
    • Problematic response rates ranged from 21.6% to 43.2% across systems.
    • Unsafe response rates ranged from 5% to 13% across systems.
    • The authors say some responses had the potential to lead to serious patient harm.
  • Humanity’s Last Exam benchmarks expert-level AI performance

    What the study found

    The study introduces Humanity’s Last Exam, a multi-modal benchmark made to test expert-level closed-ended academic questions across many subjects. The authors report that current large language models show low accuracy and calibration on this benchmark, which suggests a gap between model performance and the expert human frontier.

    Why the authors say this matters

    The authors say benchmarks are important for tracking rapid progress in large language models, but existing ones are no longer difficult enough to measure top systems well. The study suggests Humanity’s Last Exam could help researchers and policymakers understand model capabilities more clearly.

    What the researchers tested

    The researchers created a benchmark called Humanity’s Last Exam with 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. It includes multiple-choice and short-answer questions that can be automatically graded, and each question has a known, unambiguous solution that cannot be quickly answered by internet retrieval.

    What worked and what didn't

    The benchmark was designed to be suitable for automated grading and broad subject coverage. State-of-the-art large language models performed poorly on it, with low accuracy and calibration reported by the authors.

    What to keep in mind

    The abstract does not provide detailed numerical results beyond stating that accuracy and calibration were low. It also does not describe specific limitations of the benchmark itself beyond noting that it is closed-ended and that the answers are not quickly retrievable from the internet.

    • Humanity’s Last Exam is a 2,500-question benchmark covering dozens of subjects.
    • The benchmark includes multiple-choice and short-answer questions suitable for automated grading.
    • The authors say each question has a known, unambiguous solution that is not quickly found by internet retrieval.
    • State-of-the-art large language models showed low accuracy and calibration on the benchmark.
    • The study says existing benchmarks are no longer difficult enough to measure top model performance well.
  • AI feedback may not support text revision without process alignment

    What the study found

    The article argues that AI-based feedback can hinder rather than support learners’ revision when it is not aligned with how revision actually works. It identifies three tensions: timing of feedback, learner agency and motivation, and the need to connect revision to meaningful writing goals.

    Why the authors say this matters

    The authors suggest that aligning AI feedback with revision processes is important for supporting learners’ engagement with revising their own texts. They also conclude that the findings have implications for the design of AI feedback tools, writing instruction, and future empirical research.

    What the researchers tested

    The paper uses a theoretical analysis based on process-oriented writing research, which treats text revision as a sequence of interrelated sub-processes with cognitive, motivational, and strategic demands. It then analyzes two AI-based feedback tools, Khan Academy Writing Coach and FelloFish, to assess how their feedback practices fit those demands.

    What worked and what didn't

    The analysis found that the tools did not fully align with learners’ needs in three areas: feedback timing versus the need for critical distance from one’s own text, possible loss of agency and motivation when revision tasks are outsourced to AI, and weak connection between revision and broader writing purposes. The abstract does not report experimental performance outcomes or user trial results.

    What to keep in mind

    This is an analytical article, not a report of a classroom experiment or user study. The abstract does not describe sample size, participant characteristics, or empirical measurements, and limitations are not otherwise detailed in the available summary.

    • The article argues that AI-based feedback may hinder revision if it is not aligned with the revision process.
    • Three tensions are identified: timing, learner agency and motivation, and connection to meaningful writing goals.
    • The paper analyzes Khan Academy Writing Coach and FelloFish.
    • Revision is described as a sequence of interrelated sub-processes with cognitive, motivational, and strategic demands.
    • The abstract says implications are discussed for AI tool design, writing instruction, and future research.
  • Protocol compares two systems for Belgian ILI surveillance

    What the study found

    The article is a study protocol, so it does not report final findings. It describes a planned comparison of two surveillance systems for influenza-like illness (ILI, a flu-like syndrome) in Belgian general practices.

    Why the authors say this matters

    The authors say the work could help identify the most suitable alternative for effective and long-term ILI surveillance. They also conclude that the protocol could serve as a basis for validating other syndromic surveillance data from extraction-based systems in primary care.

    What the researchers tested

    The researchers are carrying out an observational retrospective study covering three influenza seasons from 2021 to 2024. They are comparing the code-based COVID-19 Barometer in General Practices, which extracts data from electronic medical records, with the questionnaire-based Belgian Sentinel General Practitioners network.

    What worked and what didn't

    The protocol says both qualitative and quantitative measures will be used to assess nine attributes: data quality, ILI incidence, sensitivity, representativeness, timeliness, acceptability, simplicity, stability, and flexibility. The study will use CDC surveillance evaluation guidelines and the Simple Multi-Attribute Rating Technique, with experts scoring and weighting three alternatives; the alternative with the higher endorsement will be considered preferable.

    What to keep in mind

    No final results are reported in the abstract because this is a protocol. The abstract does not describe limitations beyond the fact that the comparison is being planned and evaluated across the specified influenza seasons.

    • The article is a protocol, not a results paper.
    • It compares a code-based electronic medical record surveillance tool with a questionnaire-based sentinel network.
    • The focus is influenza-like illness surveillance in Belgian general practices.
    • The study uses nine evaluation attributes, including timeliness, sensitivity, and data quality.
    • Experts will score and weight three alternatives using a multi-criteria decision method.
  • Complex chromosome 6 inversions produced viable recombinant offspring

    What the study found

    The study found that a complex paracentric inversion on chromosome 6q, a type of chromosome rearrangement confined to one chromosome arm, can lead to viable recombinant chromosomes with unbalanced gains and losses. In this family, the rearrangement was transmitted across multiple generations and was associated with five affected children.

    Why the authors say this matters

    The authors conclude that these findings challenge the assumption that paracentric inversions rarely produce viable recombinant chromosomes. They also say the study underscores the importance of high-resolution genomic technologies for accurate diagnosis, understanding the molecular mechanism, and reproductive risk assessment in carriers of complex chromosomal rearrangements.

    What the researchers tested

    The researchers examined a familial intrachromosomal rearrangement involving chromosome 6q across multiple generations. They used high-resolution optical genome mapping, because conventional cytogenetic methods such as G-banded karyotype, FISH (fluorescence in situ hybridization), and chromosomal microarray were said to lack enough resolution to determine the structure.

    What worked and what didn't

    High-resolution optical genome mapping revealed a roughly 75 Mb balanced complex chromosomal rearrangement in the parent. The rearrangement appeared to consist of multiple sequential paracentric inversions and included a single roughly 13 Mb segment in the correct orientation; biased meiotic recombination within that segment was associated with recurrent unbalanced products in the children.

    What to keep in mind

    The abstract describes one family, so the findings are based on a specific familial case rather than a broad population study. The abstract does not provide additional limitations beyond the note that standard cytogenetic methods lacked sufficient resolution.

    • A complex paracentric inversion on chromosome 6q was transmitted across multiple generations.
    • Five children in the family had recombinant chromosomes with reciprocal interstitial gains and losses on 6q.
    • Optical genome mapping identified a roughly 75 Mb balanced complex rearrangement made of multiple sequential inversions.
    • A roughly 13 Mb correctly oriented segment was linked to biased meiotic recombination.
    • The authors say the case challenges the idea that paracentric inversions rarely yield viable recombinant chromosomes.
  • Guideline recommendations on self-harm tools rest on limited evidence

    What the study found

    The authors argue that the NICE guideline on self-harm made definitive recommendations against using risk assessment tools to predict repeat self-harm or suicide, even though the evidence base was very limited. They say the guideline also did not adequately acknowledge uncertainty about model impact, acceptability, and feasibility.

    Why the authors say this matters

    The authors conclude that future updates to the guideline should be informed by higher-quality evidence. They also say that well-developed and validated prediction models could have the potential to improve clinical care for people who self-harm, but should not be used in practice before adequate validation and impact assessment.

    What the researchers tested

    This perspective article examines the NICE guideline development process for self-harm and suicide prevention. The authors discuss the guideline's evidence review, the committee's role in drawing conclusions, and newer evidence on implementation and cost-effectiveness of prediction models.

    What worked and what didn't

    According to the authors, the NICE evidence review included very little evidence, which limited the basis for the recommendations. They say the recommendations relied almost entirely on committee expertise and experience, while evidence on model impact, acceptability, and feasibility was not fully addressed. The authors also note new evidence since the 2022 guideline, including international work on implementation and cost-effectiveness.

    What to keep in mind

    This is a perspective article, not a new primary study of patient outcomes. The abstract does not describe a formal new experiment or provide detailed data on effect sizes, and it does not specify the full scope of the newer evidence it mentions.

    • The authors say the NICE self-harm guideline advised against using risk assessment tools to predict repeat self-harm or suicide.
    • They argue the guideline's evidence review contained very little evidence.
    • They say the recommendations were based largely on committee expertise and experience.
    • They highlight missing discussion of model impact, acceptability, and feasibility.
    • They note newer evidence on implementation and cost-effectiveness has appeared since the 2022 guideline.
  • China study finds partial electricity-carbon price coupling

    What the study found

    The study found that China’s electricity and carbon markets are not fully coupled, but carbon prices do transmit to generator-side electricity tariffs. It also found evidence that carbon pricing can influence corporate energy transition, while several systemic barriers still limit how completely costs are reflected in prices.

    Why the authors say this matters

    The authors say this matters because, under China’s "dual-carbon" goal, the carbon market is meant to help guide the power sector toward a cleaner transition through price signals. The study suggests that improving market design and the link between policy and pricing is needed so carbon costs are reflected more transparently and efficiently.

    What the researchers tested

    The researchers used provincial data from 2013 to 2023 to examine the coupling mechanism between electricity and carbon markets, the transmission of carbon prices, and the incentive effect of carbon pricing. They estimated transmission efficiency, used rolling regression to track changes over time, built a market ecosystem overview, and examined the case of Huaneng International Group.

    What worked and what didn't

    They quantified the transmission efficiency of carbon prices to generator-side electricity tariffs at 0.765. Rolling regression showed a dynamic pass-through effect that was temporarily weakened during major institutional transitions. In the case study, carbon costs were associated with a 37% reduction in carbon intensity and a clean energy share of 31.24%, but undeducted CCERs and the carbon price’s "tidal effect" were identified as barriers to full pass-through.

    What to keep in mind

    The abstract describes one case study and provincial-level analysis in China, so the findings are specific to that context. It also notes limitations in the current market design, including undeducted CCERs, distorted grid emission factors, and incomplete cost pass-through, but it does not describe additional study limitations.

    • Provincial data from 2013 to 2023 showed partial coupling between China’s electricity and carbon markets.
    • The estimated transmission efficiency of carbon prices to generator-side electricity tariffs was 0.765.
    • Pass-through effects weakened temporarily during major institutional transitions.
    • A case study of Huaneng International Group linked carbon costs with a 37% drop in carbon intensity and a clean energy share of 31.24%.
    • Undeducted CCERs and the carbon price’s "tidal effect" were identified as barriers to full price transmission.
  • Thermal imagery did not improve aerial culling rates overall

    What the study found

    Thermal cameras could detect feral deer under a range of temperatures and canopy densities, but they did not significantly increase overall culling rate compared with unaided visual shooting. The study suggests thermal imagery may be more useful when animal densities are lower and canopy cover is dense.

    Why the authors say this matters

    The authors conclude that selective use of thermal imagery in northern Australia could provide benefits in future culling programs. They also say further assessment is needed to optimize control practices.

    What the researchers tested

    The researchers compared thermal and conventional (unaided visual) approaches during aerial control of feral deer and pigs near Collinsville, Queensland, in a warm tropical climate. They analyzed video recordings from a chital deer control program to compare culling rates, animal detections, and search times.

    What worked and what didn't

    The thermal camera successfully detected animals across a variety of temperatures and canopy densities. However, it did not significantly increase the number of animals removed per hour or per kilometer compared with conventional detection, and there were no consistent differences in search time between the two approaches.

    What to keep in mind

    The available summary does not describe detailed limitations beyond the recommendation for further assessment. The findings are specific to a warm, tropical environment in northern Australia and to the conditions studied in this chital deer control program.

    • Thermal cameras detected animals across a range of temperatures and canopy densities.
    • Overall culling rate was not significantly higher with thermal detection than with conventional visual detection.
    • Search time did not show consistent differences between thermal and conventional runs.
    • The authors suggest thermal imagery may help more when animal density is lower and canopy cover is dense.
    • The study was based on video recordings from an aerial chital deer control program near Collinsville, Queensland.
  • Acoustics research at the Technical University of Denmark began in 1935

    What the study found

    The article says that acoustics research and related activity at the Technical University of Denmark can be traced to 1935. It also says the university later established a high-quality acoustical laboratory in 1966.

    Why the authors say this matters

    The authors indicate that the 1931 building of Danish Broadcasting studios revealed a need for scientifically based knowledge on room acoustics and sound insulation. They also suggest that several early researchers from this work became important for acoustics in Denmark and internationally.

    What the researchers tested

    This is a historical research article based on the development of acoustics at the Technical University of Denmark. It traces events from 1935 onward and identifies key people, laboratories, and later institutional developments.

    What worked and what didn't

    The article reports that P.O. Pedersen started an acoustic research group in 1935 and established a laboratory of sound technology in 1941. It also says teaching in acoustics began in the university buildings, and that the 1966 acoustical laboratory had new facilities of very high quality.

    What to keep in mind

    The available summary is historical and does not provide experimental data or comparative evaluation. It also does not give detailed evidence for the claims beyond the narrative account.

    • Acoustics research at the Technical University of Denmark is traced to 1935.
    • A laboratory of sound technology was established in 1941.
    • An acoustical laboratory with very high-quality facilities was established in 1966.
    • The 1931 Danish Broadcasting studios were described as an acoustical scandal that exposed the need for scientific knowledge on room acoustics and sound insulation.
    • The article names Vilhelm Jordan, Per Brüel, and Fritz Ingerslev as important figures in the later development of acoustics.
  • Mandatory emissions disclosure in AI research is feasible

    What the study found

    The study found that mandatory, uncertainty-aware emissions disclosure for AI training runs is operationally feasible at publication time when venues use tiered requirements and light-touch verification. The authors report that a minimal disclosure template can achieve high coverage with modest added burden.

    Why the authors say this matters

    The authors conclude that their framework offers venues and policymakers decision support for comparing transparency policies without relying on proprietary telemetry or speculative large-scale estimates. They also suggest that near-universal disclosure would enable comparable, reproducible emissions reports.

    What the researchers tested

    The researchers developed a policy-level analytical framework rather than estimating emissions for specific AI models. They modeled disclosure requirements, reviewer and editorial workload, and uncertainty propagation under realistic instrumentation assumptions, and tested tiered venue policies called P0, P1, and P2 using Monte Carlo simulation.

    What worked and what didn't

    A minimal disclosure template requiring hardware, duration, energy or carbon dioxide equivalent, and an emission-factor source achieved high coverage with modest burden: median completion time was about 10.8 minutes, reviewer checklist time was about 1.6 minutes per paper, and P2 editorial audits were about 24.1 minutes per 100 submissions. Coverage rose from about 25% under P0 to about 80% under P1 and P2, and uncertainty intervals could be reported using lightweight assumptions, with median relative half-widths of about 0.33 for location-based and 0.77 for market-based reporting. Under baseline priors, H1-H3 were met, H4b was met, and H4a was narrowly missed.

    What to keep in mind

    The abstract does not describe limitations beyond the modeled assumptions and policy framework. The results are based on simulation and a policy-level model, not on direct measurement of emissions from specific training runs.

    • The paper argues that mandatory emissions disclosure for AI research venues is operationally feasible.
    • A minimal disclosure template can raise coverage to about 80% with modest review burden.
    • The modeled template included hardware, duration, energy or CO2e, and emission-factor source.
    • Uncertainty-aware emissions intervals were reported as usable under lightweight assumptions.
    • The framework compares disclosure policies without using proprietary telemetry or speculative large-number estimates.