AI Summary of Scholarly Research

This page presents an AI-generated summary of a published research paper. The original authors did not write or review this article. [See full disclosure ↓]

TRAVELER benchmark reveals weaker LLM temporal reasoning with vague references

Research area:computer-science-aiai-ml

What the study found

The study found that TRAVELER is a benchmark for testing temporal reasoning, meaning the ability to interpret time-related references such as explicit dates, implicit references like "yesterday," and vague references like "recently." When evaluated on this benchmark, the large language models tested did well on small event sets and explicit references, but their performance dropped as event sets grew and references became less explicit.

Why the authors say this matters

The authors conclude that TRAVELER helps expose limitations in current large language models' event-temporal reasoning capabilities. They also say the publicly available benchmark can be used to test other models beyond those evaluated in the study.

What the researchers tested

The researchers introduced TRAVELER, a synthetic question-answering benchmark built from past events in a household domain. It contains 3,300 English questions generated automatically with four templates from event sets containing 5 to 100 events, and it includes explicit, implicit, and vague temporal references.

What worked and what didn't

The benchmarked large language models could answer questions over event sets with only a handful of events and with explicit temporal references successfully. Performance clearly deteriorated with larger event sets and when temporal references were less explicit, and the vague question category showed the lowest performance across all models.

What to keep in mind

The benchmark uses synthetic questions from a household domain, so the reported results are tied to that setup. For vague references, ground-truth answers were established through human surveys on Prolific, and the abstract does not describe other limitations.

Key points

  • TRAVELER is a benchmark for temporal reasoning over explicit, implicit, and vague time references.
  • It includes 3,300 English questions generated from event sets with 5 to 100 events.
  • Four tested large language models performed well on small event sets and explicit references.
  • Performance dropped as event sets got larger and temporal references became less explicit.
  • Vague temporal questions had the lowest performance across all models.

Disclosure

Research title:
TRAVELER benchmark reveals weaker LLM temporal reasoning with vague references
Authors:
Svenja Kenneweg, Jörg Deigmöller, Philipp Cimiano, Julian Eggert
Institutions:
Bielefeld University, Bielefeld University, Honda (Germany), Honda (Germany)
Publication date:
2026-04-22
OpenAlex record:
View
AI provenance: This post was generated by gpt-5.4-mini (OpenAI). The original authors did not write or review this post.