What the study found
The study found that TRAVELER is a benchmark for testing temporal reasoning, meaning the ability to interpret time-related references such as explicit dates, implicit references like "yesterday," and vague references like "recently." When evaluated on this benchmark, the large language models tested did well on small event sets and explicit references, but their performance dropped as event sets grew and references became less explicit.
Why the authors say this matters
The authors conclude that TRAVELER helps expose limitations in current large language models' event-temporal reasoning capabilities. They also say the publicly available benchmark can be used to test other models beyond those evaluated in the study.
What the researchers tested
The researchers introduced TRAVELER, a synthetic question-answering benchmark built from past events in a household domain. It contains 3,300 English questions generated automatically with four templates from event sets containing 5 to 100 events, and it includes explicit, implicit, and vague temporal references.
What worked and what didn't
The benchmarked large language models could answer questions over event sets with only a handful of events and with explicit temporal references successfully. Performance clearly deteriorated with larger event sets and when temporal references were less explicit, and the vague question category showed the lowest performance across all models.
What to keep in mind
The benchmark uses synthetic questions from a household domain, so the reported results are tied to that setup. For vague references, ground-truth answers were established through human surveys on Prolific, and the abstract does not describe other limitations.
Key points
- TRAVELER is a benchmark for temporal reasoning over explicit, implicit, and vague time references.
- It includes 3,300 English questions generated from event sets with 5 to 100 events.
- Four tested large language models performed well on small event sets and explicit references.
- Performance dropped as event sets got larger and temporal references became less explicit.
- Vague temporal questions had the lowest performance across all models.
Disclosure
- Research title:
- TRAVELER benchmark reveals weaker LLM temporal reasoning with vague references
- Authors:
- Svenja Kenneweg, Jörg Deigmöller, Philipp Cimiano, Julian Eggert
- Institutions:
- Bielefeld University, Bielefeld University, Honda (Germany), Honda (Germany)
- Publication date:
- 2026-04-22
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.