Tagged articles

VitaBench 2.0

1 articles · Page 1 of 1
Meituan Technology Team
Meituan Technology Team
Jun 25, 2026 · Artificial Intelligence

Meituan LongCat’s VitaBench 2.0: A New Benchmark for Long‑Term Dynamic Agents

VitaBench 2.0, an open‑source benchmark from Meituan LongCat, evaluates large language models on long‑term, dynamic user interactions using 56 realistic users, 819 tasks, over 2 000 evolving preferences across up to 1 580 days, and reveals that even top models struggle with memory, personalization and proactive behavior.

LLM evaluationMeituan LongCatVitaBench 2.0
0 likes · 17 min read
Meituan LongCat’s VitaBench 2.0: A New Benchmark for Long‑Term Dynamic Agents