Meituan Technology Team
Jun 25, 2026 · Artificial Intelligence
Meituan LongCat’s VitaBench 2.0: A New Benchmark for Long‑Term Dynamic Agents
VitaBench 2.0, an open‑source benchmark from Meituan LongCat, evaluates large language models on long‑term, dynamic user interactions using 56 realistic users, 819 tasks, over 2 000 evolving preferences across up to 1 580 days, and reveals that even top models struggle with memory, personalization and proactive behavior.
LLM evaluationMeituan LongCatVitaBench 2.0
0 likes · 17 min read
