FEATUREDAI HOT (Curated Pool)· aihot-apiZH11:58 · 06·25
→Meituan LongCat open-sources VitaBench 2.0, a long-horizon dynamic agent benchmark
Meituan's LongCat team open-sourced VitaBench 2.0, a benchmark for testing how well agents model users over long, dynamic real-life scenarios. It includes 56 simulated users, 819 complex tasks, over 2,000 shifting preferences, and 66 executable tools—averaging 2,093 interaction events per user across roughly 1,580 days. Even the top model, Claude-Opus-4.6, barely scored above 0.5 in open-book mode. Thinking mode didn't consistently help on personalization tasks, and all models saw a sharp drop on tasks requiring proactive questions. The benchmark and tools are open-sourced.
#Meituan LongCat#Anthropic Claude Opus 4.6#Open source
why featured
Featured · importance 78 · hook + knowledge + resonance
editor take
Meituan LongCat open-sourced a long-term user modeling benchmark; Claude Opus 4.6 barely hit 0.5 in open-book, and all models tanked on proactive questions.
sharp
This one's worth a click because it turns 'does the agent remember what you said last week' into a measurable test. VitaBench 2.0 simulates 56 virtual users, each with roughly 4 years of interaction history and 2,093 events—tasks cover remembering preferences, asking proactive questions, and calling 66 tools. Even Claude Opus 4.6, the top model, barely averaged above 0.5 in open-book mode. Two details stand out: thinking mode didn't consistently help on personalization tasks—sometimes overthinking drifted away from the user's actual intent—and all models saw a cliff-like drop on tasks requiring proactive questions. Models are still much better at passive answering than at clarifying ambiguity. I'd discount this a bit since it's simulated data, and the gap from real user behavior hasn't been validated. But as an engineering benchmark, it gives teams building memory layers and agent workflows something more business-relevant than needle-in-a-haystack tests.
HKR breakdown
hook ✓knowledge ✓resonance ✓