כתבה
arXiv cs.AI ·
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
תקציר מקורי באנגליתarXiv:2610.01780v3 Announce Type: replace Abstract: An AI companion that talks with someone for months should come to understand them. It should know who they are, remember what they said, and recognize when something from the past matters now. Testing this needs real conversations, but real conversations are private, so existing benchmarks use invented people and invented questions. We release RealCompanion, ten real relationships between people and an AI companion, with 27,218 messages over up to 120 days. For each person, we release the full conversation, a profile, a persona, chat test items, and question test items. Every label points to the messages that support it, and every chat label comes with the reasoning that produced it. The real data shows three things. First, people rarely
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית