כתבה
arXiv cs.CL ·
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
תקציר מקורי באנגליתarXiv:2603.29112v2 Announce Type: replace-cross Abstract: We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike traditional RecSys benchmarks that focus on item prediction accuracy, our benchmark evaluates how well LLMs can extract and verify user interests from engagement data. We propose two novel metric families: Interest Groundedness (IG), decomposed into precision and recall components to separately penalize hallucinated interest categories and reward coverage, and Interest Specificity (IS), which assesses the distinctiveness of verified LLM-predicted user profiles. We release a synthetic dataset constructed on real user interactions on a global short-form video pl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית