יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LivingArena: האם מודלים גדולים יודעים מה שאחרים לא?

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
חוקרים פיתחו את LivingArena, שיטה להערכת מודלים גדולים באמצעות שאילתות הדדיות. המודלים מציעים שאילתות זה לזה, ומקבלים פרסים על זיהוי חוסרי ידע. השיטה מאפשרת הערכה עצמאית ויעילה של מודלים.
תקציר מקורי באנגליתarXiv:2607.24780v1 Announce Type: new Abstract: Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise.
קרא במקור המקורי