כתבה
arXiv cs.AI ·
CogArena: ניתוח מרובד של יכולות קוגניטיביות בדגמי שפה גדולים
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
CogArena הוא תקן חדש לבדיקת יכולות קוגניטיביות של דגמי שפה גדולים. התקן כולל 13 תקנים שונים ומספק תמונה עשירה של יכולות הקוגניציה של דגמי שפה.
תקציר מקורי באנגליתarXiv:2607.24999v1 Announce Type: cross Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית