כתבה
arXiv cs.AI ·
NovGauge: בדיקת חדשנות למודלים GPT
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
NovGauge הוא בנק אבטיפוס לבדיקת חדשנות של מודלים גדולים. הוא מכיל 619 זוגות מאמרים ו-50 קבוצות מאמרים, ומאפשר אבחון דק של יכולת המודלים להעריך חדשנות. המחקר מראה כי מודלים כמו GPT-5.5 עדיין רחוקים מלהיות אמינים בהערכת חדשנות מדעית.
תקציר מקורי באנגליתarXiv:2609.11234v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co-citations. Instances are independently labeled along three dimensions: task, problem, and method, capturing application goals, technical challenges, and solution approaches. We propose a cascading diagnostic pipeline
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית